# Talking video

URL: https://docs.bithuman.ai/build/talking-video

> Turn one audio file into an MP4 of your avatar speaking it: with the CLI or Python on your own machine, CPU only on Linux, or with one REST call to the bitHuman cloud.

## What you'll build

An MP4 of an avatar saying the words in an audio file, with the lips in sync. Nobody watches it live, so it suits greetings, lessons and product videos.

You need:

- an audio file of speech, or a script for the REST API;
- an [API secret](https://docs.bithuman.ai/start/api-secret) on the Creator plan or higher;
- for the CLI or Python: Linux (x86_64 or arm64) or a Mac with Apple silicon, and `ffmpeg`.

*Capture: The kwame-warm-museum-guide avatar in a talking video, rendered from one audio file by bithuman render on a Linux PC. Captured on a Linux PC, Intel Core i7-13700F (Ubuntu 24.04) · CLI 2.8.2 (bithuman render) · kwame-warm-museum-guide (Essence 2) · 2026-09-27. Rendered on the CPU alone: the PC's graphics card was hidden from the process.* (https://docs.bithuman.ai/examples/talking-video-linux/clip.mp4)

## Steps

### Get the speech

Any audio file `ffmpeg` reads works on your machine; the REST API takes a file at a public URL, or text in the agent's voice. This sample is 15 seconds of speech:

```bash
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
```

Expected:

`speech.wav` in the current directory.

### Render the video

Pick where it renders. The CLI and Python render on your own machine, and on Linux they need no GPU. The REST API renders in the bitHuman cloud.

```bash tab="CLI"
curl -fsSL https://install.bithuman.ai | sh
export BITHUMAN_API_SECRET="<your API secret>"
bithuman render kwame-warm-museum-guide speech.wav -o kwame.mp4
```

```python tab="Python"
# Needs the bithuman package, ffmpeg on PATH and BITHUMAN_API_SECRET in the environment.
# The avatar file comes from `bithuman pull sofia-ramirez`, or the download URL on /platforms/python.
import bithuman
frames = bithuman.open("sofia-ramirez.imx").render("speech.wav", out_mp4="out.mp4")
print(frames, "frames written")
```

```bash tab="REST"
curl -X POST https://api.bithuman.ai/v1/video/generate \
  -H "api-secret: $BITHUMAN_API_SECRET" -H "Content-Type: application/json" \
  -d '{"model": "essence-2", "agent_code": "'"$BITHUMAN_AGENT_CODE"'", "input": {"type": "audio", "audio_url": "https://docs.bithuman.ai/samples/speech.wav"}, "wait": true}'
```

Expected:

An MP4 as long as the speech: `kwame.mp4`, or `out.mp4` with the number of frames written, on your machine; or a `video_url` in the REST response (`"status": "completed"`). A render that takes longer than the wait returns a `job_id` to [poll](https://docs.bithuman.ai/api/video#get-talking-video-status).

### Check the file

```bash
ffprobe -v error -show_entries stream=width,height -show_entries format=duration -of compact kwame.mp4
```

Expected:

An Essence 2 video is up to 1080×1920 or 1920×1080, following the portrait; an Expression 2 video is 416×720. The duration matches the speech, and the file has one video and one audio track.

## How it works

*Diagram: The engine: speech in, frames out.* Your app pushes 16 kHz mono speech into the bitHuman engine and pulls lip-synced frames out. The same engine sits inside the Swift package, the Android SDK, the Python SDK, the CLI and the web embed.

The same engine renders a live conversation and a file: speech goes in and lip-synced frames come out, then `ffmpeg` writes them with the audio into an MP4. A file rendered on your machine keeps its audio and video there; the CLI and Python sign in online when the render starts. Through the REST API, the render runs in the bitHuman cloud, in the US.

## Make it your own

- **Your avatar:** [create one](https://docs.bithuman.ai/build/create-avatar) from a portrait, then pass its agent code to `bithuman render`, or to the REST call as `agent_code`.
- **Text instead of audio:** the REST API speaks a script in the agent's voice: `"input": {"type": "text", "text": "…"}` ([Talking video API](https://docs.bithuman.ai/api/video#text-input)).
- **A batch:** loop over files with `bithuman render`; every run reuses the downloaded avatar.
- **What it costs:** on your machine, file rendering bills the length of the video it writes at the self-hosted rate; the REST API bills per minute of output, rounded up ([Pricing](https://docs.bithuman.ai/pricing#talking-video--per-minute-of-output)).

## Troubleshooting

| Symptom | Fix |
|---|---|
| `bithuman render` exits with code 77 | No API secret: export `BITHUMAN_API_SECRET`, or run `bithuman login` once. |
| The render is refused before it starts | Set `BITHUMAN_API_SECRET` for an account on the Creator plan or higher, then render again. |
| `402 INSUFFICIENT_BALANCE` from the REST API | Your balance must cover the longest render up front; the difference is refunded when it finishes ([Talking video API](https://docs.bithuman.ai/api/video)). |
| `409 MODEL_NOT_GENERATED` | The agent has no model of that kind yet: check its `supported_models`, or [add the model](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent). |
