Talking video

Turn one audio file into an MP4 of your avatar speaking it: with the CLI or Python on your own machine, CPU only on Linux, or with one REST call to the bitHuman cloud.

5 min Creator plan or higher No GPU Your servers bitHuman cloud API secret CLI 2.8.3 Python 2.11.16

What you’ll build

An MP4 of an avatar saying the words in an audio file, with the lips in sync. Nobody watches it live, so it suits greetings, lessons and product videos.

You need:

  • an audio file of speech, or a script for the REST API;
  • an API secret on the Creator plan or higher;
  • for the CLI or Python: Linux (x86_64 or arm64) or a Mac with Apple silicon, and ffmpeg.
Captured on a Linux PC, Intel Core i7-13700F (Ubuntu 24.04) · CLI 2.8.2 (bithuman render) · kwame-warm-museum-guide (Essence 2) · 2026-09-27. Rendered on the CPU alone: the PC's graphics card was hidden from the process.

Steps

3 steps

  1. Get the speech

    Any audio file ffmpeg reads works on your machine; the REST API takes a file at a public URL, or text in the agent’s voice. This sample is 15 seconds of speech:

    curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
    Expected

    speech.wav in the current directory.

  2. Render the video

    Pick where it renders. The CLI and Python render on your own machine, and on Linux they need no GPU. The REST API renders in the bitHuman cloud.

    CLI

    curl -fsSL https://install.bithuman.ai | sh
    export BITHUMAN_API_SECRET="<your API secret>"
    bithuman render kwame-warm-museum-guide speech.wav -o kwame.mp4

    Python

    # Needs the bithuman package, ffmpeg on PATH and BITHUMAN_API_SECRET in the environment.
    # The avatar file comes from `bithuman pull sofia-ramirez`, or the download URL on /platforms/python.
    import bithuman
    frames = bithuman.open("sofia-ramirez.imx").render("speech.wav", out_mp4="out.mp4")
    print(frames, "frames written")

    REST

    curl -X POST https://api.bithuman.ai/v1/video/generate \
      -H "api-secret: $BITHUMAN_API_SECRET" -H "Content-Type: application/json" \
      -d '{"model": "essence-2", "agent_code": "'"$BITHUMAN_AGENT_CODE"'", "input": {"type": "audio", "audio_url": "https://docs.bithuman.ai/samples/speech.wav"}, "wait": true}'
    Expected

    An MP4 as long as the speech: kwame.mp4, or out.mp4 with the number of frames written, on your machine; or a video_url in the REST response ("status": "completed"). A render that takes longer than the wait returns a job_id to poll.

  3. Check the file

    ffprobe -v error -show_entries stream=width,height -show_entries format=duration -of compact kwame.mp4
    Expected

    An Essence 2 video is up to 1080×1920 or 1920×1080, following the portrait; an Expression 2 video is 416×720. The duration matches the speech, and the file has one video and one audio track.

How it works

SPEECH IN16 kHz mono audioa microphone, text to speech or a WebRTC trackINSIDE EVERY SDKThe bitHuman enginerenders the avatar on the device, on yourserver or in the browserFRAMES OUTLip-synced videoat the model's own rate, drawn by your apppush audiopull framesSwift · Kotlin · Python · CLI · Web
The engine: speech in, frames out. Your app pushes 16 kHz mono speech into the bitHuman engine and pulls lip-synced frames out. The same engine sits inside the Swift package, the Android SDK, the Python SDK, the CLI and the web embed.

The same engine renders a live conversation and a file: speech goes in and lip-synced frames come out, then ffmpeg writes them with the audio into an MP4. A file rendered on your machine keeps its audio and video there; the CLI and Python sign in online when the render starts. Through the REST API, the render runs in the bitHuman cloud, in the US.

Make it your own

  • Your avatar: create one from a portrait, then pass its agent code to bithuman render, or to the REST call as agent_code.
  • Text instead of audio: the REST API speaks a script in the agent’s voice: "input": {"type": "text", "text": "…"} (Talking video API).
  • A batch: loop over files with bithuman render; every run reuses the downloaded avatar.
  • What it costs: on your machine, file rendering bills the length of the video it writes at the self-hosted rate; the REST API bills per minute of output, rounded up (Pricing).

Troubleshooting

SymptomFix
bithuman render exits with code 77No API secret: export BITHUMAN_API_SECRET, or run bithuman login once.
The render is refused before it startsSet BITHUMAN_API_SECRET for an account on the Creator plan or higher, then render again.
402 INSUFFICIENT_BALANCE from the REST APIYour balance must cover the longest render up front; the difference is refunded when it finishes (Talking video API).
409 MODEL_NOT_GENERATEDThe agent has no model of that kind yet: check its supported_models, or add the model.