Python

Render Essence 2 and Expression 2 avatars from Python: open an avatar, push audio, get frames, on macOS (Apple silicon) and Linux, where it needs no GPU.

Your servers No GPU API secret Python 2.11.16

The bithuman package renders avatars in your own Python code on your own machine: a file in and frames out, or a live stream of audio in and frames out. To run an avatar without code, use the CLI.

Note: On Linux, Python renders both models on the CPU alone, no GPU. On macOS it renders on Apple silicon. See CPU only (no GPU).

DetailExpression 2Essence 2
Rendersany character from one portraita photoreal person from one portrait
Installpip install "bithuman[expression-2]"included in the same install
FramesRGB numpy arrays, (height, width, 3) uint8the same
Captured on an Apple M4 Mac (macOS) · Python SDK 2.11.6 · sofia-ramirez (Essence 2) · 2026-09-23.

Before you start

You needCheck
Python 3.10–3.14python3 --version
macOS 14+ on Apple silicon, Linux x86_64 or Linux arm64python3 -c "import platform; print(platform.system(), platform.machine())"
An API secretYour API secret
About 1 GB of disk (570 MB package, 118–190 MB per avatar)df -h .
ffmpeg on PATH, for MP4 output onlyffmpeg -version

Install

python3 -m venv .venv
source .venv/bin/activate
pip install "bithuman[expression-2]"

Install into a virtual environment: Debian and Ubuntu refuse a system-wide pip install (externally-managed-environment). If venv is missing, run sudo apt install -y python3-venv first. The package installs no command-line tool.

Authenticate

Set BITHUMAN_API_SECRET in the shell that runs Python (bithuman.open reads it), or pass api_secret= to AsyncBithuman.create(). See Your API secret. Credits pay for session time, talking or idle, by the exact second (pricing). Downloading a sample avatar needs no account.

First frame

export BITHUMAN_API_SECRET="<your API secret>"
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
import bithuman

with bithuman.open("wise-pup.imx") as avatar:
    frames = [image for image in avatar.render("speech.wav")]
print(len(frames), "frames of", frames[0].shape)
# → 300 frames of (720, 416, 3)

render takes a path to any audio file ffmpeg reads, or already-decoded 16 kHz mono audio (int16 or float32 arrays, or raw 16-bit bytes). Frames are RGB; OpenCV expects BGR, so write one with cv2.imwrite("frame.png", image[:, :, ::-1]). The same call opens Essence 2 and Essence 1 .imx files.

To write an MP4 instead, pass out_mp4= to the same render (any model, needs ffmpeg); it returns the number of frames written. Download the sofia-ramirez Essence 2 sample first:

curl -fL -o sofia-ramirez.imx "https://api.bithuman.ai/v1/agent/A52DHS2219/model/download?model=essence-2"
import bithuman
bithuman.open("sofia-ramirez.imx").render("speech.wav", out_mp4="out.mp4")
# → out.mp4: 1080×1920 with the speech, 15.2 s

Complete example

The quickstart from the examples repository: open an avatar and watch it speak in a window.

Requirements

You needNotes
Python 3.10–3.14on macOS (Apple silicon) or Linux (x86_64, arm64)
An API secret
A desktop sessionthe example opens a window

Run it

git clone https://github.com/bithuman-product/bithuman-examples.git
cd bithuman-examples/python/quickstart
python3 -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt
export BITHUMAN_API_SECRET="<your API secret>"
python local-avatar.py

Expected output

A window titled bitHuman avatar opens and the avatar speaks the bundled speech.wav. The first run downloads the sample avatar (about 150 MB). To write an MP4 instead of opening a window, run python -m bithuman ~/.cache/bithuman/models/A52DHS2219.imx speech.wav.

Make it your own

  • Your own avatar: pass --model with your agent’s .imx, downloaded with the Agents API or bithuman pull <AGENT_CODE>.
  • Your own audio: pass any audio file; or stream microphone audio with AsyncBithuman (Integrate into your app).
  • A conversation: cloud-avatar.py in the same folder connects the avatar to an OpenAI voice agent over LiveKit (LiveKit).
  • A desktop companion: conversation.py in the same folder talks with you through OpenAI Realtime, with your OpenAI key (Companion app).
  • A web app: send the frames from render() to your own video stream, or use the web embed.

Integrate into your app

For a live conversation, AsyncBithuman takes audio as it arrives and yields frames and audio at the model’s rate:

# excerpt: show() and play() are your own display and audio output
import asyncio, soundfile as sf
from bithuman import AsyncBithuman

async def main():
    avatar = await AsyncBithuman.create(model_path="wise-pup.imx")   # reads BITHUMAN_API_SECRET
    pcm, rate = sf.read("speech.wav", dtype="int16")

    async def speak():
        for i in range(0, len(pcm), rate // 10):                     # 100 ms chunks, as they arrive
            await avatar.push_audio(pcm[i:i + rate // 10].tobytes(), rate, last_chunk=False)
        await avatar.flush()                                          # end of the reply

    task = asyncio.create_task(speak())
    try:
        async for frame in avatar.run():                              # paced at the model's play rate
            if frame.has_image:
                show(frame.bgr_image)                                 # BGR numpy array
            if frame.audio_chunk:
                play(frame.audio_chunk.array)                         # audio in sync with the frame
    finally:
        task.cancel()
        await avatar.shutdown()                                       # frees the model and the credential

asyncio.run(main())
JobCall
Stream audioawait avatar.push_audio(int16_bytes, sample_rate, last_chunk=False)
End of a replyawait avatar.flush()
Interrupt the replyavatar.interrupt()
Idle between replieskeep reading run(): it yields idle frames when there is no speech
Stopawait avatar.shutdown() in a finally (stop() keeps the model loaded)

A voice agent on your own LiveKit server

The LiveKit plugin runs AsyncBithuman inside a LiveKit Agents worker: pass model_path and the avatar renders in the worker’s own process, next to an OpenAI Realtime voice.

# excerpt: python/self-host/agent.py (bithuman-examples)
session = AgentSession(llm=openai.realtime.RealtimeModel(model="gpt-realtime-2.1-mini", voice="coral",
    turn_detection=ServerVad(type="server_vad", silence_duration_ms=500)))   # end of turn after 0.5 s of silence
avatar = bithuman.AvatarSession(model_path="wise-pup.imx",    # renders here
                                api_secret=os.environ["BITHUMAN_MASTER_SECRET"])
await avatar.start(session, room=ctx.room)
await session.start(agent=Agent(instructions="You are a friendly assistant."),
                    room=ctx.room, room_options=RoomOptions(audio_output=False))

In a LiveKit worker, name the secret BITHUMAN_MASTER_SECRET and pass it explicitly (LiveKit). The runnable example with livekit-server --dev and a browser link: Talk to an avatar on your machine.

Platform notes

  • The first Essence 2 render downloads a shared audio encoder (about 377 MB, plus about 70 MB for streaming) to ~/.bithuman/deps, once.
  • BITHUMAN_CACHE_DIR moves the download cache from ~/.cache/bithuman.
  • python -m bithuman render <AGENT_CODE> <audio> downloads your own agent’s model by code and renders it.
  • A process with no API secret opens the file and refuses at the first frame; a rejected secret refuses at open.

Performance

ConfigurationEssence 2Expression 2
Intel Core i7-13700F (x86_64) Linux · Python CPU only (no GPU)
1.9× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26
2.3× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26
Apple M4 macOS · Python
6.9× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25
8.4× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25

Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.

A finished render logs its own rate on the bithuman logger at INFO.

Troubleshooting

SymptomCauseFix
error: externally-managed-environmentpip targeted the system Pythoncreate and activate a venv
ModuleNotFoundError: No module named 'bithuman'the venv is not active in this terminalsource .venv/bin/activate
pip finds no wheelIntel Mac, Windows, musl, or Python outside 3.10–3.14use a supported platform (WSL2 on Windows)
NotSupported opening an Expression 2 filethe extra is missingpip install "bithuman[expression-2]"
NotAuthorised at the first frame: no credential was suppliedno secret in this shellexport BITHUMAN_API_SECRET=…
NotAuthorised at open: that key was not accepted (401)the secret was rejectedcreate a new one under API secrets
NotAuthorised: no API secret was found from bithuman.openno secretset BITHUMAN_API_SECRET; nothing is rendered or written
An MP4 with sound and no picturea refused render through the deprecated render_offline leaves the audio trackrender with bithuman.open(path).render(audio, out_mp4=...), which refuses before it writes anything; check the frame count it returns
Frames look blueframes are RGB and your display wants BGRimage[:, :, ::-1]
Raw audio plays slow and longdecoded audio must be 16 kHz monopass a file path, or resample to 16 kHz
404 NOT_FOUND downloading a modelnot your agent and not a sample avatarcheck the code under your agents
The example window never opens (GUI: NONE)the headless OpenCV build won the installpip install --force-reinstall --no-deps opencv-python

Reference