How it works
More ▾
How bitHuman is built: one portable engine, thin language SDKs on top, and your app on top of that. Push 16 kHz audio in, drain lip-synced frames out, the same way on every platform.
The three layers
bitHuman is one portable engine with thin language bindings on top, and your app on top of that. Every layer reads the same model file and produces the same lip-synced frames, on an iPhone, a Mac, a Linux PC, in a browser or in the bitHuman cloud.
| Layer | What it is |
|---|---|
| Apps and tools | the bitHuman CLI, your own app, LiveKit for WebRTC transport |
| Language SDKs | Python, Swift and Kotlin: thin bindings over the same engine; the browser through the web embed |
| The bitHuman engine | the avatar renderer, inside every SDK, so there is nothing separate to install: macOS, iOS, Android, Linux and the browser |
You integrate at the SDK layer. The engine is built into each SDK, so your app needs the bitHuman dependency and nothing else. To pick a platform, start at Platforms; which model runs where is on Models.
What stays true across every surface
- One model file, every surface. The same audio drives the same lip-sync on every SDK; pixels can differ slightly between hardware backends.
- A stable public API. Deprecated options keep working with a warning until the next major, and majors call out breaks explicitly.
- Surfaces mix. The Swift package in your iOS app with the Python package on your backend is supported; keep each one current — Downloads lists the current versions.
- One credential. The same key drives every surface; how it is exchanged and billed is on Authentication and pricing.
Audio in, frames out
Every SDK has the same shape — audio in, video out:
- Push audio as it arrives — a microphone, TTS, a WebRTC track.
- Drain lip-synced frames at the model’s own rate — 25 fps for Essence 2 and Essence 1, 20 fps for Expression 2.
The engine buffers between the two, so your audio source and your render loop never have to run in lockstep.
In Python
render() takes audio a chunk at a time and yields frames as it goes, so a
stream and a file are the same program:
import numpy as np
import bithuman
def chunks(path="speech16k.raw", ms=40):
"""16 kHz mono int16 audio, delivered a chunk at a time — a mic, TTS or a socket."""
pcm = np.fromfile(path, dtype=np.int16)
step = 16000 * ms // 1000
for i in range(0, len(pcm), step):
yield pcm[i:i + step]
frames = 0
with bithuman.open("wise-pup.imx") as avatar:
for image in avatar.render(chunks()): # frames come out as audio goes in
frames += 1 # image: (height, width, 3) uint8, RGB
print(frames, image.shape)
Install, the model download and the credential are on the
Python SDK page. speech16k.raw is any speech converted with
ffmpeg -i speech.wav -ac 1 -ar 16000 -f s16le speech16k.raw.
Audio format
| Property | Value |
|---|---|
| Encoding | 16-bit signed PCM (int16), or float32 in [-1, 1] |
| Channels | mono |
| Sample rate | 16 kHz for decoded samples; a file path in any format ffmpeg reads is converted for you |
| Chunk size | anything; 10–40 ms is typical |
Frame format
Frames arrive at the model’s own rate, whatever the chunk size: 25 fps for
Essence 2 (up to 1080p: the identity’s own canvas, 1080×1920 portrait for a standard identity) and 20 fps for Expression 2
(416x720). Python yields RGB uint8 arrays; the Swift package and the Android SDK hand you
their platform’s image types.
In the other SDKs
- Apple —
feed()PCM, thenpull()frames. See the Swift package. - Android —
feed()PCM, thenpull()into a reused buffer: aBitmapfor Expression 2, an RGBAByteBufferfor Essence 2. See the Android SDK. - CLI —
bithuman rendertakes an audio file;bithuman runstreams a live conversation. See the CLI.