How it works

How bitHuman is built: one portable engine, thin language SDKs on top, and your app on top of that. Push 16 kHz audio in, drain lip-synced frames out, the same way on every platform.

SPEECH IN16 kHz mono audioa microphone, text to speech or a WebRTC trackINSIDE EVERY SDKThe bitHuman enginerenders the avatar on the device, on yourserver or in the browserFRAMES OUTLip-synced videoat the model's own rate, drawn by your apppush audiopull framesSwift · Kotlin · Python · CLI · Web
The engine: speech in, frames out. Your app pushes 16 kHz mono speech into the bitHuman engine and pulls lip-synced frames out. The same engine sits inside the Swift package, the Android SDK, the Python SDK, the CLI and the web embed.

The three layers

bitHuman is one portable engine with thin language bindings on top, and your app on top of that. Every layer reads the same model file and produces the same lip-synced frames, on an iPhone, a Mac, a Linux PC, in a browser or in the bitHuman cloud.

LayerWhat it is
Apps and toolsthe bitHuman CLI, your own app, LiveKit for WebRTC transport
Language SDKsPython, Swift and Kotlin: thin bindings over the same engine; the browser through the web embed
The bitHuman enginethe avatar renderer, inside every SDK, so there is nothing separate to install: macOS, iOS, Android, Linux and the browser

You integrate at the SDK layer. The engine is built into each SDK, so your app needs the bitHuman dependency and nothing else. To pick a platform, start at Platforms; which model runs where is on Models.

What stays true across every surface

  • One model file, every surface. The same audio drives the same lip-sync on every SDK; pixels can differ slightly between hardware backends.
  • A stable public API. Deprecated options keep working with a warning until the next major, and majors call out breaks explicitly.
  • Surfaces mix. The Swift package in your iOS app with the Python package on your backend is supported; keep each one current — Downloads lists the current versions.
  • One credential. The same key drives every surface; how it is exchanged and billed is on Authentication and pricing.

Audio in, frames out

Every SDK has the same shape — audio in, video out:

  1. Push audio as it arrives — a microphone, TTS, a WebRTC track.
  2. Drain lip-synced frames at the model’s own rate — 25 fps for Essence 2 and Essence 1, 20 fps for Expression 2.

The engine buffers between the two, so your audio source and your render loop never have to run in lockstep.

In Python

render() takes audio a chunk at a time and yields frames as it goes, so a stream and a file are the same program:

import numpy as np
import bithuman

def chunks(path="speech16k.raw", ms=40):
    """16 kHz mono int16 audio, delivered a chunk at a time — a mic, TTS or a socket."""
    pcm = np.fromfile(path, dtype=np.int16)
    step = 16000 * ms // 1000
    for i in range(0, len(pcm), step):
        yield pcm[i:i + step]

frames = 0
with bithuman.open("wise-pup.imx") as avatar:
    for image in avatar.render(chunks()):      # frames come out as audio goes in
        frames += 1                            # image: (height, width, 3) uint8, RGB
print(frames, image.shape)

Install, the model download and the credential are on the Python SDK page. speech16k.raw is any speech converted with ffmpeg -i speech.wav -ac 1 -ar 16000 -f s16le speech16k.raw.

Audio format

PropertyValue
Encoding16-bit signed PCM (int16), or float32 in [-1, 1]
Channelsmono
Sample rate16 kHz for decoded samples; a file path in any format ffmpeg reads is converted for you
Chunk sizeanything; 10–40 ms is typical

Frame format

Frames arrive at the model’s own rate, whatever the chunk size: 25 fps for Essence 2 (up to 1080p: the identity’s own canvas, 1080×1920 portrait for a standard identity) and 20 fps for Expression 2 (416x720). Python yields RGB uint8 arrays; the Swift package and the Android SDK hand you their platform’s image types.

In the other SDKs

  • Apple — feed() PCM, then pull() frames. See the Swift package.
  • Android — feed() PCM, then pull() into a reused buffer: a Bitmap for Expression 2, an RGBA ByteBuffer for Essence 2. See the Android SDK.
  • CLI — bithuman render takes an audio file; bithuman run streams a live conversation. See the CLI.