Expression 2

Expression 2 renders any character from one portrait: the whole scene generated live from the audio, on the device or in the bitHuman cloud.

Renders on the device Your servers bitHuman cloud
wise-pup, the Expression 2 sample avatar

Live in your browser, no account. Up to 3 minutes; allow the microphone when asked.

The avatar renders in the bitHuman cloud, and the conversation runs on bitHuman's servers.

What it is

Expression 2 (expression-2) generates the whole avatar scene live from the audio — expressions, mouth and head movement are synthesized each session, not replayed from a base video. It animates the entire 416x720 frame with no face detector or cropping step, so it works for any character: cartoons, animals, creatures, robots, and people.

At creation the platform trains a small model of your specific identity from one photo. That per-identity model is what serves your sessions, and it is why creation takes a couple of hours.

When to choose it

  • Your character is not a photorealistic human — this is the model for it, and where model: "auto" routes such inputs.
  • You want motion generated from the audio itself, not patched onto a base video.
  • You only have a photo — one image is enough.

For a photorealistic person animated from their own footage, compare Essence 2. The side-by-side is on Models.

Where it runs

WhereExpression 2How
iPhone and iPadYesSwift package, Expression2
MacYesSwift package (macOS 13), CLI, Python
AndroidYesexpression2-android
Linux, no GPUYesCLI, Python
Browser (WebGPU)Yesrender=local
Your serversYesCLI, Python, LiveKit plugin
bitHuman cloudYesweb embed, REST API, LiveKit
Fully offline—Coming later

Complete apps: iOS Expression 2, macOS Expression 2 and Android Expression 2.

The file you download from GET /v1/agent/{code}/model/download?model=expression-2 or bithuman pull <CODE> is labelled <CODE>.imx; .avatar is the legacy extension for the same container. How fast it renders on each device is on performance.

How creation works

Create the agent with POST /v1/agent/generate and model: "expression-2", or add expression-2 to an existing agent with POST /v1/agent/{code}/models.

  • The input is a portrait image, of any subject. Without one, the platform generates a portrait from your prompt first. It also generates the agent’s 10-second idle clip and prepares a voice.
  • Training takes about 2 to 2.5 hours; an identity that needs more work gets more, so up to 4 hours is normal. Poll GET /v1/agent/status/{agent_id} until ready or failed, or wait for the completion email.
  • A run that fails is refunded; a completed creation is not, so a second generate is a second charge (failure modes).

The creation cost is on pricing.

Serving tiers

Every published configuration, including a desktop CPU with no GPU, renders faster than real time (performance). In the bitHuman cloud, the service picks the hardware for each session; to benchmark one tier, see pin a tier for a benchmark.

Idle and speaking behavior

During silences the avatar plays its 10-second idle clip, generated from the identity at creation, looping forward-only without a seam. When speech starts, the engine hands off to generated frames with a per-identity color match, so the two stay visually continuous; idle resumes only after sustained silence, not in pauses inside a sentence. A running session bills talking and idle time alike (pricing).

Speech onset. The engine renders in fixed audio chunks; the moving idle clip covers the start of each reply.

Limits and expectations

  • Output is the full 416×720 scene, playing at 20 frames a second.
  • A clear, frontal, well-lit photo gives the best result. The identity is fixed at creation — to change the face, create a new agent.
  • The first session on a new agent can take longer to connect while its model is provisioned; later sessions reuse it.
  • Before training completes, a launch that requests this model is refused with 409 MODEL_NOT_GENERATED. Once ready, expression-2 appears in the agent’s supported_models.

Next steps