Essence 2
More ▾
Essence 2 renders a photoreal person from one portrait: the identity's own footage, lip-synced live, on the device or in the bitHuman cloud.
The avatar renders in the bitHuman cloud, and the conversation runs on bitHuman's servers.
What it is
Essence 2 (essence-2) renders a photoreal person from one portrait, up to
1080p: the identity’s own canvas, 1080×1920 portrait for a standard identity. From your portrait the platform generates a 10-second
identity video; the model then animates lip-sync and expression over it live,
with a sharp mouth and teeth taken from that video.
When to choose it
- A photorealistic person — start here.
- Always-on displays — kiosks, lobby screens and 24/7 assistants.
- On your own hardware — every SDK platform runs it.
For a stylized character, or a scene generated from one photo, choose Expression 2. The side-by-side is on Models.
Where it runs
| Where | Essence 2 | How |
|---|---|---|
| iPhone and iPad | Yes | Swift package, Essence2Kit (iOS 26) |
| Mac | Yes | Swift package (macOS 26, M3 or newer), CLI, Python |
| Android | Yes | essence2-android |
| Linux, no GPU | Yes | CLI, Python |
| Browser (WebGPU) | Yes | render=local, for identities with a browser build |
| Your servers | Yes | CLI, Python, LiveKit plugin |
| bitHuman cloud | Yes | web embed, REST API, LiveKit |
| Fully offline | — | Coming later |
A complete app for iPhone and iPad is the iOS Essence 2 example; for Android, the Android Essence 2 example.
The file you download is <CODE>.imx, from
GET /v1/agent/{code}/model/download?model=essence-2
or bithuman pull <CODE> --model essence-2. How fast it renders on each device
is on performance.
How creation works
Create the agent with POST /v1/agent/generate
and model: "essence-2", or add essence-2 to an existing agent with
POST /v1/agent/{code}/models.
- The input is a portrait image of a photorealistic human. A stylized or
non-human input is refused with
422 MODEL_SUBJECT_MISMATCHbefore anything is billed;model: "auto"routes it to Expression 2 instead. - The platform generates the identity video from the image, then trains the
identity. Poll
GET /v1/agent/status/{agent_id}untilready; allow about 2 to 2.5 hours. readyserves before it downloads. The downloadable file is published a little later; until then the download endpoint answers a retryable404 MODEL_ARTIFACT_NOT_READY.
The creation cost is on pricing.
Serving tiers
In the bitHuman cloud, the service picks the hardware for each session. To benchmark one tier, see pin a tier for a benchmark; in production, let the service choose.
Idle and speaking behavior
The identity video plays continuously and loops forward-only: at its last frame it wraps to the first, and it never plays in reverse. While idle it is pure playback of your footage; while talking, the animated face is rendered over the same frames. A running session bills talking and idle time alike (pricing).
Limits and expectations
- Output plays at 25 frames a second everywhere it runs. How fast a platform renders is on performance.
- The downloadable file is about 140–160 MB, varying per identity — read
Content-Lengthrather than assuming a size. - The identity is fixed at creation. To change the face, create a new agent.
- The first session on a new agent can take longer to connect while the identity is provisioned; later sessions reuse it.
- Before training completes, a launch that requests this model is refused
with
409 MODEL_NOT_GENERATED. Once ready,essence-2appears in the agent’ssupported_models.
Next steps
- Models — the four models side by side
- Agents API — create, poll, download
- Embed widget — a live session in minutes
- Video API — render an MP4 with
model: "essence-2" - Session behavior & troubleshooting