Essence 2

Essence 2 renders a photoreal person from one portrait: the identity's own footage, lip-synced live, on the device or in the bitHuman cloud.

Renders on the device Your servers bitHuman cloud
sofia-ramirez, the Essence 2 sample avatar

Live in your browser, no account. Up to 3 minutes; allow the microphone when asked.

The avatar renders in the bitHuman cloud, and the conversation runs on bitHuman's servers.

What it is

Essence 2 (essence-2) renders a photoreal person from one portrait, up to 1080p: the identity’s own canvas, 1080×1920 portrait for a standard identity. From your portrait the platform generates a 10-second identity video; the model then animates lip-sync and expression over it live, with a sharp mouth and teeth taken from that video.

When to choose it

  • A photorealistic person — start here.
  • Always-on displays — kiosks, lobby screens and 24/7 assistants.
  • On your own hardware — every SDK platform runs it.

For a stylized character, or a scene generated from one photo, choose Expression 2. The side-by-side is on Models.

Where it runs

WhereEssence 2How
iPhone and iPadYesSwift package, Essence2Kit (iOS 26)
MacYesSwift package (macOS 26, M3 or newer), CLI, Python
AndroidYesessence2-android
Linux, no GPUYesCLI, Python
Browser (WebGPU)Yesrender=local, for identities with a browser build
Your serversYesCLI, Python, LiveKit plugin
bitHuman cloudYesweb embed, REST API, LiveKit
Fully offline—Coming later

A complete app for iPhone and iPad is the iOS Essence 2 example; for Android, the Android Essence 2 example.

The file you download is <CODE>.imx, from GET /v1/agent/{code}/model/download?model=essence-2 or bithuman pull <CODE> --model essence-2. How fast it renders on each device is on performance.

How creation works

Create the agent with POST /v1/agent/generate and model: "essence-2", or add essence-2 to an existing agent with POST /v1/agent/{code}/models.

  • The input is a portrait image of a photorealistic human. A stylized or non-human input is refused with 422 MODEL_SUBJECT_MISMATCH before anything is billed; model: "auto" routes it to Expression 2 instead.
  • The platform generates the identity video from the image, then trains the identity. Poll GET /v1/agent/status/{agent_id} until ready; allow about 2 to 2.5 hours.
  • ready serves before it downloads. The downloadable file is published a little later; until then the download endpoint answers a retryable 404 MODEL_ARTIFACT_NOT_READY.

The creation cost is on pricing.

Serving tiers

In the bitHuman cloud, the service picks the hardware for each session. To benchmark one tier, see pin a tier for a benchmark; in production, let the service choose.

Idle and speaking behavior

The identity video plays continuously and loops forward-only: at its last frame it wraps to the first, and it never plays in reverse. While idle it is pure playback of your footage; while talking, the animated face is rendered over the same frames. A running session bills talking and idle time alike (pricing).

Limits and expectations

  • Output plays at 25 frames a second everywhere it runs. How fast a platform renders is on performance.
  • The downloadable file is about 140–160 MB, varying per identity — read Content-Length rather than assuming a size.
  • The identity is fixed at creation. To change the face, create a new agent.
  • The first session on a new agent can take longer to connect while the identity is provisioned; later sessions reuse it.
  • Before training completes, a launch that requests this model is refused with 409 MODEL_NOT_GENERATED. Once ready, essence-2 appears in the agent’s supported_models.

Next steps