What is a real-time avatar API?
Speech in, lip-synced video out: what it does and which model fits.
A real-time avatar API turns speech audio into lip-synced video of a face while the speech plays, so a voice agent, an app or a kiosk can show a talking character. With bitHuman, your code pushes speech into an SDK and pulls video frames out; the avatar renders on the device, in the browser, on your own computer or in the bitHuman cloud.
Pick Essence 2 for a real person and Expression 2 for any other character.
Every published configuration renders faster than real time, including a desktop CPU with no GPU:
| Configuration | Essence 2 | Expression 2 |
|---|---|---|
| NVIDIA RTX 4090 Cloud API · GPU | 4.1× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-27 | 17.0× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-23 |
| Intel Core i7-13700F (x86_64) Linux · CLI CPU only (no GPU) | 2.0× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 | 2.2× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 |
| iPhone 15 iPhone · Swift package | 2.1× real timeiPhone 15 · Swift package 2.17.3 · measured 2026-09-27 | 5.5× real timeiPhone 15 · Swift package 2.18.0 · measured 2026-09-27 |
| Samsung Galaxy S25+ Android | 2.0× real timeSamsung Galaxy S25+ · essence2-android 0.7.0 · measured 2026-09-25 | 2.4× real timeSamsung Galaxy S25+ · expression2-android 0.4.10 · measured 2026-09-23 |
Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.
What it does
- One portrait makes an avatar. You upload a photo or a drawing; the avatar is created in the bitHuman cloud in about 2 to 2.5 hours (Create your own avatar).
- Speech in, frames out. A session takes audio as it arrives and returns frames with the lips on the words; between replies the avatar idles (How it works).
- The conversation is yours or ours. The SDKs render the speech your own voice stack produces. A bitHuman agent runs the whole conversation in the web embed, and with the web embed, the conversation runs on bitHuman’s servers, even when the avatar renders in the tab.
Which model to use
| Your avatar | Model | What it renders | Where it runs |
|---|---|---|---|
| A real person | Essence 2 (essence-2) | a photoreal person, up to 1080×1920 | every SDK platform and the bitHuman cloud |
| Any other character: a cartoon, an animal, a mascot, a robot | Expression 2 (expression-2) | the whole 416×720 frame, with motion generated from the audio | every SDK platform and the bitHuman cloud |
| An existing first-generation agent | Essence 1 or Expression 1 | a real person’s face | Essence 1: the cloud, Python, the CLI and the browser; Expression 1: the cloud only |
Not sure: create with "model": "auto", which sends a real person to Essence 2 and anything else to Expression 2. Each model’s places, side by side: Where each model runs.
Which SDK to use
| You are building | Use | The avatar renders |
|---|---|---|
| A website | the web embed (one iframe) | in the bitHuman cloud, or in the tab with WebGPU |
| An iPhone, iPad or Mac app | the Swift package | on the device |
| An Android app | the Android SDK, or the Flutter plugin | on the device |
| A LiveKit or Pipecat voice agent | the LiveKit plugin or pipecat-bithuman | in the bitHuman cloud, or on your server |
| A backend or a script | the Python SDK or the REST API | on your machine, or in the bitHuman cloud |
| A video file, no code | the CLI: bithuman render | on your Mac or Linux machine |
Adding a face to an existing voice agent: Add a talking avatar to a voice agent.
First call
The quickest way to see one is the wise-pup sample in the web embed. It needs no account or install:
<iframe src="https://www.bithuman.ai/embed/A23WJF0199" allow="microphone *" style="width:100%;height:600px;border:0"></iframe>
Open the page, allow the microphone and talk: the avatar greets you and answers. A sample session ends after a minute. To render on your own machine instead, follow the Quickstart. To try a bitHuman character inside ChatGPT or Claude, add the hosted MCP server (MCP).
What it costs
You need the Creator plan or higher to use the API and the SDKs. A real-time session bills active session time, talking or idle, to the second; rendering on your own hardware or on the device bills a lower rate than the bitHuman cloud. Creating your own avatar is a one-time charge in credits. Every rate is on Pricing.