Docs index: /llms.txt · every page as Markdown: add .md

‹ Overview

What is a real-time avatar API?

Speech in, lip-synced video out: what it does and which model fits.

Creator plan or higher

A real-time avatar API turns speech audio into lip-synced video of a face while the speech plays, so a voice agent, an app or a kiosk can show a talking character. With bitHuman, your code pushes speech into an SDK and pulls video frames out; the avatar renders on the device, in the browser, on your own computer or in the bitHuman cloud.

Pick Essence 2 for a real person and Expression 2 for any other character.

Every published configuration renders faster than real time, including a desktop CPU with no GPU:

ConfigurationEssence 2Expression 2
NVIDIA RTX 4090 Cloud API · GPU
4.1× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-27
17.0× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-23
Intel Core i7-13700F (x86_64) Linux · CLI CPU only (no GPU)
2.0× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27
2.2× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27
iPhone 15 iPhone · Swift package
2.1× real timeiPhone 15 · Swift package 2.17.3 · measured 2026-09-27
5.5× real timeiPhone 15 · Swift package 2.18.0 · measured 2026-09-27
Samsung Galaxy S25+ Android
2.0× real timeSamsung Galaxy S25+ · essence2-android 0.7.0 · measured 2026-09-25
2.4× real timeSamsung Galaxy S25+ · expression2-android 0.4.10 · measured 2026-09-23

Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.

What it does

  • One portrait makes an avatar. You upload a photo or a drawing; the avatar is created in the bitHuman cloud in about 2 to 2.5 hours (Create your own avatar).
  • Speech in, frames out. A session takes audio as it arrives and returns frames with the lips on the words; between replies the avatar idles (How it works).
  • The conversation is yours or ours. The SDKs render the speech your own voice stack produces. A bitHuman agent runs the whole conversation in the web embed, and with the web embed, the conversation runs on bitHuman’s servers, even when the avatar renders in the tab.

Which model to use

Your avatarModelWhat it rendersWhere it runs
A real personEssence 2 (essence-2)a photoreal person, up to 1080×1920every SDK platform and the bitHuman cloud
Any other character: a cartoon, an animal, a mascot, a robotExpression 2 (expression-2)the whole 416×720 frame, with motion generated from the audioevery SDK platform and the bitHuman cloud
An existing first-generation agentEssence 1 or Expression 1a real person’s faceEssence 1: the cloud, Python, the CLI and the browser; Expression 1: the cloud only

Not sure: create with "model": "auto", which sends a real person to Essence 2 and anything else to Expression 2. Each model’s places, side by side: Where each model runs.

Which SDK to use

You are buildingUseThe avatar renders
A websitethe web embed (one iframe)in the bitHuman cloud, or in the tab with WebGPU
An iPhone, iPad or Mac appthe Swift packageon the device
An Android appthe Android SDK, or the Flutter pluginon the device
A LiveKit or Pipecat voice agentthe LiveKit plugin or pipecat-bithumanin the bitHuman cloud, or on your server
A backend or a scriptthe Python SDK or the REST APIon your machine, or in the bitHuman cloud
A video file, no codethe CLI: bithuman renderon your Mac or Linux machine

Adding a face to an existing voice agent: Add a talking avatar to a voice agent.

First call

The quickest way to see one is the wise-pup sample in the web embed. It needs no account or install:

<iframe src="https://www.bithuman.ai/embed/A23WJF0199" allow="microphone *" style="width:100%;height:600px;border:0"></iframe>

Open the page, allow the microphone and talk: the avatar greets you and answers. A sample session ends after a minute. To render on your own machine instead, follow the Quickstart. To try a bitHuman character inside ChatGPT or Claude, add the hosted MCP server (MCP).

What it costs

You need the Creator plan or higher to use the API and the SDKs. A real-time session bills active session time, talking or idle, to the second; rendering on your own hardware or on the device bills a lower rate than the bitHuman cloud. Creating your own avatar is a one-time charge in credits. Every rate is on Pricing.