# What is a real-time avatar API?

URL: https://docs.bithuman.ai/start/real-time-avatar-api

> Speech in, lip-synced video out: what it does and which model fits.

A real-time avatar API turns speech audio into lip-synced video of a face while the speech plays, so a voice agent, an app or a kiosk can show a talking character. With bitHuman, your code pushes speech into an SDK and pulls video frames out; the avatar renders on the device, in the browser, on your own computer or in the bitHuman cloud.

Pick Essence 2 for a real person and Expression 2 for any other character.

Every published configuration renders faster than real time, including a desktop CPU with no GPU:

| Configuration | Hardware | Essence 2 | Expression 2 |
|---|---|---|---|
| Cloud API · GPU | NVIDIA RTX 4090 | 4.1× real time | 17.0× real time |
| Linux · CLI · CPU only (no GPU) | Intel Core i7-13700F (x86_64) | 2.0× real time | 2.2× real time |
| iPhone · Swift package | iPhone 15 | 2.1× real time | 5.5× real time |
| Android | Samsung Galaxy S25+ | 2.0× real time | 2.4× real time |

× real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance)).

## What it does

- **One portrait makes an avatar.** You upload a photo or a drawing; the avatar is created in the bitHuman cloud in about 2 to 2.5 hours ([Create your own avatar](https://docs.bithuman.ai/build/create-avatar)).
- **Speech in, frames out.** A session takes audio as it arrives and returns frames with the lips on the words; between replies the avatar idles ([How it works](https://docs.bithuman.ai/models/how-it-works#audio-in-frames-out)).
- **The conversation is yours or ours.** The SDKs render the speech your own voice stack produces. A bitHuman agent runs the whole conversation in the web embed, and with the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab.

## Which model to use

| Your avatar | Model | What it renders | Where it runs |
|---|---|---|---|
| A real person | [Essence 2](https://docs.bithuman.ai/models/essence-2) (`essence-2`) | a photoreal person, up to 1080×1920 | every SDK platform and the bitHuman cloud |
| Any other character: a cartoon, an animal, a mascot, a robot | [Expression 2](https://docs.bithuman.ai/models/expression-2) (`expression-2`) | the whole 416×720 frame, with motion generated from the audio | every SDK platform and the bitHuman cloud |
| An existing first-generation agent | [Essence 1 or Expression 1](https://docs.bithuman.ai/models/first-generation) | a real person's face | Essence 1: the cloud, Python, the CLI and the browser; Expression 1: the cloud only |

Not sure: create with `"model": "auto"`, which sends a real person to Essence 2 and anything else to Expression 2. Each model's places, side by side: [Where each model runs](https://docs.bithuman.ai/models#where-each-model-runs).

## Which SDK to use

| You are building | Use | The avatar renders |
|---|---|---|
| A website | the [web embed](https://docs.bithuman.ai/platforms/web) (one iframe) | in the bitHuman cloud, or in the tab with WebGPU |
| An iPhone, iPad or Mac app | the [Swift package](https://docs.bithuman.ai/platforms/ios) | on the device |
| An Android app | the [Android SDK](https://docs.bithuman.ai/platforms/android), or the [Flutter plugin](https://docs.bithuman.ai/platforms/flutter) | on the device |
| A LiveKit or Pipecat voice agent | the [LiveKit plugin](https://docs.bithuman.ai/platforms/livekit) or [`pipecat-bithuman`](https://docs.bithuman.ai/platforms/pipecat) | in the bitHuman cloud, or on your server |
| A backend or a script | the [Python SDK](https://docs.bithuman.ai/platforms/python) or the [REST API](https://docs.bithuman.ai/platforms/rest) | on your machine, or in the bitHuman cloud |
| A video file, no code | the [CLI](https://docs.bithuman.ai/platforms/cli): `bithuman render` | on your Mac or Linux machine |

Adding a face to an existing voice agent: [Add a talking avatar to a voice agent](https://docs.bithuman.ai/build/how-to/voice-agent-avatar).

## First call

The quickest way to see one is the `wise-pup` sample in the web embed. It needs no account or install:

```html
<iframe src="https://www.bithuman.ai/embed/A23WJF0199" allow="microphone *" style="width:100%;height:600px;border:0"></iframe>
```

Open the page, allow the microphone and talk: the avatar greets you and answers. A sample session ends after a minute. To render on your own machine instead, follow the [Quickstart](https://docs.bithuman.ai/start). To try a bitHuman character inside ChatGPT or Claude, add the hosted MCP server ([MCP](https://docs.bithuman.ai/build/mcp#in-chatgpt-or-claude-with-no-install)).

## What it costs

You need the Creator plan or higher to use the API and the SDKs. A real-time session bills active session time, talking or idle, to the second; rendering on your own hardware or on the device bills a lower rate than the bitHuman cloud. Creating your own avatar is a one-time charge in credits. Every rate is on [Pricing](https://docs.bithuman.ai/pricing).
