bitHuman cloud
More ▾
bitHuman renders the avatar on its servers and streams it to your page, app or LiveKit room: the web embed, the REST API or the LiveKit plugin.
The avatar renders in the bitHuman cloud, and the conversation runs on bitHuman's servers.
What it is
The fewest moving parts: bitHuman renders the avatar in the cloud and streams its video and audio to a web page, an app or a LiveKit room. You provision no GPU and install nothing. The live demo above is this mode.
Where it renders
| Question | bitHuman cloud |
|---|---|
| Where the avatar renders | on bitHuman's servers, in the US |
| Where the conversation runs | on bitHuman's voice service, or with your own provider keys |
| What reaches bitHuman | the session's audio and conversation, to run it |
| Network | required for the whole session |
A managed agent’s conversation runs on bitHuman’s voice service with your persona, or with the voice and language providers whose keys you connect (Voices, Providers). Traffic is encrypted in transit: HTTPS, and WebRTC media over DTLS-SRTP.
Models available here
| Model | bitHuman cloud | How |
|---|---|---|
| Essence 2 | Yes | web embed, REST API, LiveKit |
| Expression 2 | Yes | web embed, REST API, LiveKit |
| Essence 1 | Yes | web embed, REST API, LiveKit |
| Expression 1 | Yes | web embed, REST API, LiveKit |
Speed
The bitHuman cloud serves from several kinds of hardware; each is measured:
| Configuration | Essence 2 | Expression 2 |
|---|---|---|
| NVIDIA RTX 4090 Cloud API · GPU | 4.1× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-27 | 17.0× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-23 |
| Apple M4 Max Cloud API · Apple silicon | 2.8× real timeApple M4 Max · cloud API · measured 2026-09-26 | 5.5× real timeApple M4 Max · cloud API · measured 2026-09-23 |
| x86 server CPU Cloud API · CPU CPU only (no GPU) | 1.1× real timex86 server CPU, CPU only (no GPU) · cloud API · measured 2026-09-25 | 1.3× real timex86 server CPU, CPU only (no GPU) · cloud API · measured 2026-09-24 |
Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.
Price
4 credits per minute of active session time for Essence 2 and Expression 2, about $0.04 a minute at the top-up rate of $1 = 100 credits. A managed agent's voice chat bills 10 credits per minute, all-inclusive: the avatar is part of it. Realtime usage bills active session time, talking or idle, to the second. Every rate: Pricing and credits.
Limits
bitHuman cloud sessions are limited per plan: Creator 3, Pro 10, Business 50, Enterprise 200 concurrent sessions. On-device and self-hosted sessions are limited by credits (plans).
A session over your plan’s limit is refused with 403 CONCURRENCY_LIMIT_REACHED (Rate limits). The network is needed for the whole session.
First command
Web embed
<iframe src="https://www.bithuman.ai/embed/A23WJF0199" allow="microphone *"
style="width:100%;height:600px;border:0"></iframe>
REST API
curl -s -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET"
# → {"valid":true}
LiveKit
pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv
# then pass avatar_id= to bithuman.AvatarSession: /platforms/livekit
Next steps: Web, REST API, LiveKit, or a cloud avatar in your own room without the plugin.
Choosing between modes
- Keep audio and video on your own machines: Your servers.
- Render inside your app on the phone, Mac or browser: On the device.
- A Linux PC with no GPU: CPU only (no GPU).
- No internet at the site: Fully offline.
- All five side by side: Deployment options.