bitHuman cloud

bitHuman renders the avatar on its servers and streams it to your page, app or LiveKit room: the web embed, the REST API or the LiveKit plugin.

Creator plan or higher bitHuman cloud
sofia-ramirez, the Essence 2 sample avatar

Live in your browser, no account. Up to 3 minutes; allow the microphone when asked.

The avatar renders in the bitHuman cloud, and the conversation runs on bitHuman's servers.

What it is

The fewest moving parts: bitHuman renders the avatar in the cloud and streams its video and audio to a web page, an app or a LiveKit room. You provision no GPU and install nothing. The live demo above is this mode.

Where it renders

QuestionbitHuman cloud
Where the avatar renderson bitHuman's servers, in the US
Where the conversation runson bitHuman's voice service, or with your own provider keys
What reaches bitHumanthe session's audio and conversation, to run it
Networkrequired for the whole session
YOUR APP OR SITEA browser or appsends the microphone, shows the avatarBITHUMAN CLOUD · USThe avatar rendersThe conversation runsbitHuman's voice service, or the provider keysyou connectmicrophoneaudiovideo and voiceYour API secret stays on your server; browsers getscoped embed tokens.
bitHuman cloud. In the bitHuman cloud the avatar renders on bitHuman's servers, in the US, and the conversation runs on bitHuman's voice service or with the provider keys you connect. The browser or app sends the microphone and shows the video. Your API secret stays on your server; browsers get scoped embed tokens.

A managed agent’s conversation runs on bitHuman’s voice service with your persona, or with the voice and language providers whose keys you connect (Voices, Providers). Traffic is encrypted in transit: HTTPS, and WebRTC media over DTLS-SRTP.

Models available here

ModelbitHuman cloudHow
Essence 2Yesweb embed, REST API, LiveKit
Expression 2Yesweb embed, REST API, LiveKit
Essence 1Yesweb embed, REST API, LiveKit
Expression 1Yesweb embed, REST API, LiveKit

Speed

The bitHuman cloud serves from several kinds of hardware; each is measured:

ConfigurationEssence 2Expression 2
NVIDIA RTX 4090 Cloud API · GPU
4.1× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-27
17.0× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-23
Apple M4 Max Cloud API · Apple silicon
2.8× real timeApple M4 Max · cloud API · measured 2026-09-26
5.5× real timeApple M4 Max · cloud API · measured 2026-09-23
x86 server CPU Cloud API · CPU CPU only (no GPU)
1.1× real timex86 server CPU, CPU only (no GPU) · cloud API · measured 2026-09-25
1.3× real timex86 server CPU, CPU only (no GPU) · cloud API · measured 2026-09-24

Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.

Price

4 credits per minute of active session time for Essence 2 and Expression 2, about $0.04 a minute at the top-up rate of $1 = 100 credits. A managed agent's voice chat bills 10 credits per minute, all-inclusive: the avatar is part of it. Realtime usage bills active session time, talking or idle, to the second. Every rate: Pricing and credits.

Limits

bitHuman cloud sessions are limited per plan: Creator 3, Pro 10, Business 50, Enterprise 200 concurrent sessions. On-device and self-hosted sessions are limited by credits (plans).

A session over your plan’s limit is refused with 403 CONCURRENCY_LIMIT_REACHED (Rate limits). The network is needed for the whole session.

First command

Web embed

<iframe src="https://www.bithuman.ai/embed/A23WJF0199" allow="microphone *"
        style="width:100%;height:600px;border:0"></iframe>

REST API

curl -s -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET"
# → {"valid":true}

LiveKit

pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv
# then pass avatar_id= to bithuman.AvatarSession: /platforms/livekit

Next steps: Web, REST API, LiveKit, or a cloud avatar in your own room without the plugin.

Choosing between modes