Deployment options
More ▾
The avatar renders in one of five places. Pick the one that matches where your users are and what may leave their device.
Every mode uses the same avatar and the same API secret. What changes is where the avatar renders, where the conversation runs, and what reaches bitHuman.
Compare the modes
- bitHuman cloudCreator plan or higher
- The avatar renders
- on bitHuman's servers, in the US
- The conversation runs
- on bitHuman's voice service, or with your own provider keys
- What reaches bitHuman
- the session's audio and conversation, to run it
- Network
- required for the whole session
- Price
- 4 credits per minute (Essence 2, Expression 2)
- Products
- web embed, REST API, LiveKit plugin, cloud avatar in your room
- Models
- Essence 2, Expression 2, Essence 1, Expression 1
- Your serversCreator plan or higher
- The avatar renders
- on your own Mac or Linux machines
- The conversation runs
- your choice: the CLI's local conversation brain, your own services, or bitHuman's
- What reaches bitHuman
- a credential check, the avatar download, and usage reports with no audio, video or text
- Network
- to start; rendering continues through a drop of up to 5 minutes
- Price
- 2 credits per minute (Essence 2, Expression 2)
- Products
- CLI, Python SDK, LiveKit plugin (
model_path) - Models
- Essence 2, Expression 2, Essence 1
- On the deviceCreator plan or higher
- The avatar renders
- on the iPhone, iPad, Mac or Android phone in front of the user, or in a WebGPU browser tab
- The conversation runs
- your app's choice; with the web embed, on bitHuman's servers
- What reaches bitHuman
- with your own voice and language services, usage metering only, never audio, video or conversation text
- Network
- to start; rendering continues through a drop of up to 5 minutes
- Price
- 2 credits per minute (Essence 2, Expression 2)
- Products
- Swift package, Android SDK, Flutter plugin, web embed with
render=local - Models
- Essence 2, Expression 2
- CPU only (no GPU)Creator plan or higher
- The avatar renders
- on a standard Linux PC's CPU, with no GPU
- The conversation runs
- your choice: the CLI's local conversation brain, your own services, or bitHuman's
- What reaches bitHuman
- a credential check, the avatar download, and usage reports with no audio, video or text
- Network
- to start; rendering continues through a drop of up to 5 minutes
- Price
- 2 credits per minute (Essence 2, Expression 2)
- Products
- CLI, Python SDK on Linux x86_64 or arm64
- Models
- Essence 2, Expression 2, Essence 1
- Fully offlineBusiness & Enterprise
- The avatar renders
- on your Linux PCs and terminals
- The conversation runs
- agreed with sales for your site
- What reaches bitHuman
- usage is metered on the machine; no reconnection is required
- Network
- off the internet; creating the avatar happens online first
- Price
- from 100,000 credits; arranged through sales
- Products
- Linux PCs and terminals
- Models
- Essence 1 (Linux x86_64 and ARM64); Essence 2 and Expression 2 later
Every mode bills active session time, talking or idle, to the second (pricing).
bitHuman cloud
bitHuman renders the avatar in the US and streams it to your page, app or LiveKit room. bitHuman cloud
Your servers (self-hosted)
The CLI, Python or the LiveKit plugin on your own Mac or Linux machines. When the avatar renders on your hardware, its audio and video stay there. Your servers
On the device
Essence 2 and Expression 2 render on iPhone, iPad, Mac and Android, or in a WebGPU browser tab. On the device
CPU only (no GPU)
Both models run live on a standard Linux PC with no GPU. CPU only (no GPU)
Fully offline
Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. Fully offline
Choosing a mode
- Building for the web, or want the fewest moving parts? The bitHuman cloud, starting with the web embed.
- Must audio and video stay on your network? Your servers or on the device, with the local conversation brain or your own models. What reaches bitHuman in each: Data flows & privacy.
- Want the lowest per-minute rate? Render on the device or on your servers (rates).
- No reliable internet at the site? Fully offline, on the Business and Enterprise plans.
- Planning for a branch, a clinic or a show floor? Start from Use cases.
Try it
Talk to a sample avatar streamed from the bitHuman cloud: the same web embed a customer puts on a site.
The avatar renders in the bitHuman cloud, and the conversation runs on bitHuman's servers.