Your servers (self-hosted)
More ▾
Run the CLI, the Python SDK or the LiveKit plugin on your own Mac or Linux machines. When the avatar renders on your hardware, its audio and video stay there.
What it is
The avatar renders on machines you run: a Mac with Apple silicon, or a Linux PC or server on x86_64 or arm64, including one with no GPU. You choose where the conversation runs. There is no license to buy for online self-hosting: it needs the Creator plan or higher and bills credits at the self-hosted rate.
| You want | Use | Models |
|---|---|---|
| A talking avatar or an MP4, no code | CLI | Essence 2 and Expression 2 (run, render); Essence 1 (run) |
| Frames or MP4 clips from your own code | Python SDK | Essence 2, Expression 2, Essence 1 |
| A voice agent in your own LiveKit rooms, rendered on your machine | LiveKit plugin with model_path= (guide) | Essence 2, Expression 2 |
Where it renders
| Question | Your servers |
|---|---|
| Where the avatar renders | on your own Mac or Linux machines |
| Where the conversation runs | your choice: the CLI's local conversation brain, your own services, or bitHuman's |
| What reaches bitHuman | a credential check, the avatar download, and usage reports with no audio, video or text |
| Network | to start; rendering continues through a drop of up to 5 minutes |
When the avatar renders on your hardware, its audio and video stay there. For the conversation, the CLI’s local conversation brain keeps speech recognition, the language model and the voice on the machine, or you bring any OpenAI-compatible language model, including one in your own network (Providers). Self-hosted sessions store no transcript at bitHuman.
Models available here
| Model | Your servers | How |
|---|---|---|
| Essence 2 | Yes | CLI, Python, LiveKit plugin |
| Expression 2 | Yes | CLI, Python, LiveKit plugin |
| Essence 1 | Yes | CLI (run), Python |
| Expression 1 | — |
Speed
| Configuration | Essence 2 | Expression 2 |
|---|---|---|
| Intel Core i7-13700F (x86_64) Linux · CLI CPU only (no GPU) | 2.0× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 | 2.2× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 |
| Intel Core i7-13700F (x86_64) Linux · Python CPU only (no GPU) | 1.9× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26 | 2.3× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26 |
| Apple M4 macOS · CLI | 4.2× real timeApple M4 · CLI 2.8.1 · measured 2026-09-27 | 8.4× real timeApple M4 · CLI 2.8.1 · measured 2026-09-27 |
| Apple M4 macOS · Python | 6.9× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25 | 8.4× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25 |
Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.
Price
2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: Pricing and credits.
Rendering an MP4 (bithuman render, or render() in Python) bills the length of the video it writes, at the same rate.
Limits
- Credential: rendering needs a credential. Sign in with
bithuman login, or setBITHUMAN_API_SECRET(Your API secret). - Network: a session checks your credential when it starts and keeps rendering through a network drop of up to 5 minutes. Usage reports carry no audio, video, images or conversation text.
- Sessions: self-hosted sessions are limited by credits.
- Operating systems: macOS on Apple silicon; Linux on x86_64 or arm64. On Windows, use WSL2.
- Off the internet: see Fully offline.
First command
CLI
curl -fsSL https://install.bithuman.ai | sh
bithuman login # in CI, export BITHUMAN_API_SECRET instead
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
bithuman render wise-pup speech.wav -o out.mp4
Python
python3 -m venv .venv && source .venv/bin/activate
pip install "bithuman[expression-2]"
export BITHUMAN_API_SECRET="<your API secret>"
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
LiveKit
pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv
# pass model_path="wise-pup.imx" to bithuman.AvatarSession: /build/voice-agent
out.mp4 is the wise-pup sample avatar speaking the 15-second sample. The CLI’s ffmpeg and live-session setup is on CLI.
Choosing between modes
- A Linux PC with no GPU: CPU only (no GPU).
- Inside an app on the phone, Mac or browser: On the device.
- Nothing to run yourself: bitHuman cloud.
- No internet at the site: Fully offline.
- All five side by side: Deployment options.