Run a talking avatar on your laptop
Render a lip-synced avatar on a Linux, Mac or Windows laptop's processor.
What you’ll build
Install the bitHuman CLI on a Linux laptop (x86_64 or arm64) or an Apple silicon Mac, sign in, and run bithuman render wise-pup speech.wav -o out.mp4 for a video or bithuman run wise-pup to talk to it. Essence 2 and Expression 2 render on the laptop; on Linux they need no GPU. On Windows 11 x86_64, the Python package renders both on the CPU.
The published speed was measured on a desktop CPU and an Apple M4, not on a laptop, so the steps time your own machine:
| Configuration | Essence 2 | Expression 2 |
|---|---|---|
| Intel Core i7-13700F (x86_64) Linux · CLI CPU only (no GPU) | 2.0× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 | 2.2× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 |
| Apple M4 macOS · Python | 6.9× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25 | 8.4× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25 |
| Intel Core i7-13700F (x86_64), 8 threads Windows · Python CPU only (no GPU) | 1.2× real timeIntel Core i7-13700F (x86_64), 8 threads, CPU only (no GPU) · bithuman 2.11.18 · measured 2026-09-29 | 1.0× real timeIntel Core i7-13700F (x86_64), 8 threads, CPU only (no GPU) · bithuman 2.11.18 · measured 2026-09-29 |
Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.
You need an API secret or a bitHuman sign-in, on the Creator plan or higher.
Steps
4 steps
Check your laptop
Laptop Use Check Linux on x86_64 or arm64 the CLI or the Python SDK uname -smmacOS 14 or newer on Apple silicon the CLI or the Python SDK uname -smprintsDarwin arm64Windows 11 on x86_64 the Python SDK on Windows; the CLI on Windows renders in the cloud python -c "import platform; print(platform.printssystem(), platform. machine())" Windows AMD64Intel Macs and Windows on Arm have no package.
Expected
One row that matches your laptop.
Install and sign in
On Linux or macOS, install the CLI and sign in:
# Linux, x86_64 or arm64 (Debian/Ubuntu) sudo apt install -y ffmpeg python3-venv curl -fsSL https://install.bithuman.ai | sh # macOS (Apple silicon): also installs ffmpeg, livekit-server and Python brew tap bithuman/bithuman https://gitlab.com/bithuman/sdk/homebrew-bithuman && brew install bithuman/bithuman/bithuman-cli bithuman login # opens a browser and stores a credential for this device bithuman account # exit 0 when signed inOn Windows 11, in PowerShell:
python -m venv .venv .venv\Scripts\Activate.ps1 pip install "bithuman[expression-2]" $env:BITHUMAN_API_SECRET = "<your API secret>"Expected
bithuman accountexits 0 on Linux or macOS; on Windows,pipfinishes with no error.Render the sample and time it
On Linux or macOS:
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav time bithuman render wise-pup speech.wav -o out.mp4 # → out.mp4: 416×720, 300 frames, 15.0 sOn Windows 11:
curl.exe -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2" curl.exe -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav Measure-Command { python -c "import bithuman; bithuman.open('wise-pup.imx').render('speech.wav', out_mp4='out.mp4')" }The sample is 15 seconds of speech. Run the render twice, because the first run also downloads the avatar. A second run that takes less than 15 seconds means your laptop renders Expression 2 faster than real time. For the photoreal model, render
sofia-ramirezthe same way.Expected
out.mp4:wise-pupspeaking the sample, 416×720, with its lips in sync.Talk to it live
On Linux or macOS,
bithuman runstarts a local LiveKit server, a voice agent and the avatar, then opens the page in your browser. It needs Python 3.11 or newer for its voice agent:bithuman run wise-pup # → open the printed http://127.0.0.1:8088/<CODE> and allow the microphoneThe voice runs on OpenAI Realtime, with your
OPENAI_API_KEYor on your bitHuman account (voice settings). For a conversation that also runs on the laptop, use the local conversation brain:BITHUMAN_LOCAL=1 bithuman run sofia-ramirez. On Windows, stream audio intoAsyncBithumanfrom your own code (Python: Integrate into your app).Expected
Say "hi": the avatar answers out loud with its lips in sync, and stops when you talk over it.
All steps done. Next: make it your own.
How it works
The avatar renders on the laptop, from the audio to the frames, so its audio and video stay there. A session checks your credential when it starts, keeps rendering through a network drop of up to 5 minutes, and sends usage reports with no audio, video, images or conversation text. Rendering on your own hardware bills at the self-hosted rate (pricing).
Make it your own
- Your own avatar: create one from a portrait, then
bithuman pull <AGENT_CODE>and render it the same way. - Your own words: any audio file
ffmpegreads, from a recording or a text-to-speech tool: Talking video. - A screen that runs all day: a Linux PC with no GPU: CPU only (no GPU).
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
bithuman render exits with code 77 | no credential | run bithuman login, or export BITHUMAN_ |
pip finds no wheel on Windows | 32-bit Python, Windows on Arm, or Python outside 3.10–3.14 | install 64-bit Python 3.10–3.14 on an x86_64 PC |
'curl' is not recognized in PowerShell | curl in Windows PowerShell 5.1 is an alias | type curl.exe, as above |
| The second render takes longer than the speech | the laptop’s processor is slower than the measured hardware | render files with bithuman render, or use the bitHuman cloud for live sessions |
More on CLI: Troubleshooting.