Local conversation brain
More ▾
The CLI's local conversation brain: speech recognition, the language model and the voice run on your Mac or Linux machine with one environment variable (BITHUMAN_LOCAL=1). Audio and transcripts stay on the machine.
BITHUMAN_LOCAL=1 bithuman run replaces the cloud conversation brain with one that runs on your machine: whisper.cpp for speech recognition, llama.cpp for the language model, Supertonic for speech, and Silero for voice detection. The command, the browser URL and the avatar stay the same.
Audio, transcripts and generated speech never leave the machine. The avatar session is still reported to your account, so run needs a sign-in and a network connection. Running realtime avatars off the internet is a separate arrangement: the offline license (Business and Enterprise).
Before you start
- The CLI, signed in (
bithuman login, orBITHUMAN_API_SECRET). - About 1 GB of disk and 1.5 GB of free memory.
cmakeand a C++ compiler:llama-cpp-pythonbuilds from source (several minutes).
1. Download an avatar and start the brain once
bithuman pull sofia-ramirez
bithuman run sofia-ramirez # press Ctrl-C once the URL prints
The first run creates the brain’s Python environment at ~/.cache/bithuman/brain-venv.
2. Install the local conversation brain into that environment
~/.cache/bithuman/brain-venv/bin/python -m pip install \
'livekit-agents[silero]~=1.5' supertonic pywhispercpp llama-cpp-python soxr
Install into that interpreter, not your system Python: the brain runs from it, and Debian and Ubuntu refuse a system-wide pip install.
3. Run it
BITHUMAN_LOCAL=1 bithuman run sofia-ramirez
# → open the printed http://127.0.0.1:8088/<CODE> and talk
Check it worked
bithuman doctor lists the brain packages and their versions once they import. The first local run downloads the brain models (about 860 MB) into ~/.cache/huggingface and ~/.cache/supertonic, once; later runs start in about a second.
Tuning
| Variable | Default | Effect |
|---|---|---|
BITHUMAN_LOCAL | unset | 1 uses the local conversation brain |
BITHUMAN_LOCAL_WHISPER | tiny.en | Speech-recognition model: tiny.en, base.en, or multilingual tiny, base, small, medium, large-v3-turbo |
BITHUMAN_LOCAL_LLM | Qwen/Qwen2.5-0.5B-Instruct-GGUF | Any Hugging Face GGUF chat model |
BITHUMAN_LOCAL_LLM_FILE | qwen2.5-0.5b-instruct-q4_k_m.gguf | The GGUF file in that repository |
BITHUMAN_LOCAL_VOICE | M1 | Voice preset: M1–M5, F1–F5 |
BITHUMAN_LOCAL_LANG | en | Speech language; 31 are supported (en, ko, ja, es, de, fr, zh, hi, ar and others) |
BITHUMAN_INSTRUCTIONS | a short default | The system prompt |
A larger language model answers better and uses more memory. For example, Qwen 2.5 1.5B (about 1.5 GB of memory):
export BITHUMAN_LOCAL_LLM="Qwen/Qwen2.5-1.5B-Instruct-GGUF"
export BITHUMAN_LOCAL_LLM_FILE="qwen2.5-1.5b-instruct-q4_k_m.gguf"
BITHUMAN_LOCAL=1 bithuman run sofia-ramirez
For another language, pair a multilingual speech model with a voice in that language:
export BITHUMAN_LOCAL_WHISPER=small BITHUMAN_LOCAL_LANG=ko BITHUMAN_LOCAL_VOICE=F1
BITHUMAN_LOCAL=1 bithuman run sofia-ramirez
Cloud brain or local conversation brain
| Detail | Cloud brain (default) | Local conversation brain |
|---|---|---|
| Setup | sign in (or set OPENAI_API_KEY) | three steps above |
| Cost | the managed voice-chat rate, 10 credits/min (pricing); with OPENAI_API_KEY, your OpenAI bill instead | no brain charge |
| Where audio goes | to the speech and language service | stays on the machine |
| Memory | about 300 MB (the avatar) | about 1.5 GB |
| Languages | the service’s | 31 for speech output |
| Tool calling | yes | needs a 3B or larger model |
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
BITHUMAN_LOCAL=1 fails asking for local extras | the packages are not in the brain’s environment | repeat step 2 with that interpreter |
error: externally-managed-environment | pip targeted the system Python | use ~/.cache/bithuman/brain-venv/bin/python -m pip |
bithuman doctor says the brain venv is not bootstrapped | step 1 has not run | run bithuman run once |
| The first run seems stuck | it is downloading about 860 MB of brain models | wait; later runs are fast |
| Replies are low quality | the default model is small | set a larger BITHUMAN_LOCAL_LLM |
| Speech is in the wrong language | the voice and language settings differ | set BITHUMAN_LOCAL_LANG and a matching voice |