Local conversation brain

The CLI's local conversation brain: speech recognition, the language model and the voice run on your Mac or Linux machine with one environment variable (BITHUMAN_LOCAL=1). Audio and transcripts stay on the machine.

Your servers No GPU CLI 2.8.3

BITHUMAN_LOCAL=1 bithuman run replaces the cloud conversation brain with one that runs on your machine: whisper.cpp for speech recognition, llama.cpp for the language model, Supertonic for speech, and Silero for voice detection. The command, the browser URL and the avatar stay the same.

Audio, transcripts and generated speech never leave the machine. The avatar session is still reported to your account, so run needs a sign-in and a network connection. Running realtime avatars off the internet is a separate arrangement: the offline license (Business and Enterprise).

Before you start

  • The CLI, signed in (bithuman login, or BITHUMAN_API_SECRET).
  • About 1 GB of disk and 1.5 GB of free memory.
  • cmake and a C++ compiler: llama-cpp-python builds from source (several minutes).

1. Download an avatar and start the brain once

bithuman pull sofia-ramirez
bithuman run sofia-ramirez      # press Ctrl-C once the URL prints

The first run creates the brain’s Python environment at ~/.cache/bithuman/brain-venv.

2. Install the local conversation brain into that environment

~/.cache/bithuman/brain-venv/bin/python -m pip install \
  'livekit-agents[silero]~=1.5' supertonic pywhispercpp llama-cpp-python soxr

Install into that interpreter, not your system Python: the brain runs from it, and Debian and Ubuntu refuse a system-wide pip install.

3. Run it

BITHUMAN_LOCAL=1 bithuman run sofia-ramirez
# → open the printed http://127.0.0.1:8088/<CODE> and talk

Check it worked

bithuman doctor lists the brain packages and their versions once they import. The first local run downloads the brain models (about 860 MB) into ~/.cache/huggingface and ~/.cache/supertonic, once; later runs start in about a second.

Tuning

VariableDefaultEffect
BITHUMAN_LOCALunset1 uses the local conversation brain
BITHUMAN_LOCAL_WHISPERtiny.enSpeech-recognition model: tiny.en, base.en, or multilingual tiny, base, small, medium, large-v3-turbo
BITHUMAN_LOCAL_LLMQwen/Qwen2.5-0.5B-Instruct-GGUFAny Hugging Face GGUF chat model
BITHUMAN_LOCAL_LLM_FILEqwen2.5-0.5b-instruct-q4_k_m.ggufThe GGUF file in that repository
BITHUMAN_LOCAL_VOICEM1Voice preset: M1–M5, F1–F5
BITHUMAN_LOCAL_LANGenSpeech language; 31 are supported (en, ko, ja, es, de, fr, zh, hi, ar and others)
BITHUMAN_INSTRUCTIONSa short defaultThe system prompt

A larger language model answers better and uses more memory. For example, Qwen 2.5 1.5B (about 1.5 GB of memory):

export BITHUMAN_LOCAL_LLM="Qwen/Qwen2.5-1.5B-Instruct-GGUF"
export BITHUMAN_LOCAL_LLM_FILE="qwen2.5-1.5b-instruct-q4_k_m.gguf"
BITHUMAN_LOCAL=1 bithuman run sofia-ramirez

For another language, pair a multilingual speech model with a voice in that language:

export BITHUMAN_LOCAL_WHISPER=small BITHUMAN_LOCAL_LANG=ko BITHUMAN_LOCAL_VOICE=F1
BITHUMAN_LOCAL=1 bithuman run sofia-ramirez

Cloud brain or local conversation brain

DetailCloud brain (default)Local conversation brain
Setupsign in (or set OPENAI_API_KEY)three steps above
Costthe managed voice-chat rate, 10 credits/min (pricing); with OPENAI_API_KEY, your OpenAI bill insteadno brain charge
Where audio goesto the speech and language servicestays on the machine
Memoryabout 300 MB (the avatar)about 1.5 GB
Languagesthe service’s31 for speech output
Tool callingyesneeds a 3B or larger model

Troubleshooting

SymptomCauseFix
BITHUMAN_LOCAL=1 fails asking for local extrasthe packages are not in the brain’s environmentrepeat step 2 with that interpreter
error: externally-managed-environmentpip targeted the system Pythonuse ~/.cache/bithuman/brain-venv/bin/python -m pip
bithuman doctor says the brain venv is not bootstrappedstep 1 has not runrun bithuman run once
The first run seems stuckit is downloading about 860 MB of brain modelswait; later runs are fast
Replies are low qualitythe default model is smallset a larger BITHUMAN_LOCAL_LLM
Speech is in the wrong languagethe voice and language settings differset BITHUMAN_LOCAL_LANG and a matching voice

Next