Pipecat
Give a Pipecat voice bot a lip-synced avatar with pipecat-bithuman.
Put BitHumanVideoService after your TTS service: the bot’s speech goes in, and lip-synced avatar video with the matching audio comes out.
Before you start
pipecat-bithuman is a community integration for Pipecat, maintained by bitHuman, not by the Pipecat team. The avatar renders inside your bot’s own process through the Python SDK, on your own machine. Your pipeline keeps its own speech-to-text, LLM and TTS services.
| Detail | Expression 2 | Essence 2 |
|---|---|---|
| Renders | any character from one portrait | a photoreal person from one portrait |
| Install | pip install "pipecat-bithuman[expression-2]" | included in the same install |
| Avatar file | an Expression 2 .imx | an Essence 2 .imx |
| You need | Check |
|---|---|
| Python 3.11–3.14 | python3 --version |
| macOS 14+ on Apple silicon, or Linux x86_64 or arm64 (a PC with no GPU renders both models) | python3 -c "import platform; print(platform. |
| An API secret | Your API secret |
| A paid plan (Creator or higher): usage bills per second while the avatar runs | Pricing |
ffmpeg and git, for the demo below | ffmpeg -version, git --version |
Install
python3 -m venv .venv
source .venv/bin/activate
pip install "pipecat-bithuman[expression-2]"
The package brings Pipecat (pipecat-ai 1.12.0 or newer) and the bitHuman Python SDK (bithuman) with it.
Authenticate
Set BITHUMAN_API_SECRET in the shell that runs the bot, or pass api_secret= to BitHumanVideoService. The service never logs the secret and removes it from error text. Name the avatar file with model_path=, or set BITHUMAN_MODEL_PATH.
The avatar opens on StartFrame and closes on EndFrame, CancelFrame or cleanup.
Run your first avatar
The package’s demo script sends a WAV file through a Pipecat pipeline as TTSAudioRawFrame chunks, as a TTS service would, and writes the avatar frames that come out to an MP4. In the virtual environment from Install:
export BITHUMAN_API_SECRET="<your API secret>"
git clone https://github.com/bithuman-product/pipecat-bithuman
cd pipecat-bithuman
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
BITHUMAN_MODEL_PATH=wise-pup.imx python examples/render_demo.py speech.wav demo.mp4
# → prints the frame count and size, and demo.mp4 shows wise-pup speaking the sample
To see barge-in, add --interrupt-at 10 --repeat 2: the reply is interrupted after 10 s, and a new one follows. The 37-second demo in the repository is that run with wise-pup.
Performance
The avatar renders through the Python SDK inside your bot’s process. The Python SDK’s speed on each machine:
| Configuration | Essence 2 | Expression 2 |
|---|---|---|
| Intel Core i7-13700F (x86_64) Linux · Python CPU only (no GPU) | 1.9× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26 | 2.3× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26 |
| Apple M4 macOS · Python | 6.9× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25 | 8.4× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25 |
Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.
This section moved: docs.bithuman.ai/platforms/pipecat#run-your-first-avatar