# Pipecat

URL: https://docs.bithuman.ai/platforms/pipecat

> Give a Pipecat voice bot a lip-synced avatar with pipecat-bithuman.

Put `BitHumanVideoService` after your TTS service: the bot's speech goes in, and lip-synced avatar video with the matching audio comes out.

## Before you start

`pipecat-bithuman` is a community integration for Pipecat, maintained by bitHuman, not by the Pipecat team. The avatar renders inside your bot's own process through the [Python SDK](https://docs.bithuman.ai/platforms/python), on your own machine. Your pipeline keeps its own speech-to-text, LLM and TTS services.

| Detail | Expression 2 | Essence 2 |
|---|---|---|
| **Renders** | [any character from one portrait](https://docs.bithuman.ai/models/expression-2) | [a photoreal person from one portrait](https://docs.bithuman.ai/models/essence-2) |
| **Install** | `pip install "pipecat-bithuman[expression-2]"` | included in the same install |
| **Avatar file** | an Expression 2 `.imx` | an Essence 2 `.imx` |

| You need | Check |
|---|---|
| Python 3.11–3.14 | `python3 --version` |
| macOS 14+ on Apple silicon, or Linux x86_64 or arm64 (a PC with no GPU renders both models) | `python3 -c "import platform; print(platform.system(), platform.machine())"` |
| An API secret | [Your API secret](https://docs.bithuman.ai/start/api-secret) |
| A paid plan (Creator or higher): usage bills per second while the avatar runs | [Pricing](https://docs.bithuman.ai/pricing) |
| `ffmpeg` and `git`, for the demo below | `ffmpeg -version`, `git --version` |

## Install

```bash
python3 -m venv .venv
source .venv/bin/activate
pip install "pipecat-bithuman[expression-2]"
```

The package brings Pipecat (`pipecat-ai` 1.12.0 or newer) and the bitHuman Python SDK (`bithuman`) with it.

## Authenticate

Set `BITHUMAN_API_SECRET` in the shell that runs the bot, or pass `api_secret=` to `BitHumanVideoService`. The service never logs the secret and removes it from error text. Name the avatar file with `model_path=`, or set `BITHUMAN_MODEL_PATH`.

The avatar opens on `StartFrame` and closes on `EndFrame`, `CancelFrame` or cleanup.

## Run your first avatar

The package's demo script sends a WAV file through a Pipecat pipeline as `TTSAudioRawFrame` chunks, as a TTS service would, and writes the avatar frames that come out to an MP4. In the virtual environment from Install:

```bash
export BITHUMAN_API_SECRET="<your API secret>"
git clone https://github.com/bithuman-product/pipecat-bithuman
cd pipecat-bithuman
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav
BITHUMAN_MODEL_PATH=wise-pup.imx python examples/render_demo.py speech.wav demo.mp4
# → prints the frame count and size, and demo.mp4 shows wise-pup speaking the sample
```

To see barge-in, add `--interrupt-at 10 --repeat 2`: the reply is interrupted after 10 s, and a new one follows. The [37-second demo](https://github.com/bithuman-product/pipecat-bithuman/blob/main/docs/demo.mp4) in the repository is that run with `wise-pup`.

## Performance

The avatar renders through the Python SDK inside your bot's process. The Python SDK's speed on each machine:

| Configuration | Hardware | Essence 2 | Expression 2 |
|---|---|---|---|
| Linux · Python · CPU only (no GPU) | Intel Core i7-13700F (x86_64) | 1.9× real time | 2.3× real time |
| macOS · Python | Apple M4 | 6.9× real time | 8.4× real time |

× real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance#desktop)).

## Continue

The rest of this platform, inlined so one fetch covers it: [Build a Pipecat bot](https://docs.bithuman.ai/platforms/pipecat/app.md) · [Pipecat troubleshooting](https://docs.bithuman.ai/platforms/pipecat/troubleshooting.md).

---

# Build a Pipecat bot

URL: https://docs.bithuman.ai/platforms/pipecat/app

> Fit BitHumanVideoService into your Pipecat pipeline, transport and error handling.

## Integrate into your app

- **Where it goes.** After the TTS service, or a speech-to-speech LLM, and before `transport.output()`.
- **The transport.** Turn on video out with `video_out_enabled=True`. The service logs the avatar's frame size on the first frame; set `video_out_width` and `video_out_height` to it to avoid resizing.
- **Voice and picture together.** The service does not forward your TTS audio. It pushes the avatar's copy of the speech, paired with each picture. Each image carries `sync_with_audio`, so keep the transport's `video_out_is_live` off (its default).
- **Barge-in.** On `InterruptionFrame` the avatar drops the reply in flight and goes back to idle. See [Barge-in](https://docs.bithuman.ai/build/barge-in).
- **The end of a reply.** `TTSStoppedFrame` is held until the avatar has finished speaking the reply.
- **Session time.** The avatar opens on `StartFrame` and closes on `EndFrame`, `CancelFrame` or cleanup; it is billed while open, talking or idle. End the pipeline when the user leaves, as the example below does.
- **When the avatar fails.** The service pushes one `ErrorFrame` upstream and, by default, passes TTS audio through unchanged, so the bot keeps talking without video. `audio_passthrough_on_error=False` drops the audio instead.
- **Metrics.** With `enable_metrics=True` in `PipelineParams`, the service reports TTFB: from a reply's first audio in to its first voiced avatar frame out.

### What the service does with frames

| In | Out |
|---|---|
| `TTSAudioRawFrame` | sent to the avatar, not forwarded as is |
| (avatar frame) | `OutputImageRawFrame` (RGB), while talking and while idle |
| (avatar audio) | `TTSAudioRawFrame` (16 kHz mono), paired with each picture |
| `TTSStoppedFrame` | held until the avatar has finished speaking the reply |
| `InterruptionFrame` | the avatar drops the reply in flight and goes back to idle |
| anything else | passed on unchanged |

## Complete example

`examples/bot.py` in [pipecat-bithuman](https://github.com/bithuman-product/pipecat-bithuman) is a complete voice bot in a Daily room, with Deepgram speech-to-text, an OpenAI LLM and Cartesia TTS. The avatar sits between the TTS and the output transport:

```python
# excerpt: the transport and the pipeline, from pipecat-bithuman's examples/bot.py
transport = DailyTransport(
    os.environ["DAILY_ROOM_URL"],
    None,
    "Pip",
    DailyParams(
        audio_in_enabled=True,
        audio_out_enabled=True,
        video_out_enabled=True,
        video_out_width=416,   # the wise-pup sample's frame size; the service logs yours
        video_out_height=720,
    ),
)
# …
avatar = BitHumanVideoService()  # BITHUMAN_MODEL_PATH + BITHUMAN_API_SECRET
# …
pipeline = Pipeline(
    [
        transport.input(),
        stt,
        aggregators.user(),
        llm,
        tts,
        avatar,
        transport.output(),
        aggregators.assistant(),
    ]
)
# …
@transport.event_handler("on_participant_left")
async def on_left(transport, participant, reason):
    await worker.cancel()  # close the avatar: session time stops
```

Run it from the repository, with the avatar file from [Pipecat](https://docs.bithuman.ai/platforms/pipecat#run-your-first-avatar). `examples/bot.py` sets the transport to the `wise-pup` frame size, 416×720; for another avatar, use the size the service logs on the first frame:

```bash
pip install "pipecat-bithuman[expression-2]" "pipecat-ai[daily,deepgram,openai,cartesia,silero]"
export BITHUMAN_API_SECRET="<your API secret>" BITHUMAN_MODEL_PATH=wise-pup.imx
export DAILY_ROOM_URL=… DEEPGRAM_API_KEY=… OPENAI_API_KEY=… CARTESIA_API_KEY=… CARTESIA_VOICE_ID=…
python examples/bot.py
# → join the Daily room; the avatar says hello first, then answers
```

## Platform notes

- Requires Python 3.11–3.14 and `pipecat-ai` 1.12.0 or newer.
- Your own avatar: [download an agent's model](https://docs.bithuman.ai/api/agents#download-an-agents-model) and pass the `.imx` file as `model_path=`.
- The service has no settings to change at runtime. The avatar is chosen when the service is built.
- For tests, or to wrap the SDK, pass `runtime_factory=`: an async callable that returns an object implementing `BitHumanRuntime`. The package's own tests run on fakes, with no network, secret or model file.
- Report bugs in the [pipecat-bithuman issues](https://github.com/bithuman-product/pipecat-bithuman/issues). The Pipecat team does not maintain this package.

## Reference

### `BitHumanVideoService` options

| Argument | Default | Meaning |
|---|---|---|
| `model_path` | `BITHUMAN_MODEL_PATH` | the `.imx` avatar file |
| `api_secret` | `BITHUMAN_API_SECRET` | the API secret, read by the SDK when not passed |
| `sync_video_to_audio` | `True` | sets `sync_with_audio` on each image |
| `audio_passthrough_on_error` | `True` | keep the voice if the avatar fails |
| `stop_frame_timeout_s` | `2.0` | release a held `TTSStoppedFrame` after this much quiet |
| `end_drain_timeout_s` | `30.0` | the longest wait on `EndFrame` for queued speech |
| `runtime_factory` | the SDK | advanced: your own `BitHumanRuntime` |

Properties: `is_avatar_ready` (the avatar is open and rendering) and `frame_size` (`(width, height)` of the first frame, or `None` before it).

### How it maps to the Python SDK

| Pipecat | `bithuman.AsyncBithuman` |
|---|---|
| `StartFrame` | `AsyncBithuman.create(model_path=..., api_secret=...)`, then `run()` |
| `TTSAudioRawFrame` | `push_audio(pcm, sample_rate, last_chunk=False)` |
| `TTSStoppedFrame` | `flush()` |
| `InterruptionFrame` | `interrupt()` |
| `EndFrame`, `CancelFrame`, cleanup | `shutdown()` |

The SDK side of each call: [Python reference](https://docs.bithuman.ai/platforms/python/reference).

---

# Pipecat troubleshooting

URL: https://docs.bithuman.ai/platforms/pipecat/troubleshooting

> Fix Pipecat avatar start, video and session errors by symptom.

| Symptom | Cause | Fix |
|---|---|---|
| `ValueError: BitHumanVideoService needs an avatar model` | no `model_path=` and no `BITHUMAN_MODEL_PATH` | pass `model_path="avatar.imx"` or `export BITHUMAN_MODEL_PATH=…` |
| `ErrorFrame`: *the bitHuman avatar could not start: ImportError: The bitHuman Python SDK is not installed* | `bithuman` is missing from this environment | activate the venv, then `pip install "pipecat-bithuman[expression-2]"` |
| `ErrorFrame`: *could not start: … api_secret is required* | no secret in the bot's environment | `export BITHUMAN_API_SECRET=…`, or pass `api_secret=` |
| `ErrorFrame`: *could not start: … that API secret was not accepted* | the secret was revoked or mistyped | create a new one under [API secrets](https://www.bithuman.ai/developer/api-keys) |
| `ErrorFrame`: *could not start: … API and SDK access starts at the Creator plan* | from 2026-10-12, a Free account | [choose a plan](https://www.bithuman.ai/pricing?from=docs) |
| `ErrorFrame`: *could not start: … this avatar needs one more package* with an Expression 2 file | the `expression-2` extra is missing | `pip install "pipecat-bithuman[expression-2]"` |
| The bot speaks but shows no video, after an `ErrorFrame` | the avatar failed, and TTS audio passes through by default | fix what the `ErrorFrame` names; `audio_passthrough_on_error=False` drops the voice too |
| No video and no `ErrorFrame` | video out is off in the transport | set `video_out_enabled=True` in the transport's params |
| The picture is resized or stretched | the transport's video size differs from the avatar's | set `video_out_width` and `video_out_height` to the size the service logs on the first frame |
| Log: *queued speech did not finish in … s; closing* | at `EndFrame` the service waits for the queued speech plus `stop_frame_timeout_s` and 1 s, at most `end_drain_timeout_s` (30 s by default), then closes the avatar | if the logged wait equals `end_drain_timeout_s`, raise it |
| Session time keeps running after the user leaves | the pipeline still runs, so the avatar stays open | cancel the pipeline when the participant leaves (`await worker.cancel()`), or end it with `EndFrame` |
| `404 NOT_FOUND` downloading a model | not your agent and not a sample avatar | check the code under [your agents](https://docs.bithuman.ai/api/agents) |

Errors from the Python SDK itself: [Python troubleshooting](https://docs.bithuman.ai/platforms/python/troubleshooting).
