Docs index: /llms.txt · every page as Markdown: add .md

‹ Platforms

Build a Pipecat bot

Fit BitHumanVideoService into your Pipecat pipeline, transport and error handling.

Integrate into your app

  • Where it goes. After the TTS service, or a speech-to-speech LLM, and before transport.output().
  • The transport. Turn on video out with video_out_enabled=True. The service logs the avatar’s frame size on the first frame; set video_out_width and video_out_height to it to avoid resizing.
  • Voice and picture together. The service does not forward your TTS audio. It pushes the avatar’s copy of the speech, paired with each picture. Each image carries sync_with_audio, so keep the transport’s video_out_is_live off (its default).
  • Barge-in. On InterruptionFrame the avatar drops the reply in flight and goes back to idle. See Barge-in.
  • The end of a reply. TTSStoppedFrame is held until the avatar has finished speaking the reply.
  • Session time. The avatar opens on StartFrame and closes on EndFrame, CancelFrame or cleanup; it is billed while open, talking or idle. End the pipeline when the user leaves, as the example below does.
  • When the avatar fails. The service pushes one ErrorFrame upstream and, by default, passes TTS audio through unchanged, so the bot keeps talking without video. audio_passthrough_on_error=False drops the audio instead.
  • Metrics. With enable_metrics=True in PipelineParams, the service reports TTFB: from a reply’s first audio in to its first voiced avatar frame out.

What the service does with frames

InOut
TTSAudioRawFramesent to the avatar, not forwarded as is
(avatar frame)OutputImageRawFrame (RGB), while talking and while idle
(avatar audio)TTSAudioRawFrame (16 kHz mono), paired with each picture
TTSStoppedFrameheld until the avatar has finished speaking the reply
InterruptionFramethe avatar drops the reply in flight and goes back to idle
anything elsepassed on unchanged

Complete example

examples/bot.py in pipecat-bithuman is a complete voice bot in a Daily room, with Deepgram speech-to-text, an OpenAI LLM and Cartesia TTS. The avatar sits between the TTS and the output transport:

# excerpt: the transport and the pipeline, from pipecat-bithuman's examples/bot.py
transport = DailyTransport(
    os.environ["DAILY_ROOM_URL"],
    None,
    "Pip",
    DailyParams(
        audio_in_enabled=True,
        audio_out_enabled=True,
        video_out_enabled=True,
        video_out_width=416,   # the wise-pup sample's frame size; the service logs yours
        video_out_height=720,
    ),
)
# …
avatar = BitHumanVideoService()  # BITHUMAN_MODEL_PATH + BITHUMAN_API_SECRET
# …
pipeline = Pipeline(
    [
        transport.input(),
        stt,
        aggregators.user(),
        llm,
        tts,
        avatar,
        transport.output(),
        aggregators.assistant(),
    ]
)
# …
@transport.event_handler("on_participant_left")
async def on_left(transport, participant, reason):
    await worker.cancel()  # close the avatar: session time stops

Run it from the repository, with the avatar file from Pipecat. examples/bot.py sets the transport to the wise-pup frame size, 416×720; for another avatar, use the size the service logs on the first frame:

pip install "pipecat-bithuman[expression-2]" "pipecat-ai[daily,deepgram,openai,cartesia,silero]"
export BITHUMAN_API_SECRET="<your API secret>" BITHUMAN_MODEL_PATH=wise-pup.imx
export DAILY_ROOM_URL=… DEEPGRAM_API_KEY=… OPENAI_API_KEY=… CARTESIA_API_KEY=… CARTESIA_VOICE_ID=…
python examples/bot.py
# → join the Daily room; the avatar says hello first, then answers

Platform notes

  • Requires Python 3.11–3.14 and pipecat-ai 1.12.0 or newer.
  • Your own avatar: download an agent’s model and pass the .imx file as model_path=.
  • The service has no settings to change at runtime. The avatar is chosen when the service is built.
  • For tests, or to wrap the SDK, pass runtime_factory=: an async callable that returns an object implementing BitHumanRuntime. The package’s own tests run on fakes, with no network, secret or model file.
  • Report bugs in the pipecat-bithuman issues. The Pipecat team does not maintain this package.

Reference

BitHumanVideoService options

ArgumentDefaultMeaning
model_pathBITHUMAN_MODEL_PATHthe .imx avatar file
api_secretBITHUMAN_API_SECRETthe API secret, read by the SDK when not passed
sync_video_to_audioTruesets sync_with_audio on each image
audio_passthrough_on_errorTruekeep the voice if the avatar fails
stop_frame_timeout_s2.0release a held TTSStoppedFrame after this much quiet
end_drain_timeout_s30.0the longest wait on EndFrame for queued speech
runtime_factorythe SDKadvanced: your own BitHumanRuntime

Properties: is_avatar_ready (the avatar is open and rendering) and frame_size ((width, height) of the first frame, or None before it).

How it maps to the Python SDK

Pipecatbithuman.AsyncBithuman
StartFrameAsyncBithuman.create(model_path=..., api_secret=...), then run()
TTSAudioRawFramepush_audio(pcm, sample_rate, last_chunk=False)
TTSStoppedFrameflush()
InterruptionFrameinterrupt()
EndFrame, CancelFrame, cleanupshutdown()

The SDK side of each call: Python reference.