Docs index: /llms.txt · every page as Markdown: add .md

‹ Guides

Add a talking avatar to a voice agent

Give a LiveKit, Pipecat or OpenAI Realtime voice agent a lip-synced face.

Creator plan or higherbitHuman cloudYour computer or serverOn the deviceAPI secretLiveKit plugin 1.8.6Python 2.11.21

What you’ll build

bitHuman turns your voice agent’s reply audio into lip-synced video of an avatar; the agent keeps its own speech recognition, language model and voice. In LiveKit, start a bithuman.AvatarSession beside the AgentSession. In Pipecat, put BitHumanVideoService after the TTS service. With OpenAI Realtime or any other stack, push the reply audio into the Python, Swift or Android SDK.

All three take Essence 2 (a real person) and Expression 2 (any character). You need:

  • a voice agent that produces spoken replies, or the LiveKit or Pipecat example on its platform page;
  • an API secret for an account on the Creator plan or higher;
  • Python 3.10–3.14 for LiveKit, or 3.11–3.14 for Pipecat.

Steps

5 steps

  1. Choose the integration for your stack

    Your stackUseThe avatar renders
    LiveKit Agentslivekit-plugins-bithumanin the bitHuman cloud (avatar_id=), or inside your worker (model_path=)
    Pipecatpipecat-bithumaninside your bot’s process
    OpenAI Realtime or any stack, in Pythonthe Python SDK, AsyncBithumanon your Mac, Linux or Windows machine
    Any stack, in an iPhone, iPad, Mac or Android appthe Swift package or the Android SDKon the device
    No voice stack yeta bitHuman agent in the web embedin the bitHuman cloud, or in the browser tab

    With the web embed, the conversation runs on bitHuman’s servers, even when the avatar renders in the tab.

    Expected

    One row, and the package it names.

  2. Pick the avatar

    Use the wise-pup sample (Expression 2, agent code A23WJF0199) or sofia-ramirez (Essence 2, A52DHS2219) while you build, or create your own. A LiveKit cloud avatar takes the agent code. Pipecat and the Python SDK render from an avatar file, and the sample’s download needs no account:

    curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
    Expected

    An agent code, or wise-pup.imx in the current directory.

  3. LiveKit: start an avatar session beside the agent session

    pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv

    In the worker, keep your API secret as BITHUMAN_MASTER_SECRET and pass the plugin a minted one-hour token, so the secret never reaches the room. The complete agent.py, with the mint call, is on LiveKit: Run your first avatar:

    # excerpt: the avatar lines of agent.py on the LiveKit page
    avatar = bithuman.AvatarSession(
        avatar_id=agent_code,
        api_secret=await livekit_cloud_token(agent_code, ctx.room.name),
    )
    await avatar.start(session, room=ctx.room)
    await session.start(
        agent=Agent(instructions="You are a friendly assistant. Keep answers short."),
        room=ctx.room,
        room_options=RoomOptions(audio_output=False),   # the avatar publishes the audio
    )

    To render inside your worker instead, pass model_path="wise-pup.imx" and install "bithuman[expression-2]" (Voice agent in Python).

    Expected

    python agent.py dev, then join the room from the LiveKit Agents Playground: the avatar appears and answers with its lips in sync.

  4. Pipecat: put the avatar after the TTS service

    pip install "pipecat-bithuman[expression-2]"

    Set BITHUMAN_API_SECRET and BITHUMAN_MODEL_PATH=wise-pup.imx, turn on video_out_enabled=True in your transport, and add the service between tts and transport.output():

    # excerpt: the pipeline, from pipecat-bithuman's examples/bot.py
    avatar = BitHumanVideoService()  # BITHUMAN_MODEL_PATH + BITHUMAN_API_SECRET
    # …
    pipeline = Pipeline(
        [
            transport.input(),
            stt,
            aggregators.user(),
            llm,
            tts,
            avatar,
            transport.output(),
            aggregators.assistant(),
        ]
    )

    The transport and the rest of the bot are on Build a Pipecat bot.

    Expected

    The bot answers with the avatar's video and the avatar's copy of the speech, paired frame by frame.

  5. OpenAI Realtime or any other stack: push the reply audio

    In Python, AsyncBithuman takes each chunk of reply audio with its sample rate and yields frames with the matching audio. Feed it the PCM audio your stack plays, from OpenAI Realtime or a TTS service:

    # excerpt: show() and play() are your own display and audio output
    import asyncio, soundfile as sf
    from bithuman import AsyncBithuman
    
    async def main():
        avatar = await AsyncBithuman.create(model_path="wise-pup.imx")   # reads BITHUMAN_API_SECRET
        pcm, rate = sf.read("speech.wav", dtype="int16")
    
        async def speak():
            for i in range(0, len(pcm), rate // 10):                     # 100 ms chunks, as they arrive
                await avatar.push_audio(pcm[i:i + rate // 10].tobytes(), rate, last_chunk=False)
            await avatar.flush()                                          # end of the reply
    
        task = asyncio.create_task(speak())
        try:
            async for frame in avatar.run():                              # paced at the model's play rate
                if frame.has_image:
                    show(frame.bgr_image)                                 # BGR numpy array
                if frame.audio_chunk:
                    play(frame.audio_chunk.array)                         # audio in sync with the frame
        finally:
            task.cancel()
            await avatar.shutdown()                                       # frees the model and the credential
    
    asyncio.run(main())

    Call flush() when a reply ends and interrupt() when the user talks over it. In an iPhone, iPad, Mac or Android app, resample the reply to 16 kHz mono first: Resample speech to 16 kHz. To run OpenAI Realtime with no OpenAI key of your own, use the Realtime relay.

    Expected

    Idle frames until the reply arrives, then the lips follow its words, with the audio in step.

How it works

The avatar does not listen or think: it renders the speech your agent already produces. With a LiveKit cloud avatar, the audio reaches bitHuman and the avatar renders in the bitHuman cloud, in the US. When it renders in your worker, bot or app, its audio and video stay with you; bitHuman receives a credential check and usage reports. A session bills active session time, talking or idle (pricing).

Make it your own

  • Interruptions: stop the avatar mid-sentence when the user talks: Barge-in.
  • Your own character: one portrait makes an avatar: Create your own avatar.
  • Everything on one computer: a local LiveKit server and the CLI: Voice agent.

Troubleshooting

SymptomCauseFix
Two voices play in LiveKitthe agent session also publishes audioset room_options=RoomOptions(audio_output=False)
The mint call returns 401a missing or invalid API secretcheck BITHUMAN_MASTER_SECRET against your API secret
The Pipecat bot speaks with no video and no ErrorFramevideo out is off in the transportset video_out_enabled=True in the transport’s params
The avatar refuses to startno API secret, or a plan below Creatorset the secret for an account on the Creator plan or higher
The mouth runs slow in an appthe speech is not 16 kHzresample it before you feed it

More on LiveKit: Troubleshooting and Pipecat: Troubleshooting.