# Add a talking avatar to a voice agent

URL: https://docs.bithuman.ai/build/how-to/voice-agent-avatar

> Give a LiveKit, Pipecat or OpenAI Realtime voice agent a lip-synced face.

## What you'll build

bitHuman turns your voice agent's reply audio into lip-synced video of an avatar; the agent keeps its own speech recognition, language model and voice. In LiveKit, start a `bithuman.AvatarSession` beside the `AgentSession`. In Pipecat, put `BitHumanVideoService` after the TTS service. With OpenAI Realtime or any other stack, push the reply audio into the Python, Swift or Android SDK.

All three take Essence 2 (a real person) and Expression 2 (any character). You need:

- a voice agent that produces spoken replies, or the LiveKit or Pipecat example on its platform page;
- an [API secret](https://docs.bithuman.ai/start/api-secret) for an account on the Creator plan or higher;
- Python 3.10–3.14 for LiveKit, or 3.11–3.14 for Pipecat.

## Steps

### Choose the integration for your stack

| Your stack | Use | The avatar renders |
|---|---|---|
| [LiveKit Agents](https://docs.bithuman.ai/platforms/livekit) | `livekit-plugins-bithuman` | in the bitHuman cloud (`avatar_id=`), or inside your worker (`model_path=`) |
| [Pipecat](https://docs.bithuman.ai/platforms/pipecat) | `pipecat-bithuman` | inside your bot's process |
| OpenAI Realtime or any stack, in [Python](https://docs.bithuman.ai/platforms/python/app) | the Python SDK, `AsyncBithuman` | on your Mac, Linux or Windows machine |
| Any stack, in an [iPhone, iPad, Mac or Android app](https://docs.bithuman.ai/build/how-to/voice-assistant-face) | the Swift package or the Android SDK | on the device |
| No voice stack yet | a bitHuman agent in the [web embed](https://docs.bithuman.ai/platforms/web) | in the bitHuman cloud, or in the browser tab |

With the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab.

Expected:

One row, and the package it names.

### Pick the avatar

Use the `wise-pup` sample (Expression 2, agent code `A23WJF0199`) or `sofia-ramirez` (Essence 2, `A52DHS2219`) while you build, or [create your own](https://docs.bithuman.ai/build/create-avatar). A LiveKit cloud avatar takes the agent code. Pipecat and the Python SDK render from an avatar file, and the sample's download needs no account:

```bash
curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2"
```

Expected:

An agent code, or `wise-pup.imx` in the current directory.

### LiveKit: start an avatar session beside the agent session

```bash
pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv
```

In the worker, keep your API secret as `BITHUMAN_MASTER_SECRET` and pass the plugin a minted one-hour token, so the secret never reaches the room. The complete `agent.py`, with the mint call, is on [LiveKit: Run your first avatar](https://docs.bithuman.ai/platforms/livekit#run-your-first-avatar):

```python
# excerpt: the avatar lines of agent.py on the LiveKit page
avatar = bithuman.AvatarSession(
    avatar_id=agent_code,
    api_secret=await livekit_cloud_token(agent_code, ctx.room.name),
)
await avatar.start(session, room=ctx.room)
await session.start(
    agent=Agent(instructions="You are a friendly assistant. Keep answers short."),
    room=ctx.room,
    room_options=RoomOptions(audio_output=False),   # the avatar publishes the audio
)
```

To render inside your worker instead, pass `model_path="wise-pup.imx"` and install `"bithuman[expression-2]"` ([Voice agent in Python](https://docs.bithuman.ai/build/voice-agent/python)).

Expected:

`python agent.py dev`, then join the room from the LiveKit Agents Playground: the avatar appears and answers with its lips in sync.

### Pipecat: put the avatar after the TTS service

```bash
pip install "pipecat-bithuman[expression-2]"
```

Set `BITHUMAN_API_SECRET` and `BITHUMAN_MODEL_PATH=wise-pup.imx`, turn on `video_out_enabled=True` in your transport, and add the service between `tts` and `transport.output()`:

```python
# excerpt: the pipeline, from pipecat-bithuman's examples/bot.py
avatar = BitHumanVideoService()  # BITHUMAN_MODEL_PATH + BITHUMAN_API_SECRET
# …
pipeline = Pipeline(
    [
        transport.input(),
        stt,
        aggregators.user(),
        llm,
        tts,
        avatar,
        transport.output(),
        aggregators.assistant(),
    ]
)
```

The transport and the rest of the bot are on [Build a Pipecat bot](https://docs.bithuman.ai/platforms/pipecat/app#complete-example).

Expected:

The bot answers with the avatar's video and the avatar's copy of the speech, paired frame by frame.

### OpenAI Realtime or any other stack: push the reply audio

In Python, `AsyncBithuman` takes each chunk of reply audio with its sample rate and yields frames with the matching audio. Feed it the PCM audio your stack plays, from OpenAI Realtime or a TTS service:

```python
# excerpt: show() and play() are your own display and audio output
import asyncio, soundfile as sf
from bithuman import AsyncBithuman

async def main():
    avatar = await AsyncBithuman.create(model_path="wise-pup.imx")   # reads BITHUMAN_API_SECRET
    pcm, rate = sf.read("speech.wav", dtype="int16")

    async def speak():
        for i in range(0, len(pcm), rate // 10):                     # 100 ms chunks, as they arrive
            await avatar.push_audio(pcm[i:i + rate // 10].tobytes(), rate, last_chunk=False)
        await avatar.flush()                                          # end of the reply

    task = asyncio.create_task(speak())
    try:
        async for frame in avatar.run():                              # paced at the model's play rate
            if frame.has_image:
                show(frame.bgr_image)                                 # BGR numpy array
            if frame.audio_chunk:
                play(frame.audio_chunk.array)                         # audio in sync with the frame
    finally:
        task.cancel()
        await avatar.shutdown()                                       # frees the model and the credential

asyncio.run(main())
```

Call `flush()` when a reply ends and `interrupt()` when the user talks over it. In an iPhone, iPad, Mac or Android app, resample the reply to 16 kHz mono first: [Resample speech to 16 kHz](https://docs.bithuman.ai/build/how-to/voice-assistant-face#resample-speech-to-16-khz). To run OpenAI Realtime with no OpenAI key of your own, use the [Realtime relay](https://docs.bithuman.ai/api/realtime).

Expected:

Idle frames until the reply arrives, then the lips follow its words, with the audio in step.

## How it works

The avatar does not listen or think: it renders the speech your agent already produces. With a LiveKit cloud avatar, the audio reaches bitHuman and the avatar renders in the bitHuman cloud, in the US. When it renders in your worker, bot or app, its audio and video stay with you; bitHuman receives a credential check and usage reports. A session bills active session time, talking or idle ([pricing](https://docs.bithuman.ai/pricing)).

## Make it your own

- **Interruptions:** stop the avatar mid-sentence when the user talks: [Barge-in](https://docs.bithuman.ai/build/barge-in).
- **Your own character:** one portrait makes an avatar: [Create your own avatar](https://docs.bithuman.ai/build/create-avatar).
- **Everything on one computer:** a local LiveKit server and the CLI: [Voice agent](https://docs.bithuman.ai/build/voice-agent).

## Troubleshooting

| Symptom | Cause | Fix |
|---|---|---|
| Two voices play in LiveKit | the agent session also publishes audio | set `room_options=RoomOptions(audio_output=False)` |
| The mint call returns `401` | a missing or invalid API secret | check `BITHUMAN_MASTER_SECRET` against your [API secret](https://docs.bithuman.ai/start/api-secret) |
| The Pipecat bot speaks with no video and no `ErrorFrame` | video out is off in the transport | set `video_out_enabled=True` in the transport's params |
| The avatar refuses to start | no API secret, or a plan below Creator | set the secret for an account on the Creator plan or higher |
| The mouth runs slow in an app | the speech is not 16 kHz | [resample it](https://docs.bithuman.ai/build/how-to/voice-assistant-face#resample-speech-to-16-khz) before you feed it |

More on [LiveKit: Troubleshooting](https://docs.bithuman.ai/platforms/livekit/troubleshooting) and [Pipecat: Troubleshooting](https://docs.bithuman.ai/platforms/pipecat/troubleshooting).
