LiveKit

Give a LiveKit voice agent a face with the bitHuman Python plugin: a cloud-rendered avatar in your room, or an avatar rendered on your own server.

bitHuman cloud Your servers API secret LiveKit plugin 1.8.4

livekit-plugins-bithuman adds a bitHuman avatar to any LiveKit Agents worker, on LiveKit Cloud or your own LiveKit server. It is a Python plugin; there is no Node.js plugin. The avatar joins the room as a participant, in one of two ways:

A bitHuman cloud avatarThe avatar on your server
You passavatar_id= (an agent code)model_path= (an avatar file)
The avatar rendersin the bitHuman cloud, in the USinside your worker’s process
Its audio and videoare published into your room by bitHumanstay on your machine until your worker publishes them
Pricethe cloud rate (pricing)the self-hosted rate
Guidethis pageVoice agent

Either way, the plugin serves the agent’s own model, Essence 2 or Expression 2.

Before you start

  • Python 3.10–3.14 and a LiveKit project (its URL and credentials).
  • A bitHuman agent code: A23WJF0199 (the wise-pup sample) or your own from Agents.
  • An API secret, and an OpenAI key for the voice model in the example.

Install

pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv

Authenticate

In a LiveKit worker, name the secret BITHUMAN_MASTER_SECRET and pass a minted token. For a cloud avatar, the worker mints a one-hour token that can start only this agent’s avatar in this room (POST /v1/runtime-tokens/mint with "scope": "livekit-cloud"), as in the worker below. The plugin’s PyPI page, which LiveKit maintains, sets BITHUMAN_API_SECRET instead; for a cloud avatar, use the minted token shown here.

export BITHUMAN_MASTER_SECRET="<your API secret>"
export BITHUMAN_AGENT_ID=A23WJF0199
export LIVEKIT_URL=wss://your-project.livekit.cloud
export LIVEKIT_API_KEY=… LIVEKIT_API_SECRET=…
export OPENAI_API_KEY=…
YOUR SERVERYour agent workerBITHUMAN_MASTER_SECRET stays in this processBITHUMAN CLOUDThe cloud avatarstarts with a one-hour token for this agent androomYOUR LIVEKIT ROOMYour appjoins with a room token from your server, nevera bitHuman secretmint a tokenvideo and voice
LiveKit: the secret stays on your server. Your LiveKit agent worker holds the API secret as BITHUMAN_MASTER_SECRET and mints a one-hour runtime token for each session. The bitHuman cloud avatar joins your room with that token and publishes its video and voice. Your app joins the room with a room token from your server and never holds a bitHuman secret.

First frame

A complete worker (agent.py):

import os

import aiohttp
from dotenv import load_dotenv
from livekit.agents import Agent, AgentSession, JobContext, WorkerOptions, cli
from livekit.agents.voice.room_io import RoomOptions
from livekit.plugins import bithuman, openai, silero
from openai.types.realtime.realtime_audio_input_turn_detection import ServerVad

load_dotenv()


async def livekit_cloud_token(agent_code: str, room_name: str) -> str:
    """A one-hour token that can only start this agent's avatar in this room."""
    async with aiohttp.ClientSession() as http:
        async with http.post(
            "https://api.bithuman.ai/v1/runtime-tokens/mint",
            headers={"api-secret": os.environ["BITHUMAN_MASTER_SECRET"]},
            json={"agent_code": agent_code, "scope": "livekit-cloud",
                  "room_name": room_name, "livekit_url": os.environ["LIVEKIT_URL"]},
        ) as resp:
            resp.raise_for_status()
            return (await resp.json())["scoped_token"]


async def entrypoint(ctx: JobContext):
    await ctx.connect()
    await ctx.wait_for_participant()
    agent_code = os.environ["BITHUMAN_AGENT_ID"]

    session = AgentSession(
        llm=openai.realtime.RealtimeModel(
            voice="coral",
            # end the user's turn after 0.5 s of silence (the default waits longer)
            turn_detection=ServerVad(type="server_vad", silence_duration_ms=500, create_response=True, interrupt_response=True),
        ),
        vad=silero.VAD.load(),
    )
    avatar = bithuman.AvatarSession(
        avatar_id=agent_code,
        api_secret=await livekit_cloud_token(agent_code, ctx.room.name),
    )
    await avatar.start(session, room=ctx.room)
    await session.start(
        agent=Agent(instructions="You are a friendly assistant. Keep answers short."),
        room=ctx.room,
        room_options=RoomOptions(audio_output=False),   # the avatar publishes the audio
    )


if __name__ == "__main__":
    cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
python agent.py dev
# → join the room from the LiveKit Agents Playground; the avatar appears and answers

Integrate into your app

  • Your own client. The avatar is a normal LiveKit participant: any LiveKit client SDK (JavaScript, Swift, Kotlin) subscribes to its video and audio tracks. The app takes a room token from your server, never a bitHuman secret.
  • Choosing a model. The plugin serves the agent’s own model. Do not pass model=; create the agent with the model you want (Models).
  • A photo instead of an agent. avatar_image= with no avatar_id animates the photo on Expression 1 only. On every other model the launch is refused with 400 VALIDATION_ERROR before anything is billed: create an agent from the photo and pass its code as avatar_id.
  • Rendering on your own machine. Pass model_path= (an avatar file) instead of avatar_id=, and the secret explicitly: api_secret=os.environ["BITHUMAN_MASTER_SECRET"]. The avatar renders inside the worker’s process, and the secret stays in that process. Runnable example: Voice agent.
  • Without the plugin. If your voice pipeline is not a LiveKit Agents worker, start the avatar with one REST call and stream it your audio: Cloud avatar without the plugin.
  • Several agents in one room. The avatar lip-syncs the agent that calls AvatarSession.start() and ignores other agents’ audio.
  • Gestures. Trigger avatar actions from your agent: Gestures.

The cloud example packages this worker with a web UI and Docker Compose.

Platform notes

  • The mint call is one per session. The token starts that session only; it expires after an hour, and a session that runs longer continues.
  • livekit_url in the mint call must be the URL the plugin connects to (LIVEKIT_URL, unless you pass livekit_url= to AvatarSession.start()).
  • If you run your own LiveKit server, use livekit-server 1.9.12 or newer. With older servers, browsers leave and rejoin the room every 15 s, and the video stalls each time.

Performance

A cloud avatar renders in the bitHuman cloud; with model_path= the avatar renders where your worker runs, as Python does:

ConfigurationEssence 2Expression 2
NVIDIA RTX 4090 Cloud API · GPU
4.1× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-27
17.0× real timeNVIDIA RTX 4090 · cloud API · measured 2026-09-23
Intel Core i7-13700F (x86_64) Linux · Python CPU only (no GPU)
1.9× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26
2.3× real timeIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26
Apple M4 macOS · Python
6.9× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25
8.4× real timeApple M4 · bithuman 2.11.12 · measured 2026-09-25

Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.

Troubleshooting

SymptomCauseFix
The mint call returns 403the token was minted for another agent, room or LiveKit URLmint with the same agent_code, room_name and livekit_url the plugin uses
The mint call returns 401a missing or invalid API secretcheck BITHUMAN_MASTER_SECRET
The avatar speaks with the wrong modelthe plugin serves the agent’s own modelcreate an agent with the model you want
The video stalls for 1–2 s every 15 s, or a LiveKit Meet tile goes blacklivekit-server older than 1.9.12: the browser leaves and rejoins the room every 15 sbrew upgrade livekit (macOS) or curl -sSL https://get.livekit.io | bash (Linux), then restart livekit-server
Two voices playthe agent session also publishes audioset room_options=RoomOptions(audio_output=False)

Reference