Cloud avatar in your room
More ▾
Start a bitHuman cloud avatar in your own LiveKit room with one REST call, then stream it your TTS audio. No LiveKit plugin or agent framework needed.
Overview
A cloud avatar is a participant in your LiveKit room. It lip-syncs the audio you send it, and it publishes the video and that audio as its own tracks. The LiveKit plugin does all of this for a Python LiveKit Agents worker. Use this page when your voice pipeline is something else. You will:
- Mint a LiveKit join token for the avatar.
- Start the session with
POST /v1/runtime-tokens/request. The avatar then joins your room. - Send each reply’s audio to the avatar as a LiveKit byte stream.
Your clients subscribe to the avatar like any other participant. Call the endpoint from your server, because it takes your API secret.
Mint the avatar’s join token
Mint the token with your LiveKit credentials, not with a bitHuman credential. It needs:
- Identity:
bithuman-avatar-agent. - Kind:
agent. - Grant: permission to join the room.
- Attribute
lk.publish_on_behalf: the identity of the participant that sends the audio. The avatar takes audio only from that participant. If you leave this attribute out, the avatar takes audio from the first agent participant in the room.
from livekit import api
avatar_token = (
api.AccessToken(LIVEKIT_API_KEY, LIVEKIT_API_SECRET)
.with_identity("bithuman-avatar-agent")
.with_kind("agent")
.with_grants(api.VideoGrants(room_join=True, room=room_name))
.with_attributes({"lk.publish_on_behalf": sender_identity})
.to_jwt()
)
Don’t put a bitHuman secret in the token or in its attributes. Everyone in the room can read attributes.
Join your sender to the room as an agent participant too (kind agent). The avatar treats standard participants as users, and it stays in the room while any user is there.
Start the session
POST https://api.bithuman.ai/v1/runtime-tokens/request with the api-secret header:
curl
curl -X POST https://api.bithuman.ai/v1/runtime-tokens/request \
-H "api-secret: $BITHUMAN_API_SECRET" \
-H "content-type: application/json" \
-d '{
"mode": "gpu",
"model": "essence-2",
"agent_id": "<your agent code>",
"livekit_url": "wss://your-project.livekit.cloud",
"livekit_token": "<the avatar join token>",
"room_name": "your-room"
}'
# → {"avatar_session_started": true, "model": "essence-2", ...}
Python
import os, requests
resp = requests.post(
"https://api.bithuman.ai/v1/runtime-tokens/request",
headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]},
json={
"mode": "gpu",
"model": "essence-2",
"agent_id": "<your agent code>",
"livekit_url": "wss://your-project.livekit.cloud",
"livekit_token": "<the avatar join token>",
"room_name": "your-room",
},
)
print(resp.json())
Node
const resp = await fetch("https://api.bithuman.ai/v1/runtime-tokens/request", {
method: "POST",
headers: {
"api-secret": process.env.BITHUMAN_API_SECRET,
"content-type": "application/json",
},
body: JSON.stringify({
mode: "gpu",
model: "essence-2",
agent_id: "<your agent code>",
livekit_url: "wss://your-project.livekit.cloud",
livekit_token: "<the avatar join token>",
room_name: "your-room",
}),
});
console.log(await resp.json());
| Field | Required | What it is |
|---|---|---|
mode | yes | "gpu". This field is what makes the call start an avatar; without it, the endpoint only issues a runtime token. The older value "cpu" also works but requires agent_id. |
model | recommended | essence-1, essence-2, expression-1 or expression-2. If you leave it out, the server picks the model from mode and the agent’s models, so check model in the response. |
agent_id | yes, unless image is sent | The agent code. |
image | for photo sessions | A portrait, for Expression 1 only. See A photo instead of an agent. |
livekit_url | yes | Your LiveKit server URL. |
livekit_token | yes | The avatar’s join token from the step above. |
room_name | yes | The room the avatar joins. |
A 200 response means the session is starting. avatar_session_started is true, and model is the model the session launched as. The response also echoes mode, agent_id and image, and includes your user_id.
The session is billed at the model’s cloud rate for as long as it runs (Pricing). It ends in any of these cases:
- The last user leaves the room.
- You remove
bithuman-avatar-agentfrom the room. - The room closes.
Send the audio
Send each reply as one LiveKit byte stream:
- Topic:
lk.audio_stream. - Destination identity:
bithuman-avatar-agent. - Attributes:
sample_rateandnum_channels, as strings. - Payload: raw 16-bit little-endian PCM.
Closing the stream marks the end of the reply. Mono 16 kHz is the native format. Other sample rates are resampled, and stereo is mixed down to mono.
writer = await room.local_participant.stream_bytes(
name="reply-1",
topic="lk.audio_stream",
destination_identities=["bithuman-avatar-agent"],
attributes={"sample_rate": "16000", "num_channels": "1"},
)
async for pcm in tts_audio(): # 16-bit mono PCM chunks, as the TTS produces them
await writer.write(pcm)
await writer.aclose() # end of this reply
- First reply: wait until
bithuman-avatar-agenthas joined before you send it. Audio sent before the avatar joins is dropped. - Interrupting: perform the RPC
lk.clear_bufferonbithuman-avatar-agent, then close the stream. - Playback events (optional): the avatar calls the RPCs
lk.playback_startedandlk.playback_finishedon your sender. Register handlers for them if you want to know when it speaks. If you don’t register them, the avatar skips these calls.
Keep latency low
- Write each TTS chunk as soon as it exists. Open the stream on the first chunk. Don’t pace the audio to real time: the avatar buffers it and plays it at the right speed. The size of each write doesn’t matter.
- Close the stream as soon as the reply ends. The avatar renders speech in blocks of up to about a second of audio. A reply shorter than one block, like “Sure!”, waits for the close.
- Keep one session for the whole conversation. Starting a session connects to the room and loads the avatar, so don’t start a new one for each turn.
- Run your pipeline in the US. Cloud avatars render in the US. Put your LiveKit server or LiveKit Cloud project and your STT, LLM and TTS in a US region; US East is a good default.
- Speed up the voice pipeline. Most of a slow turn is usually spent there: end-of-speech detection, the LLM’s first token and the TTS’s first audio. Streaming STT and streaming TTS help the most.
A photo instead of an agent
Expression 1 can animate a photo without an agent. Send image, leave out agent_id, and set "model": "expression-1".
image is a string in one of these forms:
- A public
http(s)URL. The server fetches it with a plain GET, so it must return200without authentication. A pre-signed URL works. - The image as base64, either raw or as a
data:image/jpeg;base64,…URI.
The endpoint takes JSON only.
Photo requirements:
- A JPEG or PNG.
- One person, facing the camera, with the face clearly visible.
- Head and shoulders, with room around the head.
The server finds the face and crops a square about twice the face’s width, with a little extra room above the head. It scales the crop to 512×512, so a face at least ~256 px wide keeps full detail. A face near the edge of the photo gets a clipped crop. If no face is found, the whole photo is used as it is, and a photo that isn’t square looks stretched.
Photo sessions work on Expression 1 only. On every other model, the request is refused with 400 VALIDATION_ERROR, and nothing is launched or billed. To use another model, create an agent from the photo first.
Errors
| Status | Code | Cause |
|---|---|---|
400 | VALIDATION_ERROR | One of these: a field is missing or invalid; image was sent without agent_id for a model other than Expression 1; an Essence 1 session was started without livekit_url, livekit_token and room_name. Nothing is launched or billed. |
401 | MISSING_AUTH / UNAUTHORIZED | The api-secret header is missing or invalid. |
403 | PLAN_REQUIRED | The model is not in your plan. |
403 | CONCURRENCY_LIMIT_REACHED | The session would exceed your plan’s concurrent sessions (Rate limits). |
404 | NOT_FOUND | No agent with this code that you can start. |
409 | VALIDATION_ERROR | The agent can’t be served as the requested model, or its own model isn’t ready yet. |
503 | SERVICE_UNAVAILABLE | No capacity right now, or a temporary failure. Retry after the Retry-After header. |
All error codes are listed on Errors.