# bitHuman — Build (voice agent, companion app, kiosk, talking video, avatars, personas, voices, gestures, MCP, troubleshooting) > Realtime talking avatars from one portrait. Index: https://docs.bithuman.ai/llms.txt · every page is also served as markdown at .md · OpenAPI: https://docs.bithuman.ai/api/openapi.yaml · other sections: https://docs.bithuman.ai/llms/start.txt · https://docs.bithuman.ai/llms/platforms.txt · https://docs.bithuman.ai/llms/deploy.txt · https://docs.bithuman.ai/llms/models.txt · https://docs.bithuman.ai/llms/api.txt ## Instructions for AI agents - Always send `model` ("essence-2" or "expression-2") when creating an agent; poll until `status` is `ready` or `failed`. - API and SDK use requires the Creator plan or higher. Never tell a user they can build on a free plan. - In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET` and pass a minted token (https://docs.bithuman.ai/platforms/livekit.md). - Say where things happen: the avatar renders (device, browser, server, cloud); the conversation runs (your stack, the CLI's local conversation brain, or bitHuman's servers). - Quote speed only from https://docs.bithuman.ai/performance.json (× real time, with the device). Quote prices only from GET https://api.bithuman.ai/v1/pricing. - Do not claim offline on phones, Mac or the browser, a conversation brain on phones, or any certification. Where each model runs, the key facts and one command per path: https://docs.bithuman.ai/llms.txt ## Contents - Create your own avatar — https://docs.bithuman.ai/build/create-avatar - Persona — https://docs.bithuman.ai/build/persona - Voices — https://docs.bithuman.ai/build/voices - Gestures (Essence 1) — https://docs.bithuman.ai/build/gestures - Voice agent — https://docs.bithuman.ai/build/voice-agent - Website widget — https://docs.bithuman.ai/build/website-widget - Companion app — https://docs.bithuman.ai/build/companion-app - Kiosk on a Linux PC — https://docs.bithuman.ai/build/kiosk - Talking video — https://docs.bithuman.ai/build/talking-video - Claude & Cursor (MCP) — https://docs.bithuman.ai/build/mcp - Troubleshooting — https://docs.bithuman.ai/resources/troubleshooting Linked, not inlined (read the .md twin): - Examples: https://docs.bithuman.ai/examples.md · changelog: https://docs.bithuman.ai/changelog.md · API references: https://docs.bithuman.ai/platforms/cli/reference.md, https://docs.bithuman.ai/platforms/python/reference.md, https://docs.bithuman.ai/platforms/swift/reference.md, https://docs.bithuman.ai/platforms/android/reference.md --- # Create your own avatar URL: https://docs.bithuman.ai/build/create-avatar > Turn a portrait, a voice sample and a prompt into your own avatar: what to upload, how to write the persona, and the API call. An avatar is a face, a voice and a personality, packaged as one agent with a short code. Use a [sample avatar](https://docs.bithuman.ai/examples#ready-made-avatars) to start, or create your own from one portrait. *Diagram: Creating an avatar.* You upload one portrait. Avatar creation happens in the bitHuman cloud and takes about 2 to 2.5 hours. The finished avatar model then runs on your devices (iPhone, iPad, Mac, Android, a Linux PC or a WebGPU browser) or in the bitHuman cloud. ## Before you start - An [API secret](https://docs.bithuman.ai/start/api-secret) on the Creator plan or higher, and credits for the creation ([pricing](https://docs.bithuman.ai/pricing#creation--one-time-credits)). - A portrait image at a public URL (or create one from a prompt). - Optionally, 3–10 seconds of clean speech for voice cloning. ## 1. Choose a model | You have | Use | |---|---| | A photo of a real person | `essence-2`: photoreal, up to 1080p | | A character, animal, robot or illustration | `expression-2`: any subject | | Not sure | `auto`: people go to `essence-2`, everything else to `expression-2` | More on the difference: [Models](https://docs.bithuman.ai/models). ## 2. Pick the inputs | Input | Use for | Limits | |---|---|---| | Image | the face | under 10 MB; one clear figure, neutral expression, facing the camera, face unobstructed | | Voice | voice cloning | under 1 minute of clean speech (MP3, WAV or M4A), no music | | Prompt | the personality | required when there is no image | ### What makes a good photo One subject, in focus, facing the camera with a resting expression and the whole face visible (eyes, nose and mouth). Photos are not checked; these are the conditions the models are built for. For an animal, use a well-lit, front-facing photo with the face filling the frame. Without a voice sample, a voice is generated to match the persona. Without a prompt, a persona is generated from the image. ## 3. Create the agent ```bash curl -X POST https://api.bithuman.ai/v1/agent/generate \ -H "Content-Type: application/json" -H "api-secret: $BITHUMAN_API_SECRET" \ -d '{"model": "auto", "prompt": "You are a friendly receptionist.", "image": "https://example.com/headshot.jpg"}' ``` The response carries an `agent_id`. Creation takes about 2–2.5 hours for the second-generation models. You can also create an agent in the dashboard at [bithuman.ai](https://www.bithuman.ai/explore), which starts on Expression 2. ## 4. Wait until it is ready Poll [`GET /v1/agent/status/{agent_id}`](https://docs.bithuman.ai/api/agents#poll-status) every few seconds until `status` is `ready` or `failed`. ## Check it worked Open `https://www.bithuman.ai/embed/` in a browser and talk to it, or `bithuman run `. ## Variations - **Write a better persona** with [the persona guide](https://docs.bithuman.ai/build/persona). - **Add another model** later without re-creating the agent: [add a model](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent). - **Change the prompt** at any time: [update an agent](https://docs.bithuman.ai/api/agents#update-an-agent). ## Troubleshooting | Symptom | Cause | Fix | |---|---|---| | `402 INSUFFICIENT_BALANCE` | not enough credits | top up or choose a plan | | `422 MODEL_SUBJECT_MISMATCH` | `essence-2` for a subject that is not a photoreal person | use `expression-2` or `auto` | | `failed` with an image error | the image URL is not publicly fetchable | host the image publicly and create again (the failed creation is refunded) | | The likeness is off | a side profile, several people or poor light | crop to one front-facing person in good light | | The voice sounds noisy | background noise or music in the sample | re-record in a quiet room | ## Next - [Agents API](https://docs.bithuman.ai/api/agents) · [Persona](https://docs.bithuman.ai/build/persona) · [Voices](https://docs.bithuman.ai/build/voices) --- # Persona URL: https://docs.bithuman.ai/build/persona > Write the system prompt that gives an avatar its personality, with the CO-STAR framework and a worked example. The persona is the avatar's system prompt: who it is, what it is for, and how it answers. Set it as `prompt` when you [create the agent](https://docs.bithuman.ai/build/create-avatar), or change it later by sending it as `system_prompt` ([update an agent](https://docs.bithuman.ai/api/agents#update-an-agent)): ```bash curl -X POST https://api.bithuman.ai/v1/agent/ -H "Content-Type: application/json" \ -H "api-secret: $BITHUMAN_API_SECRET" -d '{"system_prompt": "CONTEXT: ..."}' ``` ## Before you start - An agent, or one you are about to create. ## 1. Fill in the six fields | Field | Defines | Weak → strong | |---|---|---| | **C**ontext | setting and situation | "customer service" → "second-level support for a cloud product, handling escalated cases" | | **O**bjective | the goal | "be helpful" → "resolve the issue in the first conversation" | | **S**tyle | how it communicates | "professional" → "like a patient in-store technician who uses analogies" | | **T**one | emotional attitude | patient, empathetic, calm under frustration | | **A**udience | who it talks to | "everyone" → "everyday users, beginner to intermediate" | | **R**esponse | the shape of answers | "acknowledge → clarify → step by step → confirm → offer more" | ## 2. Write it as a prompt ```text CONTEXT: Online tutor helping high-school students with exam-season math. OBJECTIVE: Explain concepts clearly, solve specific problems, build confidence. STYLE: Like an award-winning teacher: real-world examples, step by step. TONE: Encouraging and patient; mistakes are part of learning. AUDIENCE: Ages 14–18, mixed ability, some test anxiety. RESPONSE: Acknowledge, break into steps, encourage, use an analogy, close with confidence. ``` Keep spoken answers short: an avatar speaks every word, so two or three sentences per turn feel natural. ## Check it worked Talk to the agent (`https://www.bithuman.ai/embed/`) and ask three questions your users would ask. Adjust the field that produced the weakest answer. ## Troubleshooting | Symptom | Fix | |---|---| | Answers are long and read like text | add "Answer in two or three spoken sentences" to RESPONSE | | The tone swings between formal and casual | pick one tone; remove the conflicting words | | Answers are generic | make CONTEXT and AUDIENCE specific | ## Next - [Create your own avatar](https://docs.bithuman.ai/build/create-avatar) · [Knowledge](https://docs.bithuman.ai/api/knowledge) · [Voices](https://docs.bithuman.ai/build/voices) --- # Voices URL: https://docs.bithuman.ai/build/voices > Every agent speaks every language on bitHuman's built-in voice pipeline, included in the voice chat rate — or bring your own OpenAI, Grok, ElevenLabs, or Cartesia key for premium voices. ## Two ways to give your agent a voice | Option | What you get | Cost | |------|--------------|------| | **bitHuman default** | The built-in voice pipeline. Your agent detects the caller's language and replies in it. No setup, no keys. | Included in the voice chat rate ([pricing](https://docs.bithuman.ai/pricing)) | | **Bring your own provider** | Connect your **own** OpenAI, Grok (xAI), ElevenLabs, or Cartesia key and pick that provider's premium voices — including low-latency speech-to-speech (realtime). | The same voice chat rate, plus your provider's charges on your key | You never *have* to bring a key. The default pipeline already speaks every language. Bring your own only when you want a specific premium voice or a provider's realtime engine. ## Use the default (nothing to do) Open any agent's voice settings at [bithuman.ai](https://www.bithuman.ai/explore). The **bitHuman voice** section is marked *Included* — design a voice, clone one, or pick from the gallery. It's multilingual automatically, so there's no language toggle to manage. ## Bring your own voice provider ### 1. Connect your key Go to **Developer → Integrations**, add your provider, and paste that provider's key; bitHuman checks the key there before saving and rejects an invalid one. The [Providers API](https://docs.bithuman.ai/api/providers) stores a key as sent, without a check, so start one session to confirm it works. Keys are encrypted at rest and never leave the platform in plaintext. ### 2. Pick a premium voice Back in the agent's voice settings, the premium providers you've connected unlock. Choose a voice; a provider you haven't connected stays locked with a shortcut to Integrations. Your selection runs that voice on your key. ## Supported providers | Provider | Voices you can select | Realtime (speech-to-speech) | |----------|-----------------------|:---------------------------:| | **OpenAI** | Realtime voices (alloy, ash, ballad, cedar, …) | ✓ | | **Grok (xAI)** | Grok voices (ara, eve, rex, …) | ✓ | | **ElevenLabs** | Your ElevenLabs voice library | — | | **Cartesia** | Cartesia voices | — | > **Tip** Realtime providers (OpenAI, Grok) give the lowest-latency, most expressive speech-to-speech — great for kiosks and live demos. ElevenLabs and Cartesia give you a specific voice on the standard pipeline. ## How billing works - Every managed-agent conversation bills the voice chat rate, whichever voice it uses ([pricing](https://docs.bithuman.ai/pricing)). - A bring-your-own voice or realtime model is also billed by your provider on your key. If a bring-your-own key ever fails or is removed, the agent automatically falls back to the built-in multilingual pipeline — it never silently stops talking. --- # Gestures (Essence 1) URL: https://docs.bithuman.ai/build/gestures > Play a named gesture (wave, nod, clap) on an avatar exactly when your code asks, on a cloud avatar or a self-hosted one. Gestures are named clips baked into an avatar, such as `mini_wave_hello` or `clap_cheer`. Your code plays one by name, when it chooses: on an app event, a timer, or an allow-listed tool call. Nothing plays at random. Gestures are an Essence 1 feature; Essence 2 and Expression 2 avatars have no gesture clips. ## Before you start - An Essence 1 agent of yours with gestures generated ([Gestures API](https://docs.bithuman.ai/api/dynamics)). - A LiveKit agent worker with the bitHuman plugin ([LiveKit](https://docs.bithuman.ai/platforms/livekit)). ## 1. List the gesture names ```bash curl -s https://api.bithuman.ai/v1/dynamics/$AGENT_CODE -H "api-secret: $BITHUMAN_API_SECRET" # → {"success": true, "data": {"status": "ready", "gestures": {"mini_wave_hello": "…", "clap_cheer": "…"}}} ``` The keys of `gestures` are the names you play, for example `mini_wave_hello` or `clap_cheer`. ## 2. Play one Cloud avatar (`AvatarSession(avatar_id=…)`): send the avatar participant a `trigger_dynamics` call. ```python import json resp = await ctx.room.local_participant.perform_rpc( destination_identity=avatar.avatar_identity, method="trigger_dynamics", payload=json.dumps({"action": "mini_wave_hello"}), ) # → {"action": "mini_wave_hello", "animation_triggered": true, "status": "success"} ``` Self-hosted avatar (`AvatarSession(model_path="avatar.imx")`): push a `VideoControl` to the runtime. ```python from bithuman import VideoControl await avatar.runtime.push(VideoControl(action="mini_wave_hello")) ``` ## 3. Wire it to your events As an allow-listed tool, so the language model can ask for a gesture only from your set. This tool is for a self-hosted avatar (`AvatarSession(model_path=…)`). For a cloud avatar, put the `perform_rpc` call from step 2 in the tool body instead. ```python from livekit.agents import function_tool, RunContext ALLOWED = {"mini_wave_hello", "clap_cheer", "thumbs_up_pulse"} @function_tool() async def play_gesture(context: RunContext, gesture: str) -> str: if gesture not in ALLOWED: return f"Unknown gesture. Options: {sorted(ALLOWED)}" await avatar.runtime.push(VideoControl(action=gesture, force_action=True)) return f"Played {gesture}" ``` Other triggers work the same way: a LiveKit data message, a participant event or a timer calls the same line. ## Check it worked The cloud call returns `animation_triggered: true`; the avatar plays the clip, then returns to idle or speech. ## Variations | `VideoControl` field | Effect | |---|---| | `force_action=True` | play even if another gesture is running | | `stop_on_user_speech=True` | stop the gesture when the user starts talking | | `stop_on_agent_speech=True` | stop the gesture when the agent starts talking | | `target_video=""` | switch the idle loop to another base clip | `avatar.runtime.interrupt()` stops a gesture at once. ## Troubleshooting | Symptom | Cause | Fix | |---|---|---| | `animation_triggered: false` | the name is not one of the avatar's gestures, or it has none | list the names (step 1); generate gestures with the [Gestures API](https://docs.bithuman.ai/api/dynamics) | | Nothing plays on a self-hosted avatar | the `.imx` was downloaded before gestures were generated | download the model again after generation | ## Next - [Gestures API](https://docs.bithuman.ai/api/dynamics) · [LiveKit](https://docs.bithuman.ai/platforms/livekit) · [Python](https://docs.bithuman.ai/platforms/python) --- # Voice agent URL: https://docs.bithuman.ai/build/voice-agent > A voice avatar on your own computer: LiveKit runs locally, OpenAI Realtime listens and speaks on your own key, and the bitHuman avatar renders on your CPU. One CLI command, or a short Python agent. Everything except the voice model runs on your computer. LiveKit is the stock `livekit-server`, OpenAI Realtime listens, thinks and speaks on your own `OPENAI_API_KEY`, and the bitHuman avatar renders on your CPU — no GPU needed. The avatar name picks the model: `wise-pup` is Expression 2, `sofia-ramirez` is Essence 2, or pass your own agent code or avatar file. You need two secrets: your bitHuman API secret (both models refuse to render without it) and `OPENAI_API_KEY`. *Diagram: A voice agent on your machine.* Your browser, a local livekit-server and the agent all run on your machine; the agent renders the avatar from the reply's voice, and the API secret stays in its process. The agent sends your speech to OpenAI Realtime with your key, and sends bitHuman a credential check and usage reports with no audio, video, images or conversation text. Use [the CLI](#with-the-cli) for a talking avatar with no code, [Python](#with-python) for your own agent code, or the [Python voice conversation](#python-voice-conversation) for a desktop window with no LiveKit server. ## With the CLI The CLI starts `livekit-server`, the voice agent and the avatar, then opens its page in your browser. ```bash # macOS (Apple silicon): installs livekit-server too brew install bithuman-product/bithuman/bithuman-cli # Linux x86_64 or arm64: the download includes livekit-server sudo apt install -y ffmpeg python3-venv # Ubuntu/Debian; macOS gets both from brew curl -fsSL https://install.bithuman.ai | sh bithuman login # opens your browser; in CI, export BITHUMAN_API_SECRET instead export OPENAI_API_KEY="" bithuman run wise-pup # Expression 2; `bithuman run sofia-ramirez` for Essence 2 ``` Expected output: ```text http://127.0.0.1:8088/WISEPUP Opening your browser… (Ctrl-C to stop.) ``` The first run downloads the avatar and sets up the voice agent, which takes a minute or two and needs Python 3.11 or newer. Allow the microphone and say "hi": the avatar answers, lip-synced, and stops when you talk over it. The voice settings are on [CLI](https://docs.bithuman.ai/platforms/cli#voice-settings). ## With Python A plain [LiveKit Agents](https://docs.livekit.io/agents/) program: you run `livekit-server` (1.9.12 or newer; check with `livekit-server --version`), and the avatar renders inside `agent.py`. Use Python 3.10–3.14. The CLI checks its `livekit-server` version for you. ```bash # 1. LiveKit server brew install livekit python@3.13 # macOS curl -sSL https://get.livekit.io | bash # Linux (Ubuntu 24.04 also: sudo apt install python3.12-venv) # 2. The example git clone https://github.com/bithuman-product/bithuman-examples cd bithuman-examples/python/self-host python3.13 -m venv .venv && . .venv/bin/activate # Ubuntu 24.04: python3.12 pip install -r requirements.txt # 3. Keys: in .env, never on the command line cp .env.example .env # fill BITHUMAN_MASTER_SECRET and OPENAI_API_KEY # 4. Run, in two terminals livekit-server --dev # terminal 1 python agent.py dev # terminal 2 ``` Expected output in terminal 2: ```text Open in Chrome: http://localhost:8089/?liveKitUrl=ws%3A%2F%2Flocalhost%3A7880&token=… ``` Open that link, click **Start** and allow the microphone. The heart of `agent.py`: ```python # excerpt: python/self-host/agent.py @server.rtc_session() async def entrypoint(ctx: JobContext): await ctx.connect(auto_subscribe=AutoSubscribe.AUDIO_ONLY) # the agent listens; it never needs your camera session = AgentSession(llm=openai.realtime.RealtimeModel( model=os.getenv("BITHUMAN_REALTIME_MODEL", "gpt-realtime-2.1-mini"), voice=os.getenv("BITHUMAN_VOICE", "coral"), # reply 0.5 s after you stop (the plugin's default semantic VAD can wait ~4 s) turn_detection=ServerVad(type="server_vad", silence_duration_ms=500, create_response=True, interrupt_response=True))) # Local mode: the avatar renders in this process and publishes the lip-synced video AND audio. avatar = bithuman.AvatarSession(model_path=os.environ["BITHUMAN_MODEL_PATH"], api_secret=os.environ["BITHUMAN_MASTER_SECRET"]) await avatar.start(session, room=ctx.room) await session.start( agent=Agent(instructions=os.getenv("BITHUMAN_INSTRUCTIONS", "You are a friendly assistant. Keep answers short.")), room=ctx.room, room_options=RoomOptions(audio_output=False, close_on_disconnect=False)) ``` Swap the `RealtimeModel` for any LiveKit speech-to-text, LLM and text-to-speech plugins; the avatar lines stay the same. ## Python voice conversation The same conversation in a desktop window, with no LiveKit server and no browser: your microphone goes to OpenAI Realtime, and the reply's voice drives a bitHuman avatar rendered on your machine. ### Requirements - A bitHuman API secret ([Developer → API Secrets](https://www.bithuman.ai/developer/api-keys)) and an `OPENAI_API_KEY`. - Python 3.10–3.14 in a virtualenv, a microphone and speakers. - On Linux, the PortAudio library: `sudo apt install libportaudio2`. ### Get the code ```bash git clone https://github.com/bithuman-product/bithuman-examples.git cd bithuman-examples/python/quickstart python3 -m venv .venv && . .venv/bin/activate pip install -r requirements.txt bithuman pull sofia-ramirez 2>/dev/null || curl -fsSL --create-dirs -o ~/.cache/bithuman/showcase/sofia-ramirez.imx \ "https://api.bithuman.ai/v1/agent/A52DHS2219/model/download?model=essence-2" ``` `sofia-ramirez` is a sample avatar (Essence 2, about 148 MB); downloading it needs no credential. ```bash export BITHUMAN_API_SECRET="" OPENAI_API_KEY="" ``` ### Run it ```bash python conversation.py --model ~/.cache/bithuman/showcase/sofia-ramirez.imx ``` Speak into your microphone; press `Q` in the window to quit. A window opens with the avatar at rest. When you stop speaking, the avatar answers, lip-synced, and you hear the reply through your speakers. ### How it works The pipeline: microphone → OpenAI Realtime (24 kHz PCM16) → `push_audio`/`flush` into the runtime → lip-synced frames and audio out. The heart of `conversation.py`: ```python # excerpt: python/quickstart/conversation.py async with client.realtime.connect(model="gpt-realtime-2.1-mini") as conn: await conn.session.update(session={ "type": "realtime", "instructions": "You are a friendly AI assistant. Keep responses concise.", "output_modalities": ["audio"], "audio": { "input": {"format": {"type": "audio/pcm", "rate": OPENAI_SAMPLE_RATE}, "turn_detection": {"type": "server_vad"}}, "output": {"format": {"type": "audio/pcm", "rate": OPENAI_SAMPLE_RATE}, "voice": args.voice}, }, }) # … async for event in conn: if event.type == "response.output_audio.delta": await ai_audio_queue.put(base64.b64decode(event.delta)) elif event.type == "response.output_audio.done": await ai_audio_queue.put(None) # … async def push_to_bithuman(): while True: data = await ai_audio_queue.get() if data is None: await runtime.flush() else: await runtime.push_audio(data, OPENAI_SAMPLE_RATE, last_chunk=False) # … async for frame in runtime.run(): if frame.has_image: cv2.imshow(WINDOW, frame.bgr_image) if cv2.waitKey(1) & 0xFF == ord("q"): break if frame.audio_chunk: with speaker_lock: speaker_buf.extend(frame.audio_chunk.array.tobytes()) ``` - **Personality:** edit the `instructions` string. - **Voice:** pass `--voice` with any OpenAI Realtime voice. - **Another avatar:** `bithuman list` prints every sample slug; pass the file with `--model`. ## Barge-in Talk over the avatar and it stops mid-sentence, then listens: that is barge-in. - **The CLI and the Python agent:** the voice model's turn detection hears you start talking and cancels its reply; LiveKit Agents then clears the avatar's buffered audio and frames, so the mouth stops with the voice. In the Python agent it is `interrupt_response=True` on the turn detection, shown above. - **Your own loop in Python:** when your speech detection fires, call `interrupt()` on the runtime and clear any reply audio you still hold. `run()` carries on with idle frames. With OpenAI Realtime, the event to watch is `input_audio_buffer.speech_started`: ```python # excerpt: barge-in, in the event loop of conversation.py elif event.type == "input_audio_buffer.speech_started": # the user started talking runtime.interrupt() # drop the rest of the reply ``` - **On the device:** `interrupt()` in Swift and Flutter; `resetState(true)` for Expression 2 and `resetAudio()` for Essence 2 on Android ([Companion app](https://docs.bithuman.ai/build/companion-app#let-the-user-interrupt)). - **A cloud avatar without the plugin:** perform the RPC `lk.clear_buffer` on the avatar ([Cloud avatar](https://docs.bithuman.ai/api/cloud-avatar)). Barge-in does not change billing: a session bills active session time, talking or idle, to the second. ## Pick the avatar `bithuman run `, or `BITHUMAN_AVATAR=` in the Python example's `.env`: | Value | You get | |---|---| | `wise-pup` | Expression 2, a sample avatar | | `sofia-ramirez` | Essence 2, a sample avatar | | your agent code, e.g. `A24EKJ8433` | your own avatar, downloaded with your API secret | | a path to an avatar file | a file you already have | ## Configuration The Python example reads these from `.env`: | Variable | Default | What it does | |---|---|---| | `BITHUMAN_MASTER_SECRET` | — | Your API secret, passed to the plugin explicitly. Rendering is metered on it. In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET`, never `BITHUMAN_API_SECRET`; `agent.py` refuses to start while `BITHUMAN_API_SECRET` is set. | | `OPENAI_API_KEY` | — | Your OpenAI key. | | `BITHUMAN_AVATAR` | `wise-pup` | Which avatar ([table above](#pick-the-avatar)). | | `BITHUMAN_REALTIME_MODEL` | `gpt-realtime-2.1-mini` | The OpenAI Realtime model. | | `BITHUMAN_VOICE` | `coral` | Any OpenAI Realtime voice. | | `BITHUMAN_INSTRUCTIONS` | a short assistant prompt | The agent's system prompt. | | `LIVEKIT_URL`, `LIVEKIT_API_KEY`, `LIVEKIT_API_SECRET` | `ws://localhost:7880`, `devkey`, `secret` | Your LiveKit server (`livekit-server --dev`). | ## What you pay Session time, talking or idle, bills bitHuman credits; OpenAI bills your own key. See [Pricing](https://docs.bithuman.ai/pricing). ## Troubleshooting | Symptom | Cause | Fix | |---|---|---| | `bithuman run` exits 69: `livekit-server 1.8.0 at …/livekit-server is too old for `bithuman run`` | The CLI needs livekit-server 1.13 or newer | `brew upgrade livekit` (macOS), or reinstall the CLI (Linux: its download includes one) | | The video stalls for 1–2 s every 15 s, or a LiveKit Meet tile goes black | `livekit-server` older than 1.9.12: the browser leaves and rejoins the room every 15 s | `brew upgrade livekit` (macOS) or `curl -sSL https://get.livekit.io \| bash` (Linux), then restart `livekit-server` | | `livekit-server not found` (exit 69) | LiveKit is not installed | `brew install livekit` (macOS) or `curl -sSL https://get.livekit.io \| bash` (Linux) | | The avatar never appears | No or invalid API secret | CLI: `bithuman login`. Python example: set `BITHUMAN_MASTER_SECRET` in `.env` | | The avatar never appears; the terminal shows `essence-2: ffmpeg not found` | Essence 2 unpacks its avatar with `ffmpeg` | `sudo apt install -y ffmpeg`, then run again | | The page says it could not connect | `livekit-server --dev` is not running | Start it, then click **Start** again | | Nothing happens after joining | `livekit-server --dev` or `agent.py` is not running | Start both, `livekit-server` first | | Another device on your network cannot join | `--dev` listens on `localhost` only | `livekit-server --dev --bind 0.0.0.0 --node-ip `; other browsers also need HTTPS for the microphone | | On a Mac, your own page with no microphone, on the same machine as the avatar, fails with `could not establish pc connection` | Chrome hides the machine's local addresses until the page has microphone permission | call `navigator.mediaDevices.getUserMedia({ audio: true })` before connecting, or open the page from another device | | `OSError: PortAudio library not found` | the system library is missing | `sudo apt install libportaudio2` (Debian, Ubuntu) | | `cv2.error: … The function is not implemented` | the headless OpenCV build won the install | `pip install --force-reinstall --no-deps opencv-python` | | `error: externally-managed-environment` | outside the virtualenv | `. .venv/bin/activate` | | No microphone input on macOS | the terminal has no microphone permission | System Settings → Privacy & Security → Microphone | --- # Website widget URL: https://docs.bithuman.ai/build/website-widget > Add a floating, talking avatar to any website with one script tag: the floating widget or the chat widget, every option, React and Next.js, and your own persona and model with no server. ## What you'll build A talking avatar in the corner of your website. Visitors open it, allow the microphone and talk to it; it answers out loud with its lips in sync. One script tag adds it to any page, in any framework. There is no npm package and no server to run. You need: - a website you can add a ` ``` Expected: A button in the bottom-right corner. Select it: the avatar opens in a floating frame you can drag and resize, asks for the microphone, and answers when you speak. ### Or add the chat widget The chat widget opens a side panel where visitors type, talk, or switch to video: ```html ``` Expected: A floating card in the corner. Opening it shows the panel with the welcome message; the mode switch moves between text, voice and video. ### Load it in React or Next.js In Next.js, load the script with `next/script` and call `init` when it has loaded: ```tsx "use client"; import Script from "next/script"; declare global { interface Window { BitHumanGadget?: { init: (options: Record) => void } } } export function AvatarWidget({ code }: { code: string }) { return (