# bitHuman — REST API (agents, realtime, video, voice, embedding, errors, rate limits) > Realtime talking avatars from one portrait. Index: https://docs.bithuman.ai/llms.txt · every page is also served as markdown at .md · OpenAPI: https://docs.bithuman.ai/api/openapi.yaml · other sections: https://docs.bithuman.ai/llms/start.txt · https://docs.bithuman.ai/llms/platforms.txt · https://docs.bithuman.ai/llms/deploy.txt · https://docs.bithuman.ai/llms/models.txt · https://docs.bithuman.ai/llms/build.txt ## Instructions for AI agents - Always send `model` ("essence-2" or "expression-2") when creating an agent; poll until `status` is `ready` or `failed`. - API and SDK use requires the Creator plan or higher. Never tell a user they can build on a free plan. - In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET` and pass a minted token (https://docs.bithuman.ai/platforms/livekit.md). - Say where things happen: the avatar renders (device, browser, server, cloud); the conversation runs (your stack, the CLI's local conversation brain, or bitHuman's servers). - Quote speed only from https://docs.bithuman.ai/performance.json (× real time, with the device). Quote prices only from GET https://api.bithuman.ai/v1/pricing. - Do not claim offline on phones, Mac or the browser, a conversation brain on phones, or any certification. Where each model runs, the key facts and one command per path: https://docs.bithuman.ai/llms.txt ## Contents - API overview — https://docs.bithuman.ai/api - Authentication — https://docs.bithuman.ai/api/authentication - Agents — https://docs.bithuman.ai/api/agents - Text to speech — https://docs.bithuman.ai/api/text-to-speech - Talking video API — https://docs.bithuman.ai/api/video - Embedding — https://docs.bithuman.ai/api/embedding - Cloud avatar in your room — https://docs.bithuman.ai/api/cloud-avatar - Realtime relay — https://docs.bithuman.ai/api/realtime - Errors — https://docs.bithuman.ai/api/errors - Rate limits — https://docs.bithuman.ai/api/rate-limits Linked, not inlined (read the .md twin): - Knowledge — https://docs.bithuman.ai/api/knowledge.md - Files — https://docs.bithuman.ai/api/files.md - Gestures API — https://docs.bithuman.ai/api/dynamics.md - Providers (BYOK) — https://docs.bithuman.ai/api/providers.md - Webhooks — https://docs.bithuman.ai/api/webhooks.md - Runtime sessions — https://docs.bithuman.ai/api/runtime-sessions.md - API secrets API — https://docs.bithuman.ai/api/api-keys.md - Organizations — https://docs.bithuman.ai/api/organizations.md - Billing & usage — https://docs.bithuman.ai/api/billing.md - Examples: https://docs.bithuman.ai/examples.md · changelog: https://docs.bithuman.ai/changelog.md · API references: https://docs.bithuman.ai/platforms/cli/reference.md, https://docs.bithuman.ai/platforms/python/reference.md, https://docs.bithuman.ai/platforms/swift/reference.md, https://docs.bithuman.ai/platforms/android/reference.md --- # API overview URL: https://docs.bithuman.ai/api > REST API for generating avatars, synthesizing voice, driving live sessions, and embedding agents — from any language. ## What the API does The bitHuman API lets you create, manage, and drive avatar agents from any programming language. Reach for it when you don't need a native SDK — backends, CI scripts, or platforms where Python or Swift aren't a fit. Everything is plain HTTPS + JSON. One header authenticates every request. ## Base URL ```text https://api.bithuman.ai ``` All endpoints are relative to this URL and require an `api-secret` header. Every endpoint, with a live console: [API reference](https://docs.bithuman.ai/api/reference) (raw spec: https://docs.bithuman.ai/api/openapi.yaml). [Get an API secret →](https://www.bithuman.ai/developer/api-keys) ## Authentication Pass your API secret in the `api-secret` header on every request — the cheapest check is `/v1/validate`, which costs nothing: ```bash curl -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET" # {"valid":true} ``` Treat the secret like a password — never commit it to source control and never embed it in client apps. For browser-side embeds, mint a short-lived token with the [embed token flow](https://docs.bithuman.ai/api/embedding) instead. See [Authentication](https://docs.bithuman.ai/api/authentication) for the full model. ## What you can build - **Generate avatars** — turn a prompt, portrait, and voice sample into a new agent. See [Agents](https://docs.bithuman.ai/api/agents). - **Synthesize voice** — text-to-speech in 30+ languages with 10 built-in voices, plus an OpenAI-compatible drop-in. See [Text to Speech](https://docs.bithuman.ai/api/text-to-speech). - **Drive live sessions** — make a hosted agent speak or inject silent knowledge into an active room. See [Agents](https://docs.bithuman.ai/api/agents). - **Add gestures** — generate and toggle conversational animations. See [Gestures](https://docs.bithuman.ai/api/dynamics). - **Ground agents in your docs** — ingest files and URLs into knowledge bases from code. See [Knowledge](https://docs.bithuman.ai/api/knowledge). - **Add realtime voice** — mint a browser client secret for OpenAI-Realtime sessions. See [Realtime](https://docs.bithuman.ai/api/realtime). - **Bring your own keys** — use your own LLM/STT/TTS provider keys. See [Providers](https://docs.bithuman.ai/api/providers). - **Render talking videos** — generate a finished mp4 of an agent speaking, from a text script or hosted audio. See [Video API](https://docs.bithuman.ai/api/video). - **Embed in any page** — mint a token and drop an iframe. See [Embedding](https://docs.bithuman.ai/api/embedding). - **Track credits** — read balance and per-mode minute estimates. See [Billing](https://docs.bithuman.ai/api/billing). - **Manage keys & teams** — rotate [API secrets](https://docs.bithuman.ai/api/api-keys), watch [runtime sessions](https://docs.bithuman.ai/api/runtime-sessions), and run [organizations](https://docs.bithuman.ai/api/organizations) programmatically. - **Get notified** — register [webhooks](https://docs.bithuman.ai/api/webhooks) for signed `agent.ready` / `agent.failed` and `video.completed` / `video.failed` events instead of polling. - **Drive it from an AI agent** — the [MCP server](https://docs.bithuman.ai/build/mcp) (`bithuman mcp`) exposes the common endpoints (agents, speech, gestures, files, embed tokens, webhooks, balance) as tools for Claude, Cursor and other MCP clients. ## How agents are identified Every endpoint identifies an agent by its **agent code** — a short string like `A80HVD8577`. You receive one when you [generate an agent](https://docs.bithuman.ai/api/agents), or find it in your [Library](https://www.bithuman.ai/#library) (click an agent to reveal the code). > **Note** Different endpoint paths use slightly different parameter names for > the same value: `{code}`, `{agent_code}`, or `{agent_id}`. They all expect the > same string — the agent code shown in your Library. ## Next steps - [Quickstart](https://docs.bithuman.ai/platforms/rest) — make your first API call and drive a live agent. - [Authentication](https://docs.bithuman.ai/api/authentication) — get an API secret and runtime tokens. - [Models](https://docs.bithuman.ai/models) — the four models, where each runs, and which to pick. - [API reference](https://docs.bithuman.ai/api/reference) — the interactive reference for the core endpoints (the OpenAPI file is at /api/openapi.yaml). - [Errors](https://docs.bithuman.ai/api/errors) and [Rate limits](https://docs.bithuman.ai/api/rate-limits) — the operational contract. - [MCP server](https://docs.bithuman.ai/build/mcp) — call the common endpoints as tools from an AI agent. ## Status and versioning Endpoints are stable and change additively. The `/v1` and `/v2` prefixes are part of each endpoint's path, not a version switch; use each path exactly as documented. Live API status is at [status.bithuman.ai](https://status.bithuman.ai). --- # Authentication URL: https://docs.bithuman.ai/api/authentication > Authenticate REST calls with the api-secret header, check a secret with /v1/validate, and use short-lived tokens where a secret must not go. Every REST call carries your API secret in the `api-secret` header. The same secret works for the SDKs and the CLI ([Your API secret](https://docs.bithuman.ai/start/api-secret)). Create one under [Developer → API Secrets](https://www.bithuman.ai/developer/api-keys); the value is shown once. | Method | Path | Purpose | Credits | |---|---|---|---| | `POST` | `/v1/validate` | Check an API secret | free | | `POST` | `/v1/runtime-tokens/request` | Exchange the secret for a short-lived runtime token (the SDKs do this for you) | free | | `POST` | `/v1/runtime-tokens/mint` | A one-hour token scoped to one agent: for the LiveKit plugin (`"scope": "livekit-cloud"`, bound to a room when you send `room_name`/`livekit_url`) or, without `scope`, a Bearer credential for downloading that agent's model on a device | free | | `POST` | `/v1/embed-tokens/request` | A one-hour token for a browser embed ([Embedding](https://docs.bithuman.ai/api/embedding)) | free | ## POST /v1/validate Checks the secret in the `api-secret` header. Always returns `200`; read `valid`. ### Example ```bash tab="curl" curl -s -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET" ``` ```python tab="Python" import os, requests r = requests.post("https://api.bithuman.ai/v1/validate", headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]}) print(r.json()) ``` ### Response ```json {"valid": true} ``` ## POST /v1/runtime-tokens/request Exchanges the API secret for a short-lived runtime token that authorizes rendering for your account. The Python SDK and the LiveKit plugin call it for you and renew the token while a session runs; call it yourself only when you build your own runtime integration. A runtime token cannot create other tokens or call other endpoints. Sent with `mode`, the same endpoint starts a cloud avatar in your LiveKit room instead: [Cloud avatar without the plugin](https://docs.bithuman.ai/api/cloud-avatar). ## POST /v1/runtime-tokens/mint Mints a one-hour token for `"scope": "livekit-cloud"` that can only start one agent's avatar in one LiveKit room. Pass it to the LiveKit plugin instead of your secret, because the plugin writes its credential into room attributes every participant can read. Send `room_name` and `livekit_url` to bind the LiveKit token to one room and server. The request and a complete worker are on [LiveKit](https://docs.bithuman.ai/platforms/livekit#authenticate). ## Keep the secret safe - Keep it in the environment or a secrets manager, never in source, a command line or an app bundle. - Browsers get an [embed token](https://docs.bithuman.ai/api/embedding); LiveKit rooms get a minted token; shipped apps fetch a credential from your backend. - `BITHUMAN_API_KEY` is a deprecated alias of `BITHUMAN_API_SECRET`. ## Rotate a secret Create the new secret, move your services to it, then revoke the old one under [API Secrets](https://www.bithuman.ai/developer/api-keys). Revocation is immediate: sessions still using the old secret stop at their next usage report. The CLI's per-device secrets (`cli@`) revoke one machine at a time. The API equivalents are on [API secrets](https://docs.bithuman.ai/api/api-keys). ## Errors | Status | Code | Cause | Fix | |---|---|---|---| | `401` | `MISSING_AUTH` | no `api-secret` header | send the header on every request | | `401` | `UNAUTHORIZED` | the secret is invalid, or a revoked secret on a REST endpoint | check it with `/v1/validate`; create a new one | | `403` | `RUNTIME_SUSPENDED` | a revoked secret on a token endpoint (`/v1/runtime-tokens/*`, `/v1/embed-tokens/request`) | create a new secret and move your services to it | | `403` | `RUNTIME_SUSPENDED` | runtime access is suspended for the account; the message says so | contact support | | `403` | `SECRET_REVEAL_CONSOLE_ONLY` | reading a stored secret's value with an API secret | reveal secrets in the console | All codes: [Errors](https://docs.bithuman.ai/api/errors). --- # Agents URL: https://docs.bithuman.ai/api/agents > Create an avatar agent from a portrait, poll until it is ready, then manage it, download its model, and make it speak in live sessions. An agent is an avatar (face, voice and persona) identified by a short code such as `A23WJF0199`. Create one, poll until it is `ready`, then use it everywhere: the [web embed](https://docs.bithuman.ai/platforms/web), the SDKs, [talking video](https://docs.bithuman.ai/api/video) and live sessions. Creation and model adds cost credits per model ([pricing](https://docs.bithuman.ai/pricing)); everything else on this page is free. | Method | Path | Purpose | |---|---|---| | `POST` | `/v1/agent/generate` | [Create an agent](#generate-an-agent) | | `GET` | `/v1/agent/status/{agent_id}` | [Poll creation](#poll-status) | | `GET` | `/v1/agent/{code}` | [Get an agent](#get-an-agent) | | `GET` | `/v1/agents` | [List your agents](#list-your-agents) | | `POST` | `/v1/agent/{code}` | [Update the prompt or providers](#update-an-agent) | | `DELETE` | `/v1/agent/{code}` | [Delete an agent](#delete-an-agent) | | `POST` | `/v1/agent/{code}/models` | [Add a model](#add-a-model-to-an-existing-agent) | | `GET` | `/v1/agent/{code}/model/download` | [Download the model file](#download-an-agents-model) | | `GET` | `/v1/agent/{code}/sessions` | [List live sessions](#list-an-agents-live-sessions) | | `POST` | `/v1/agent/{code}/speak` | [Speak in a live session](#make-an-agent-speak) | | `POST` | `/v1/agent/{code}/add-context` | [Add knowledge to a live session](#inject-knowledge) | ## Generate an agent `POST /v1/agent/generate` starts an asynchronous creation and returns an `agent_id` at once. Credits are reserved at submit and refunded automatically if creation fails. ### Request | Parameter | Type | Required | Description | |---|---|---|---| | `model` | string | always send it | `essence-2` (a photoreal person), `expression-2` (any character), `auto` (the platform picks from the image), `essence-1` or `expression-1`. Omitted, the API still creates an `expression-1` agent and warns (`MODEL_DEFAULT_DEPRECATED`); **from 2026-12-26 `model` is required**, and the bare names `essence` and `expression` (and the `version` field) are refused with a `400` naming the replacement. | | `image` | string | no | Portrait URL (publicly fetchable) or base64. Used as a reference; a portrait is generated from `prompt` when omitted | | `prompt` | string | no | System prompt and personality | | `audio` | string | no | Voice sample URL or base64, for voice cloning | | `aspect_ratio` | string | no | `16:9` (default), `9:16` or `1:1` | | `framing` | string | no | `portrait` (default) or `full_body` | | `transparency` | boolean | no | `true` generates on a green-screen background for chroma key | | `agent_id` | string | no | Leave it out; a fresh code is generated. If it names an agent you already own, that agent is regenerated in place (it returns to `processing`); another account's code returns `404` | Essence 2 Max is available on the Enterprise plan only; other plans get `403 PLAN_REQUIRED`. [Contact sales](https://www.bithuman.ai/sales) to enable it. Headers: `api-secret`, and optionally `Idempotency-Key`: a repeated request with the same key returns the first response and starts no second creation. ### Example ```bash tab="curl" curl -X POST https://api.bithuman.ai/v1/agent/generate \ -H "Content-Type: application/json" \ -H "api-secret: $BITHUMAN_API_SECRET" \ -H "Idempotency-Key: museum-guide-1" \ -d '{"model": "expression-2", "prompt": "You are a cheerful museum guide.", "image": "https://example.com/portrait.jpg"}' ``` ```python tab="Python" import os, requests resp = requests.post( "https://api.bithuman.ai/v1/agent/generate", headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]}, json={"model": "expression-2", "prompt": "You are a cheerful museum guide.", "image": "https://example.com/portrait.jpg"}, ) print(resp.json()) ``` ### Response ```json {"success": true, "message": "Agent generation started", "agent_id": "A80HVD8577", "status": "processing"} ``` ### Notes - Creation takes minutes for `essence-1` and `expression-1`, and about 2–2.5 hours for `essence-2` and `expression-2`. Set your polling timeout per model. - `essence-2` needs a photoreal person: another subject returns `422 MODEL_SUBJECT_MISMATCH` before anything is charged. An `essence-2` creation or add also checks the photo's face before charging: a face too small in frame, or several similar-sized faces, returns `422 IMAGE_FACE_UNSUITABLE` — upload a closer photo of one person. `auto` routes people to `essence-2` and everything else to `expression-2`. - A `200` does not mean the image was fetched. An unreachable `image` fails the creation a few seconds later (refunded); poll [status](#poll-status) to confirm. Creation is image-only: a `video` field returns `400 VIDEO_INPUT_NOT_SUPPORTED`. - At most two `essence-2` creations run at once per account; a third fails at once with a capacity message and no charge. ## Poll status `GET /v1/agent/status/{agent_id}` reports a creation's progress. Poll every 5 seconds until `status` is `ready` or `failed`; every other value is intermediate. ```bash curl https://api.bithuman.ai/v1/agent/status/A80HVD8577 -H "api-secret: $BITHUMAN_API_SECRET" ``` ```json {"success": true, "data": {"agent_id": "A80HVD8577", "status": "ready", "progress": 1.0, "current_step": "done", "error_message": null, "model_url": "https://…", "supported_models": ["expression-2"], "model_status": {"expression-2": {"state": "ready", "reason": null}}, "name": "Museum Guide"}} ``` | `current_step` | Progress | Stage | |---|---|---| | `payment` | about 2% | credits reserved | | `persona` | 5–15% | persona prepared | | `voice_image` | about 20% | voice and portrait | | `video` | about 45% | identity video (Essence models) | | `lip_sync` | 70–99% | the model step; the longest for second-generation models | | `done` | 100% | `ready` | ```python import os, time, requests def wait_until_ready(agent_id, timeout_s=3 * 3600): deadline = time.time() + timeout_s while time.time() < deadline: r = requests.get(f"https://api.bithuman.ai/v1/agent/status/{agent_id}", headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]}, timeout=30) if r.ok: data = r.json()["data"] if data["status"] == "ready": return data if data["status"] == "failed": raise RuntimeError(data["error_message"]) time.sleep(5) # creation continues server-side through transient errors raise TimeoutError(agent_id) ``` ### Notes - `supported_models` lists the models the agent can be launched as, spelled as `model` values you can send back. - `model_status` gives each requested model's state (`pending`, `ready`, `failed`). Models never requested are absent. - The downloadable model file is published shortly after `ready`; until then the [download](#download-an-agents-model) returns `404 MODEL_ARTIFACT_NOT_READY`. Retry. ## Get an agent `GET /v1/agent/{code}` returns the agent's full record: persona (`system_prompt`, `name`, `language`), `voice_id`, media URLs, creation state, `model` and `supported_models`. ```bash curl https://api.bithuman.ai/v1/agent/A80HVD8577 -H "api-secret: $BITHUMAN_API_SECRET" ``` ```json {"success": true, "data": {"code": "A80HVD8577", "status": "ready", "model": "expression-2", "supported_models": ["expression-2"], "name": "Museum Guide", "system_prompt": "You are a cheerful museum guide.", "image_url": "https://…/image.jpg"}} ``` ## List your agents `GET /v1/agents` lists your agents, newest first. Query: `limit` (default 20, max 100), `offset`, `status` (for example `ready`). Items are summaries (`code`, `name`, `model`, `status`, `supported_models`, `created_at` and a few more); read the full record with [Get an agent](#get-an-agent). Deleted agents appear with `status: "deleted"`. ```bash curl "https://api.bithuman.ai/v1/agents?status=ready&limit=20" -H "api-secret: $BITHUMAN_API_SECRET" ``` ```json {"success": true, "data": [{"code": "A80HVD8577", "name": "Museum Guide", "model": "expression-2", "status": "ready"}], "pagination": {"limit": 20, "offset": 0, "total": 1, "has_more": false}} ``` ## Update an agent `POST /v1/agent/{code}` changes the `system_prompt`, the voice-provider selection (`providers`, see [Voice providers](https://docs.bithuman.ai/api/providers)), or both. Send at least one, or the call returns `400 MISSING_PARAM`. The name is generated and cannot be set. ```bash curl -X POST https://api.bithuman.ai/v1/agent/A80HVD8577 \ -H "Content-Type: application/json" -H "api-secret: $BITHUMAN_API_SECRET" \ -d '{"system_prompt": "You are a concise sales assistant."}' ``` ```json {"agent_code": "A80HVD8577", "updated": true} ``` ## Delete an agent `DELETE /v1/agent/{code}` deletes an agent you own. Usage history is kept. An unknown or unowned code returns `404`. ```bash curl -X DELETE https://api.bithuman.ai/v1/agent/A80HVD8577 -H "api-secret: $BITHUMAN_API_SECRET" ``` ```json {"success": true, "agent_code": "A80HVD8577", "deleted": true} ``` ## Add a model to an existing agent `POST /v1/agent/{code}/models` with `{"model": ""}` adds a model to a `ready` agent without re-creating it. Re-adding a model the agent has costs nothing, and a failed add is refunded. | `model` | Needs | Time | Credits | |---|---|---|---| | `expression-1` | a stored image and voice | immediate | free | | `expression-2` | a stored image | about 2–2.5 h | 2000 | | `essence-2` | a stored identity video and a photoreal person | about 2–2.5 h | 500 | | `essence-1` | a stored image or identity video | 10–20 min | 250 | ```bash curl -X POST https://api.bithuman.ai/v1/agent/A80HVD8577/models \ -H "Content-Type: application/json" -H "api-secret: $BITHUMAN_API_SECRET" \ -d '{"model": "essence-2"}' ``` ```json {"success": true, "agent_id": "A80HVD8577", "model": "essence-2", "status": "processing", "supported_models": ["expression-2"]} ``` Poll [status](#poll-status) until `model_status["essence-2"].state` is `ready` or `failed`. The top-level `status` stays `ready` during an add. A failed add is refunded, and its `reason` says why. Errors: `409 AGENT_NOT_READY`, `422 MODEL_PREREQUISITE_MISSING`, `422 MODEL_SUBJECT_MISMATCH`. ## Download an agent's model `GET /v1/agent/{code}/model/download` redirects (`302`) to the agent's model file, a `.imx` container, for the [SDKs](https://docs.bithuman.ai/platforms) and [CLI](https://docs.bithuman.ai/platforms/cli). Pass `?model=` to choose a model when the agent has several; the default is the model it was created with. Sample avatars download with no credential; your own agents need the `api-secret` header. ```bash curl -fL -o A80HVD8577.imx -H "api-secret: $BITHUMAN_API_SECRET" \ "https://api.bithuman.ai/v1/agent/A80HVD8577/model/download?model=expression-2" ``` Name the output file yourself (`-o`). Add `?redirect=false` to get the URL as JSON: ```json {"success": true, "data": {"code": "A80HVD8577", "model": "expression-2", "filename": "A80HVD8577.imx", "url": "https://…", "expires_in": 3600}} ``` | Status | Code | Meaning | |---|---|---| | `400` | `MODEL_NOT_DOWNLOADABLE` | the model has no file (`expression-1`) | | `403` | `PLAN_REQUIRED` | the model is not in your plan; the message names the plan | | `404` | `MODEL_ARTIFACT_NOT_READY` | published shortly after `ready`; retry | | `409` | `MODEL_NOT_GENERATED` | the agent does not have that model; [add it](#add-a-model-to-an-existing-agent) | ## List an agent's live sessions `GET /v1/agent/{code}/sessions` lists the agent's open sessions. Use a `room_id` whose `deliverable` is `true` with [speak](#make-an-agent-speak) or [add-context](#inject-knowledge). ```bash curl https://api.bithuman.ai/v1/agent/A80HVD8577/sessions -H "api-secret: $BITHUMAN_API_SECRET" ``` ```json {"agent_code": "A80HVD8577", "sessions": [{"room_id": "room-A80HVD8577-x1y2", "num_participants": 2, "created_at": 1788480000, "deliverable": true}]} ``` ## Make an agent speak `POST /v1/agent/{code}/speak` makes the avatar say `message` in its live sessions: one session with `room_id`, or every deliverable session without it. With no live session it returns `404 NOT_FOUND`. ```bash curl -X POST https://api.bithuman.ai/v1/agent/A80HVD8577/speak \ -H "Content-Type: application/json" -H "api-secret: $BITHUMAN_API_SECRET" \ -d '{"message": "We have a 20% discount today.", "room_id": "room-A80HVD8577-x1y2"}' ``` ```json {"agent_code": "A80HVD8577", "delivered_to_rooms": 1, "rooms": ["room-A80HVD8577-x1y2"], "rooms_skipped": [], "rooms_failed": []} ``` ## Inject knowledge `POST /v1/agent/{code}/add-context` gives a live agent background knowledge (`"type": "add_context"`, the default) or a message to say (`"type": "speak"`). `room_id` targets one session. It needs a live session, like speak. ```bash curl -X POST https://api.bithuman.ai/v1/agent/A80HVD8577/add-context \ -H "Content-Type: application/json" -H "api-secret: $BITHUMAN_API_SECRET" \ -d '{"context": "The visitor is a member. Preferred name: Alex."}' ``` ## Errors | Status | Code | When | |---|---|---| | `400` | `MISSING_PARAM` | an update with nothing to change | | `400` | `MODEL_NOT_DOWNLOADABLE` | downloading a model that has no file | | `400` | `VALIDATION_ERROR` | invalid body or `model`; the message lists accepted values | | `400` | `VIDEO_INPUT_NOT_SUPPORTED` | `video` in a creation request | | `401` | `UNAUTHORIZED` | missing or invalid `api-secret` | | `402` | `INSUFFICIENT_BALANCE` | not enough credits to create | | `403` | `PLAN_REQUIRED` | the model is not in your plan; the message names the plan | | `404` | `MODEL_ARTIFACT_NOT_READY` | the model file is not published yet; retry | | `404` | `NOT_FOUND` | unknown agent, or no live session for speak or add-context | | `409` | `AGENT_NOT_READY` | adding a model to an agent that is not `ready` | | `409` | `MODEL_NOT_GENERATED` | the agent does not have the requested model | | `410` | `AGENT_DELETED` / `AGENT_PURGED` | the agent was deleted; model download, gestures and talking video refuse it | | `422` | `MODEL_PREREQUISITE_MISSING` | the agent lacks an asset the model needs | | `422` | `IMAGE_FACE_UNSUITABLE` | the face is too small or there are several faces | | `422` | `MODEL_SUBJECT_MISMATCH` | `essence-2` for a subject that is not a photoreal person | | `503` | `MODEL_NOT_YET_AVAILABLE` | a model is paused for your account (not returned in normal operation) | All codes: [Errors](https://docs.bithuman.ai/api/errors). --- # Text to speech URL: https://docs.bithuman.ai/api/text-to-speech > Turn text into natural speech with bitHuman's real-time TTS — built-in voices, inline tuning, and shareable voice codes designed in the playground. bitHuman's text-to-speech runs the same in-house voice engine that powers live agents. One `POST` turns text into a WAV you can save or stream on the fly. It supports 30+ languages, ten built-in voices, fine-grained tuning, and **voice codes** — opaque handles for a voice you've designed in the [Voice Designer](https://www.bithuman.ai/voice). ## Authentication Every call uses your bitHuman API secret in the `api-secret` header. Get one at [Developer → API Secrets](https://www.bithuman.ai/developer/api-keys) (the Creator plan or higher), then export it so the examples below pick it up: ```bash export BITHUMAN_API_SECRET="" ``` ## Synthesize speech `POST https://api.bithuman.ai/v1/tts` returns audio bytes (a WAV by default). **curl** ```bash curl -X POST https://api.bithuman.ai/v1/tts \ -H "api-secret: $BITHUMAN_API_SECRET" \ -H "content-type: application/json" \ -d '{"text": "Hello from bitHuman.", "voice": "F1", "language": "en"}' \ --output voice.wav ``` **Python** ```python import os, requests resp = requests.post( "https://api.bithuman.ai/v1/tts", headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]}, json={"text": "Hello from bitHuman.", "voice": "F1", "language": "en"}, timeout=60, ) resp.raise_for_status() with open("voice.wav", "wb") as f: f.write(resp.content) ``` ### Request fields | Field | Type | Notes | | --- | --- | --- | | `text` | string | **Required.** Any length; multi-sentence is supported. | | `voice` | string | Built-in voice id (`M1`–`M5`, `F1`–`F5`). Defaults to `M1`. | | `voice_code` | string | A designed-voice handle (see [Voice codes](#voice-codes)). Takes precedence over `voice`. | | `axes` | object | Inline tuning — see [Tuning a voice](#tuning-a-voice). Ignored when `voice_code` is set. | | `language` | string | ISO-2 code. Call `GET /v1/voices` for the current list of languages. Defaults to `en`. | | `total_steps` | integer | Quality vs. speed: `5` fast, `8` balanced (default), `12` highest. | | `speed` | number | Playback rate, `0.7`–`2.0`. Defaults to `1.05`. | ## List voices `GET /v1/voices` returns the catalog — ten built-ins (`M1`–`M5`, `F1`–`F5`) plus any custom voices. ```bash curl https://api.bithuman.ai/v1/voices -H "api-secret: $BITHUMAN_API_SECRET" # {"voices":[{"id":"F1","kind":"builtin"}, ... ]} ``` ## Tuning a voice Shape any built-in voice with semantic `axes` — `gender`, `pitch`, `rate`, and `brightness`. Offsets are small (roughly −0.3…0.3); `0` is neutral. Call `GET /v1/studio/axes` for each axis's suggested range and per-voice anchors. ```bash curl -X POST https://api.bithuman.ai/v1/tts \ -H "api-secret: $BITHUMAN_API_SECRET" \ -H "content-type: application/json" \ -d '{ "text": "Tuned, warm, and a touch brighter.", "voice": "F3", "axes": {"gender": 0.1, "pitch": 0.05, "rate": -0.1, "brightness": 0.2} }' \ --output voice.wav ``` ## Voice codes Rather than hand-tuning axes, design a voice from a description in the [Voice Designer](https://www.bithuman.ai/voice) ("a calm meditation guide", "a gruff old captain"). When you open **Use in your app**, you get a **voice code** — a single opaque handle that already encodes the base voice and its tuning. Pass it as `voice_code` and skip `voice`/`axes` entirely: ```bash curl -X POST https://api.bithuman.ai/v1/tts \ -H "api-secret: $BITHUMAN_API_SECRET" \ -H "content-type: application/json" \ -d '{"text": "Hello from my custom voice.", "voice_code": "YOUR_VOICE_CODE"}' \ --output voice.wav ``` A voice code is a UUID (e.g. `f8fb5feb-8a19-435c-89e5-a286a03565ec`). The endpoint expands it to the underlying voice + tuning, so your integration only ever references the code — re-tune the voice in the playground without touching your code path. An unknown or revoked `voice_code` returns `404 VOICE_NOT_FOUND` — handle it rather than assuming a fallback voice. ## Stream and play on the fly `/v1/tts` returns standard WAV bytes, so you can pipe the response straight into a player instead of saving a file — handy for quick local testing: ```bash curl -sN -X POST https://api.bithuman.ai/v1/tts \ -H "api-secret: $BITHUMAN_API_SECRET" \ -H "content-type: application/json" \ -d '{"text": "Playing right away.", "voice_code": "YOUR_VOICE_CODE"}' \ | ffplay -autoexit -nodisp -i - ``` For sentence-by-sentence streaming of length-prefixed PCM frames (lowest latency for long text), set `"stream": true`. ## OpenAI-compatible endpoint Already calling OpenAI's TTS? Point existing clients at `POST /v1/audio/speech` — swap the base URL to `https://api.bithuman.ai/v1` and the auth header to `api-secret`. See the [API reference](https://docs.bithuman.ai/api/reference#tag/voice) for the full schema. ## Errors `401` means a missing or invalid `api-secret`; `400` is a malformed body; `404` (`VOICE_NOT_FOUND`) means the `voice_code` doesn't resolve to a known voice; `503` means the queue is briefly full — retry with backoff. See [Errors](https://docs.bithuman.ai/api/errors). --- # Talking video API URL: https://docs.bithuman.ai/api/video > Render a talking-video mp4 over REST — submit a text script or hosted audio, poll the async job, and receive a CDN URL. Per-minute billing, auto-refunded on failure. ## Overview `POST /v1/video/generate` renders an MP4 of one of your agents speaking, from a **text** script (in the agent's voice) or a **hosted audio** file. Submit a job and poll for the URL, or pass [`wait: true`](#blocking-mode-wait-true) to get the MP4 in the response. `essence-2` renders at up to 1080p, `1080×1920` or `1920×1080` to match the source; `expression-2` renders at `416×720`. Renders bill **per minute of output, rounded up**: 4 credits/min for `essence-2`, `expression-2` and `expression-1`, 2 for `essence-1`. A job charges the 120-second maximum up front and refunds the difference when it finishes, so your balance must cover that maximum at submit time. A failed render is refunded in full. Limits: up to **120 seconds** of output and **5000 characters** of text. ## Generate a talking video **Before you start:** you need an agent you own. List yours with `curl https://api.bithuman.ai/v1/agents -H "api-secret: $BITHUMAN_API_SECRET"` and `export BITHUMAN_AGENT_CODE=A…`. Creating one needs the Creator plan or higher ([Pricing](https://docs.bithuman.ai/pricing#plans)). `POST /v1/video/generate` returns a `job_id` with `status: "processing"`; poll [`GET /v1/video/{job_id}`](#get-talking-video-status) until it completes, or register a [webhook](https://docs.bithuman.ai/api/webhooks) for `video.completed` / `video.failed`. | Parameter | Type | Required | Description | |---|---|---|---| | `model` | string | yes | Engine: `essence-1`, `expression-1`, `expression-2`, or `essence-2`. All four render talking video today. A model outside your plan returns `403 PLAN_REQUIRED`. | | `agent_code` | string | yes | An agent you own — supplies the avatar identity (and, for text, the default voice). | | `input` | object | yes | The render source — see below. | | `input.type` | string | yes | `text` or `audio`. | | `input.text` | string | for text | Script to speak (≤ 5000 chars). | | `input.voice` | string | no | Voice id override for text input. Defaults to the agent's own voice. | | `input.audio_url` | string | for audio | Public URL to a WAV or MP3 file. | | `wait` | boolean | no | Blocking mode. `false` (default) returns a `job_id` to poll. `true` blocks until the render finishes (up to ~90s) and returns the finished `video_url` — plus `duration_seconds` and `credits_charged` — directly in this response; if it exceeds the cap you get the async `{ job_id }` to poll instead. Accepted as a JSON/multipart field or as a `?wait=true` query parameter. | ### Text input ```bash curl -X POST https://api.bithuman.ai/v1/video/generate \ -H "api-secret: $BITHUMAN_API_SECRET" -H "Content-Type: application/json" \ -d '{"model": "essence-2", "agent_code": "'"$BITHUMAN_AGENT_CODE"'", "input": {"type": "text", "text": "Hello, welcome to bitHuman."}}' ``` ```json { "success": true, "job_id": "vid_3f9a2c1b8e7d4a6f0b21", "status": "processing" } ``` ### Audio input ```bash curl -X POST https://api.bithuman.ai/v1/video/generate \ -H "api-secret: $BITHUMAN_API_SECRET" -H "Content-Type: application/json" \ -d '{"model": "expression-2", "agent_code": "'"$BITHUMAN_AGENT_CODE"'", "input": {"type": "audio", "audio_url": "https://example.com/speech.wav"}}' ``` ### Blocking mode (`wait: true`) Add `"wait": true` (a JSON/multipart field, or `?wait=true` as a query parameter) to hold the connection until the render finishes and get the mp4 back in the same response — no polling. If the render exceeds the ~90-second cap you get the async `{ job_id }` to poll instead. ```bash curl -X POST https://api.bithuman.ai/v1/video/generate \ -H "api-secret: $BITHUMAN_API_SECRET" -H "Content-Type: application/json" \ -d '{"model": "essence-2", "agent_code": "'"$BITHUMAN_AGENT_CODE"'", "input": {"type": "text", "text": "Hello, welcome to bitHuman."}, "wait": true}' ``` ```json { "success": true, "job_id": "vid_3f9a2c1b8e7d4a6f0b21", "status": "completed", "video_url": "https://assets.bithuman.ai/.../vid_3f9a2c1b8e7d4a6f0b21.mp4", "duration_seconds": 6.5, "credits_charged": 4 } ``` Errors are returned at submit time, before any charge: `402 INSUFFICIENT_BALANCE` if your balance cannot cover the up-front maximum; `400` for an invalid `model` or `input`, or text over the limit; [`409 MODEL_NOT_GENERATED`](https://docs.bithuman.ai/api/errors#model-errors) if the agent does not have that model yet. Check the agent's `supported_models` ([poll status](https://docs.bithuman.ai/api/agents#poll-status)) or [add the model](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent). ## Get talking-video status `GET /v1/video/{job_id}` — poll a render job. ```bash curl https://api.bithuman.ai/v1/video/vid_3f9a2c1b8e7d4a6f0b21 -H "api-secret: $BITHUMAN_API_SECRET" ``` While rendering: ```json { "success": true, "job_id": "vid_3f9a2c1b8e7d4a6f0b21", "status": "processing", "model": "essence-2" } ``` When complete: ```json { "success": true, "job_id": "vid_3f9a2c1b8e7d4a6f0b21", "status": "completed", "model": "essence-2", "video_url": "https://assets.bithuman.ai/.../vid_3f9a2c1b8e7d4a6f0b21.mp4", "duration_seconds": 6.5, "credits_charged": 4 } ``` | Field | Type | Description | |---|---|---| | `status` | string | `processing`, `completed`, or `failed`. | | `model` | string | The engine used. | | `video_url` | string | Public mp4 URL (present when `completed`). | | `duration_seconds` | number | Output duration (present when `completed`). | | `credits_charged` | integer | Credits charged for this render (present when `completed`). | | `error` | object | Failure detail (present when `failed`); the charge is refunded. | > **Note** Read `video_url` from the response; never construct it. The storage host can change, so allowlist what the API returns. ## Polling pattern ```python import time, requests def wait_for_video(job_id, api_secret, timeout=600): while timeout > 0: r = requests.get( f"https://api.bithuman.ai/v1/video/{job_id}", headers={"api-secret": api_secret}, ).json() if r["status"] == "completed": return r["video_url"] if r["status"] == "failed": raise RuntimeError(r.get("error")) time.sleep(3) timeout -= 3 raise TimeoutError("render did not finish in time") ``` See [pricing](https://docs.bithuman.ai/pricing) for how credits are consumed. --- # Embedding URL: https://docs.bithuman.ai/api/embedding > Mint short-lived JWT tokens from your backend and embed a talking avatar on any website via an iframe. ## Embed an avatar Drop an agent onto any page as an iframe — no SDK install required: ```html ``` `A23WJF0199` is the `wise-pup` sample avatar (Expression 2). Replace it with your agent code — find it in the [Library](https://www.bithuman.ai/#library) or the Deploy & Share dialog. > **Warning** Keep `microphone *` (and `camera *` for camera chat) in `allow`, with the `*`. The embed redirects to another origin, so a bare `allow="microphone"` leaves the microphone silently blocked. A restrictive `Permissions-Policy` on your page blocks it too. ## Production: mint a token For per-visitor session tracking and rate limiting, mint a short-lived embed token on your **backend** (never expose your API secret in frontend code) and append it to the iframe URL. `POST /v1/embed-tokens/request` | Field | Type | Required | Description | |---|---|---|---| | `agent_id` | string | yes | Your agent's code. | | `fingerprint` | string | yes | Stable per-visitor string (any format). Used for per-visitor rate limiting and — if you run your own LLM — sent to your endpoint as the OpenAI `user` field so you can tell whose call it is ([details](https://docs.bithuman.ai/api/providers#knowing-which-end-user-a-call-belongs-to)). Supply one value per end user and reuse it across their visits. | | `model` | string | no | Optional model name: `essence-1`, `expression-1`, `essence-2` or `expression-2`. To pin a serving tier see [Models](https://docs.bithuman.ai/performance#pin-a-tier-for-a-benchmark). A model outside your plan returns `403 PLAN_REQUIRED`. Validated **early**: unknown values return `400` listing the accepted names; requesting a family the agent can't be launched as (missing from its `supported_models` — a trained model that doesn't exist yet) returns [`409 MODEL_NOT_GENERATED`](https://docs.bithuman.ai/api/errors#model-errors) instead of a failed session later. Omitted → the agent's own default model. | Every entry of `supported_models` in the mint response (and in `GET /v1/agent/status/{id}`) is a model name you can send back as `model` unchanged. ```js // server: mint token (api-secret never reaches the browser) const visitorFingerprint = "3f9a2c1b8e7d4a6f0b21c4d5e6f70812"; // one stable id per visitor, persisted const res = await fetch("https://api.bithuman.ai/v1/embed-tokens/request", { method: "POST", headers: { "api-secret": process.env.BITHUMAN_API_SECRET, "content-type": "application/json", }, body: JSON.stringify({ agent_id: process.env.BITHUMAN_AGENT_CODE, // your agent's code fingerprint: visitorFingerprint, }), }); const { data: { token } } = await res.json(); ``` ### Response ```json { "status": "success", "status_code": 200, "data": { "token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...", "sid": "f3c9...", "model": "expression-2", "supported_models": ["essence-1", "expression-2"] } } ``` The `token` is a **1-hour, HS256-signed JWT**. Mint one per visitor session. `supported_models` lists the canonical model families the agent can be launched as right now (useful for building your own model picker). The response always includes the `model` baked into the token (the agent's own model when you omit it). ### Use the token in the iframe Pass it as a query string: ```html ``` ## Session events Every conversation on an embed of your agent can POST `room.join` and `chat.push` events, including the text of every message, to a URL on your account. See [session events](https://docs.bithuman.ai/api/webhooks#session-events). ## Notes - The embed token is more constrained than a [runtime token](https://docs.bithuman.ai/api/authentication) — it's purpose-built for cross-origin iframe authentication. - WebRTC requires a secure context: serve the embedding page over **HTTPS** or the browser will block microphone access (except on `localhost`). - The `fingerprint` should be generated once per device and persisted, so per-visitor rate limits track the same visitor across sessions. - An embedded session bills to the agent's owner at the rates on [pricing](https://docs.bithuman.ai/pricing). See the interactive [API reference](https://docs.bithuman.ai/api/reference) for the full request and response schema. --- # Cloud avatar in your room URL: https://docs.bithuman.ai/api/cloud-avatar > Start a bitHuman cloud avatar in your own LiveKit room with one REST call, then stream it your TTS audio. No LiveKit plugin or agent framework needed. ## Overview A cloud avatar is a participant in your LiveKit room. It lip-syncs the audio you send it, and it publishes the video and that audio as its own tracks. The [LiveKit plugin](https://docs.bithuman.ai/platforms/livekit) does all of this for a Python LiveKit Agents worker. Use this page when your voice pipeline is something else. You will: 1. Mint a LiveKit join token for the avatar. 2. Start the session with `POST /v1/runtime-tokens/request`. The avatar then joins your room. 3. Send each reply's audio to the avatar as a LiveKit byte stream. Your clients subscribe to the avatar like any other participant. Call the endpoint from your server, because it takes your API secret. ## Mint the avatar's join token Mint the token with your LiveKit credentials, not with a bitHuman credential. It needs: - **Identity:** `bithuman-avatar-agent`. - **Kind:** `agent`. - **Grant:** permission to join the room. - **Attribute `lk.publish_on_behalf`:** the identity of the participant that sends the audio. The avatar takes audio only from that participant. If you leave this attribute out, the avatar takes audio from the first agent participant in the room. ```python from livekit import api avatar_token = ( api.AccessToken(LIVEKIT_API_KEY, LIVEKIT_API_SECRET) .with_identity("bithuman-avatar-agent") .with_kind("agent") .with_grants(api.VideoGrants(room_join=True, room=room_name)) .with_attributes({"lk.publish_on_behalf": sender_identity}) .to_jwt() ) ``` Don't put a bitHuman secret in the token or in its attributes. Everyone in the room can read attributes. Join your sender to the room as an agent participant too (kind `agent`). The avatar treats standard participants as users, and it stays in the room while any user is there. ## Start the session `POST https://api.bithuman.ai/v1/runtime-tokens/request` with the `api-secret` header: ```bash curl -X POST https://api.bithuman.ai/v1/runtime-tokens/request \ -H "api-secret: $BITHUMAN_API_SECRET" \ -H "content-type: application/json" \ -d '{ "mode": "gpu", "model": "essence-2", "agent_id": "", "livekit_url": "wss://your-project.livekit.cloud", "livekit_token": "", "room_name": "your-room" }' # → {"avatar_session_started": true, "model": "essence-2", ...} ``` | Field | Required | What it is | |---|---|---| | `mode` | yes | `"gpu"`. This field is what makes the call start an avatar; without it, the endpoint only issues a runtime token. The older value `"cpu"` also works but requires `agent_id`. | | `model` | recommended | `essence-1`, `essence-2`, `expression-1` or `expression-2`. If you leave it out, the server picks the model from `mode` and the agent's models, so check `model` in the response. | | `agent_id` | yes, unless `image` is sent | The agent code. | | `image` | for photo sessions | A portrait, for Expression 1 only. See [A photo instead of an agent](#a-photo-instead-of-an-agent). | | `livekit_url` | yes | Your LiveKit server URL. | | `livekit_token` | yes | The avatar's join token from the step above. | | `room_name` | yes | The room the avatar joins. | A `200` response means the session is starting. `avatar_session_started` is `true`, and `model` is the model the session launched as. The response also echoes `mode`, `agent_id` and `image`, and includes your `user_id`. The session is billed at the model's cloud rate for as long as it runs ([Pricing](https://docs.bithuman.ai/pricing)). It ends in any of these cases: - The last user leaves the room. - You remove `bithuman-avatar-agent` from the room. - The room closes. ## Send the audio Send each reply as one LiveKit byte stream: - **Topic:** `lk.audio_stream`. - **Destination identity:** `bithuman-avatar-agent`. - **Attributes:** `sample_rate` and `num_channels`, as strings. - **Payload:** raw 16-bit little-endian PCM. Closing the stream marks the end of the reply. Mono 16 kHz is the native format. Other sample rates are resampled, and stereo is mixed down to mono. ```python writer = await room.local_participant.stream_bytes( name="reply-1", topic="lk.audio_stream", destination_identities=["bithuman-avatar-agent"], attributes={"sample_rate": "16000", "num_channels": "1"}, ) async for pcm in tts_audio(): # 16-bit mono PCM chunks, as the TTS produces them await writer.write(pcm) await writer.aclose() # end of this reply ``` - **First reply:** wait until `bithuman-avatar-agent` has joined before you send it. Audio sent before the avatar joins is dropped. - **Interrupting:** perform the RPC `lk.clear_buffer` on `bithuman-avatar-agent`, then close the stream. - **Playback events (optional):** the avatar calls the RPCs `lk.playback_started` and `lk.playback_finished` on your sender. Register handlers for them if you want to know when it speaks. If you don't register them, the avatar skips these calls. ## Keep latency low - **Write each TTS chunk as soon as it exists.** Open the stream on the first chunk. Don't pace the audio to real time: the avatar buffers it and plays it at the right speed. The size of each write doesn't matter. - **Close the stream as soon as the reply ends.** The avatar renders speech in blocks of up to about a second of audio. A reply shorter than one block, like "Sure!", waits for the close. - **Keep one session for the whole conversation.** Starting a session connects to the room and loads the avatar, so don't start a new one for each turn. - **Run your pipeline in the US.** Cloud avatars render in the US. Put your LiveKit server or LiveKit Cloud project and your STT, LLM and TTS in a US region; US East is a good default. - **Speed up the voice pipeline.** Most of a slow turn is usually spent there: end-of-speech detection, the LLM's first token and the TTS's first audio. Streaming STT and streaming TTS help the most. ## A photo instead of an agent Expression 1 can animate a photo without an agent. Send `image`, leave out `agent_id`, and set `"model": "expression-1"`. `image` is a string in one of these forms: - A public `http(s)` URL. The server fetches it with a plain GET, so it must return `200` without authentication. A pre-signed URL works. - The image as base64, either raw or as a `data:image/jpeg;base64,…` URI. The endpoint takes JSON only. Photo requirements: - A JPEG or PNG. - One person, facing the camera, with the face clearly visible. - Head and shoulders, with room around the head. The server finds the face and crops a square about twice the face's width, with a little extra room above the head. It scales the crop to 512×512, so a face at least ~256 px wide keeps full detail. A face near the edge of the photo gets a clipped crop. If no face is found, the whole photo is used as it is, and a photo that isn't square looks stretched. Photo sessions work on Expression 1 only. On every other model, the request is refused with `400 VALIDATION_ERROR`, and nothing is launched or billed. To use another model, [create an agent](https://docs.bithuman.ai/api/agents#generate-an-agent) from the photo first. ## Errors | Status | Code | Cause | |---|---|---| | `400` | `VALIDATION_ERROR` | One of these: a field is missing or invalid; `image` was sent without `agent_id` for a model other than Expression 1; an Essence 1 session was started without `livekit_url`, `livekit_token` and `room_name`. Nothing is launched or billed. | | `401` | `MISSING_AUTH` / `UNAUTHORIZED` | The `api-secret` header is missing or invalid. | | `403` | `PLAN_REQUIRED` | The model is not in your plan. | | `403` | `CONCURRENCY_LIMIT_REACHED` | The session would exceed your plan's concurrent sessions ([Rate limits](https://docs.bithuman.ai/api/rate-limits)). | | `404` | `NOT_FOUND` | No agent with this code that you can start. | | `409` | `VALIDATION_ERROR` | The agent can't be served as the requested model, or its own model isn't ready yet. | | `503` | `SERVICE_UNAVAILABLE` | No capacity right now, or a temporary failure. Retry after the `Retry-After` header. | All error codes are listed on [Errors](https://docs.bithuman.ai/api/errors). --- # Realtime relay URL: https://docs.bithuman.ai/api/realtime > Open a metered OpenAI-Realtime voice session through bitHuman: a WebSocket relay or a server-brokered WebRTC call. ## Overview The Realtime API opens an OpenAI-Realtime voice session **through bitHuman**. Your API secret never leaves your server or app, no OpenAI key is involved, and the session is billed on your bitHuman balance as the chat line (see [Limits & billing](#limits--billing)). There are two ways in: - **WebSocket relay:** `wss://api.bithuman.ai/v1/realtime`, the OpenAI Realtime WebSocket protocol, unchanged. - **WebRTC:** `POST /v1/realtime/connect` with your SDP offer; bitHuman places the call and returns the SDP answer. Authenticate with the `api-secret` header (or `Authorization: Bearer `). ## WebSocket relay `wss://api.bithuman.ai/v1/realtime?model=gpt-realtime-mini` Speak the [OpenAI Realtime events](https://platform.openai.com/docs/api-reference/realtime) exactly as you would to OpenAI: `session.update`, `input_audio_buffer.append`, `response.create`, and so on. The model is fixed when you connect: a `session.update` that repeats it is accepted, and one that changes it is answered with an `error` event whose code is `MODEL_LOCKED`. ```python import asyncio, json, os, websockets async def main(): async with websockets.connect( "wss://api.bithuman.ai/v1/realtime?model=gpt-realtime-mini", additional_headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]}, ) as ws: print(json.loads(await ws.recv())["type"]) # session.created await ws.send(json.dumps({"type": "session.update", "session": { "type": "realtime", "instructions": "Answer in one sentence.", "output_modalities": ["audio"]}})) print(json.loads(await ws.recv())["type"]) # session.updated asyncio.run(main()) ``` The Flutter plugin (2.6.20+) and the CLI (2.8.1+) connect this way for you. ## WebRTC `POST /v1/realtime/connect?model=gpt-realtime-mini`, with the SDP offer as the request body (`Content-Type: application/sdp`). A `201` carries the SDP answer; set it as the remote description on your peer connection. The `X-Bithuman-Call` response header names the call. **End the call when you are done.** Billing stops the moment you hang up: ```bash curl -X POST https://api.bithuman.ai/v1/realtime/connect \ -H "api-secret: $BITHUMAN_API_SECRET" -H "content-type: application/json" \ -d '{"action": "hangup", "call": ""}' ``` ```js // browser: start the call from your server (it holds the api-secret), then on close: pc.close(); await fetch("/your-server/hangup", { method: "POST", body: JSON.stringify({ call }) }); // your server: POST https://api.bithuman.ai/v1/realtime/connect // {"action": "hangup", "call": call} with the same api-secret that started it ``` If a client only closes its peer connection without hanging up, the call ends when the connection times out, usually 8–10 seconds later, and those seconds are billed. ## Limits & billing - **Billing:** the chat line, **10 credits per minute, all-inclusive**, for every active second of the session (talking or idle). A self-hosted avatar rendering the same session is included: it is not billed on top. - **Balance:** a session needs a positive balance to start (`402 INSUFFICIENT_BALANCE`) and is closed when the balance runs out. - **Models:** a standard API secret uses `gpt-realtime-mini`; `gpt-realtime` needs an entitlement on your account (`403 PLAN_REQUIRED` otherwise; contact sales). - **Length:** one session lasts at most one hour. - Other errors: `401` missing or invalid key · `503` the relay is at capacity (retry after the `Retry-After` seconds). ## Retired: the client-secret mint The earlier realtime token mint, which handed the client an OpenAI client secret (`ek_…`), is **retired** and answers `410 ENDPOINT_RETIRED`. Connect through the relay or `/v1/realtime/connect` instead; CLI 2.8.1 and later already do. --- # Errors URL: https://docs.bithuman.ai/api/errors > The bitHuman API error format, HTTP status codes, and the full error-code catalog with resolution steps. ## Error response format Every error follows the same structured envelope: ```json { "error": { "code": "ERROR_CODE", "message": "Human-readable description of what went wrong.", "httpStatus": 401 }, "status": "error", "status_code": 401 } ``` The HTTP status always matches `status_code` and `error.httpStatus`; there is no "200 on error". Branch on either the status or `error.code`. ### 502 and 504 may not be JSON A `502` or `504` comes either from the API (the envelope above, code `UPSTREAM_UNAVAILABLE` or `UPSTREAM_TIMEOUT`) or from the delivery network in front of it (an HTML page). Check `Content-Type` before parsing any error, and fall back to the HTTP status line when it is not `application/json`. Both are transient: retry after `Retry-After` seconds when present, otherwise back off. ## HTTP status codes | Status | Meaning | Common cause | |---|---|---| | `200` | Success | Request completed. | | `302` | Redirect | Not an error — [`GET /v1/agent/{code}/model/download`](https://docs.bithuman.ai/api/agents#download-an-agents-model) redirects to the artifact URL by default. | | `400` | Bad Request | Malformed JSON, missing required parameter (`MISSING_PARAM`), failed validation (`VALIDATION_ERROR`), or a request that can never succeed as posed (`MODEL_NOT_DOWNLOADABLE`). | | `401` | Unauthorized | Invalid `api-secret` (`UNAUTHORIZED`) or absent `api-secret` header (`MISSING_AUTH`). | | `402` | Payment Required | Insufficient credits — top up to continue. | | `403` | Forbidden | The credential is known but refused here: a revoked secret on a token endpoint (`RUNTIME_SUSPENDED`), a session limit (`CONCURRENCY_LIMIT_REACHED`, `SESSION_DURATION_LIMIT`), a model outside your plan (`PLAN_REQUIRED`), or a secret's value read with an API secret (`SECRET_REVEAL_CONSOLE_ONLY`). | | `404` | Not Found | Agent, resource, or endpoint doesn't exist — or a model artifact not published to the download store yet (`MODEL_ARTIFACT_NOT_READY`, retryable). | | `410` | Gone | The agent was deleted (`AGENT_DELETED`, `AGENT_PURGED`). | | `409` | Conflict | The request is valid but the agent's **state** doesn't allow it yet (`MODEL_NOT_GENERATED`, `AGENT_NOT_READY`) — a state change (generate/add the model, wait for `ready`) fixes it. | | `413` | Payload Too Large | File exceeds the size limit. | | `415` | Unsupported Media Type | File type not supported. | | `422` | Unprocessable Entity | The request is well-formed but semantically incompatible with the target model (`MODEL_SUBJECT_MISMATCH`, `MODEL_PREREQUISITE_MISSING`) — change the input or asset, not the request syntax. | | `429` | Rate Limited | Too many requests — see [rate limits](https://docs.bithuman.ai/api/rate-limits). | | `500` | Internal Error | Server-side error — retry or contact support. | | `502` / `504` | Bad Gateway / Gateway Timeout | Transient. JSON from the API, HTML from the delivery network ([above](#502-and-504-may-not-be-json)). Retry after `Retry-After`. | | `503` | Service Unavailable | Temporarily unavailable or at capacity (`SERVICE_UNAVAILABLE`), or a model paused for your account (`MODEL_NOT_YET_AVAILABLE`). Retry after `Retry-After`, with backoff. | ## Error codes ### Authentication | Code | HTTP | Resolution | |---|---|---| | `UNAUTHORIZED` | 401 | The `api-secret` header is present but invalid. Get a valid secret from [Developer → API Secrets](https://www.bithuman.ai/developer/api-keys). | | `MISSING_AUTH` | 401 | The `api-secret` header is absent. Add it to your request. | | `RUNTIME_SUSPENDED` | 403 | A token endpoint refused the secret: it was revoked (create a new one), or runtime access is suspended (contact support). | | `ACCOUNT_SUSPENDED` | 403 | Your balance is too far below zero. Top up; contact support if it persists. | | `PLAN_REQUIRED` | 403 | Your plan does not include this. The `message` says which case applies: **(a)** the model you named is not in your plan; the message names the plan. [Contact sales](https://www.bithuman.ai/sales). **(b)** Free-plan API/SDK use from **2026-10-12 00:00 UTC**: "API and SDK access starts at the Creator plan. Upgrade at https://www.bithuman.ai/pricing to keep using your API secret." Until then, responses to Free accounts carry a `plan_notice` field and an `X-Bithuman-Plan-Notice` header: "Free-plan API and SDK access ends on 2026-10-12. Upgrade at https://www.bithuman.ai/pricing to keep it." A Free account holding top-up credits purchased before 2026-09-27 keeps API/SDK access until those credits are used up. **(c)** Free-plan agent creation: "Creating agents starts at the Creator plan. Upgrade at https://www.bithuman.ai/pricing." **(d)** A Free-plan credit top-up: "Credit top-ups start at the Creator plan. Upgrade at https://www.bithuman.ai/pricing." [Upgrade](https://www.bithuman.ai/pricing). | | `AGENT_LIMIT_REACHED` | 403 | Creating an agent would exceed your plan's agent limit (Creator 7, Pro 40, Business 200, Enterprise unlimited): "Your {Plan} plan includes {N} agents and you have {M}. Existing agents keep working; delete one or upgrade at https://www.bithuman.ai/pricing to create more." Existing agents are never removed. | | `SECRET_REVEAL_CONSOLE_ONLY` | 403 | An API secret tried to read a stored secret's value. Reveal it in the console, or create a new secret. | | `INSUFFICIENT_BALANCE` | 402 | Top up credits at [www.bithuman.ai](https://www.bithuman.ai). | ### Agent operations | Code | HTTP | Resolution | |---|---|---| | `NOT_FOUND` | 404 | Returned both when no agent matches the code **and** when an agent has no active session for `/speak` / `/add-context`. Distinguish by the `message` string: `"Agent not found for code: "` vs `"No active rooms found for agent "`. | | `VALIDATION_ERROR` | 400 | Body failed schema validation. Include all required fields. | | `VIDEO_INPUT_NOT_SUPPORTED` | 400 | [Agent creation](https://docs.bithuman.ai/api/agents#generate-an-agent) with a `video` input. Creation is **image-only** for every model — provide a portrait `image`; bitHuman generates the 10-second identity video internally so it loops without a seam (first frame == last frame). Nothing is charged; never send `video`. | | `MISSING_PARAM` | 400 | A required parameter was not provided. | | `IMAGE_FACE_UNSUITABLE` | 422 | Essence 2 creation or add: the face is too small, missing, or one of several similar faces. Upload a closer waist-up or head-and-shoulders photo of one person. Nothing is charged. | | `AGENT_DELETED` / `AGENT_PURGED` | 410 | The agent was deleted; model download, gestures and talking video refuse it. Create a new agent. | ### Model errors The model-release surfaces — [creation](https://docs.bithuman.ai/api/agents#generate-an-agent), [model add](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent), [model download](https://docs.bithuman.ai/api/agents#download-an-agents-model), the [embed-token `model` field](https://docs.bithuman.ai/api/embedding), and [talking video](https://docs.bithuman.ai/api/video) — share these codes: | Code | HTTP | Resolution | |---|---|---| | `MODEL_NOT_GENERATED` | 409 | The requested model family isn't in the agent's `supported_models` — it can't be launched (or downloaded) as that family yet. **The message names the fix**: the exact [model-add](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent) call and its cost when this agent qualifies for it, or the missing asset when it doesn't. Trained families (`expression-2`, `essence-2`): `"agent 's model hasn't been generated yet — add it with POST /v1/agent//models …"`. `expression-1` reads `"isn't enabled on this agent yet"` instead — nothing is ever trained for it, and the add takes effect at once and is **free** (see [Add a model to an existing agent](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent)). Checked **before any charge**. | | `AGENT_NOT_READY` | 409 | [`POST /v1/agent/{code}/models`](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent) on an agent that is still generating or failed. Wait for the current generation to finish, or fix/re-create a failed agent first. | | `MODEL_SUBJECT_MISMATCH` | 422 | An explicit Essence 2 creation or add whose input is not a **photorealistic human subject** — e.g. `"essence-2 requires a photorealistic human subject; this image looks like a cartoon — use expression-2"`. Nothing is billed and no agent row is created. Use `expression-2` for stylized/non-human subjects, or `model: "auto"` to route automatically. See [the subject gate](https://docs.bithuman.ai/api/agents#generate-an-agent). | | `MODEL_PREREQUISITE_MISSING` | 422 | A [model add](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent) needs a stored asset this agent doesn't have — a stored identity video for `essence-2` (generated internally by Essence creations, never uploaded), face image for `expression-2`, image + voice for `expression-1`, stored identity video or image for `essence-1`. Add the missing image/voice asset, then retry. | | `MODEL_NOT_DOWNLOADABLE` | 400 | [Model download](https://docs.bithuman.ai/api/agents#download-an-agents-model) for a family with no per-identity artifact — `expression-1` renders server-side from the agent's image. A `400` because no state change can fix it (unlike the 409s). | | `MODEL_NOT_YET_AVAILABLE` | 503 | A model is paused for your account (not returned in normal operation). Nothing is charged; retry later or use another model. | | `MODEL_ARTIFACT_NOT_READY` | 404 | [Model download](https://docs.bithuman.ai/api/agents#download-an-agents-model) for a **supported** family whose artifact hasn't been published to the download store yet. Retryable — the message carries a per-family retry hint; poll on this code. | ### File operations | Code | HTTP | Resolution | |---|---|---| | `FILE_TOO_LARGE` | 413 | Images 10 MB, video 100 MB, audio and documents 25 MB. | | `UNSUPPORTED_TYPE` | 415 | The bytes are not a supported type, they contradict `file_type`, the filename extension is not one [File upload](https://docs.bithuman.ai/api/files) lists, or the file would run in a browser (SVG, HTML, scripted text). | | `DOWNLOAD_FAILED` | 400 | Ensure the URL is publicly accessible and returns a valid file. | ### Session & infrastructure | Code | HTTP | Resolution | |---|---|---| | `RATE_LIMITED` | 429 | Back off and retry. See [rate limits](https://docs.bithuman.ai/api/rate-limits). | | `CONCURRENCY_LIMIT_REACHED` | 403 | A new session start would exceed your plan's [concurrent avatar session allowance](https://docs.bithuman.ai/api/rate-limits#session-concurrency). End an active session or upgrade the plan, then retry — live sessions are never cut off mid-stream by this limit. | | `SESSION_DURATION_LIMIT` | 403 | One session ran past the maximum continuous length. Start a new session; your account is fine. | | `SERVICE_UNAVAILABLE` | 503 | A dependency is briefly unavailable or at capacity. Nothing was changed. Retry after `Retry-After` seconds, with backoff. | | `UPSTREAM_UNAVAILABLE` / `UPSTREAM_TIMEOUT` | 502 / 504 | Transient. Retry with backoff. | | `INTERNAL_ERROR` | 500 | Retry once. If persistent, report via [Discord](https://discord.gg/ES953n7bPA). | ### Text to speech | Code | HTTP | Resolution | |---|---|---| | `VOICE_NOT_FOUND` | 404 | Unknown voice code. List voices with `GET /v1/voices`. | ## Handling errors in Python The Python examples use [`requests`](https://pypi.org/project/requests/) (`pip install requests`). This call is free: it asks for the status of an agent that does not exist. ```python import os, requests resp = requests.get( "https://api.bithuman.ai/v1/agent/status/A00XXX0000", headers={"api-secret": os.environ["BITHUMAN_API_SECRET"]}, timeout=30, ) if resp.ok: print(resp.json()["data"]["status"]) elif resp.headers.get("content-type", "").startswith("application/json") and "error" in resp.json(): err = resp.json()["error"] print(resp.status_code, err["code"], err["message"]) # 404 NOT_FOUND here else: print(resp.status_code, "retry after", resp.headers.get("Retry-After")) ``` Retry `429` and `5xx`. Fix the request for any other `4xx`. For `429` and `5xx`, use exponential backoff with jitter — see [rate limits](https://docs.bithuman.ai/api/rate-limits) for the recommended retry strategy. --- # Rate limits URL: https://docs.bithuman.ai/api/rate-limits > Plan-tiered request limits by endpoint cost tier, the 429 / Retry-After contract, session concurrency, and a recommended retry strategy. ## Request limits Requests are limited **per account** (every API secret on the account shares the same buckets), by endpoint cost tier and plan. Each cell is requests per minute; a bucket also allows a burst of that size and refills continuously. | Cost tier | Free | Creator | Pro | Business | Enterprise* | |---|---|---|---|---|---| | **Generate** | 4 | 10 | 30 | 60 | 120 | | **Write** | 30 | 60 | 180 | 360 | 720 | | **Read** | 120 | 240 | 720 | 1440 | 2880 | \* Enterprise defaults; custom limits are available from [sales](https://www.bithuman.ai/sales). | Cost tier | Covers | |---|---| | **Generate** | heavy jobs: `POST /v1/agent/generate`, `POST /v1/dynamics/generate`, video generation | | **Write** | every other `POST`, `PUT`, `PATCH`, `DELETE`, including `POST /v1/tts`, plus `GET /v1/agent/{code}/sessions` | | **Read** | other `GET` requests, such as `GET /v1/agent/status/*` and `GET /v2/credit-summaries` | A plan change reaches the limiter within about a minute; no new secret is needed. Over the limit, the API returns `429 RATE_LIMITED` with a `Retry-After` header and the standard [error envelope](https://docs.bithuman.ai/api/errors). **Never limited:** webhook deliveries, and the runtime-token routes (`/v1/runtime-tokens*`, `/v1/runtime/*`) that keep a live session authenticated. A live session is never cut off with a `429`. **Failed authentication** is throttled per client IP at 30 failures per minute. Routes that check the secret themselves (embed-token and realtime mints, `/v1/me`, CLI sign-in) are limited per client IP, at 120 requests per minute, instead of per account. ## Session concurrency | Plan | Concurrent cloud avatar sessions | |---|---| | Free | 1 | | Creator | 3 | | Pro | 10 | | Business | 50 | | Enterprise | 200 | | Custom (contact sales) | Unlimited | A session over the allowance is refused at start with `403 CONCURRENCY_LIMIT_REACHED`; a live session is never cut off by this limit. Agent and dynamics generation jobs queue and run as capacity frees up. **Session length.** One continuous session can run up to 24 hours in the cloud and 7 days self-hosted. It then ends with `403 SESSION_DURATION_LIMIT`; start a new session to continue. For longer unattended installs (kiosks), [contact sales](https://www.bithuman.ai/sales). Sessions you render on your own hardware are limited only by your credits ([self-hosting](https://docs.bithuman.ai/deploy/self-hosted)). Credits pay for session time, talking or idle, by the exact second ([pricing](https://docs.bithuman.ai/pricing)). ## Response headers Metered endpoints return your current state, so you can slow down before a `429`: | Header | Meaning | |---|---| | `X-RateLimit-Limit` | your plan's limit for this request's cost tier | | `X-RateLimit-Remaining` | whole requests left right now | | `X-RateLimit-Reset` | Unix time when the bucket is full again | | `Retry-After` | on `429` only: seconds to wait | | `X-Request-Id` | include it when you contact support | Not every response carries them (`POST /v1/validate` has none); fall back to backoff. ## Recommended retry strategy Retry `429` and `5xx` with exponential backoff and jitter, honouring `Retry-After`: ```python import time, random, requests def call(method, url, max_retries=5, **kw): for attempt in range(max_retries): resp = requests.request(method, url, timeout=30, **kw) if resp.status_code not in (429, 502, 503, 504): return resp wait = float(resp.headers.get("Retry-After", 2 ** attempt)) time.sleep(wait + random.uniform(0, 1)) return resp ``` Usage: `call("POST", "https://api.bithuman.ai/v1/tts", headers=h, json={...})`. ## Best practices - **Use [webhooks](https://docs.bithuman.ai/api/webhooks), not polling,** for `agent.ready` and `agent.failed`. If you poll status, poll every 5 seconds or slower. - **Cache agent details** from `GET /v1/agent/{code}`; they rarely change. - **Reuse a session** for back-to-back conversations rather than starting a new one; note that an open session bills its time, talking or idle. - **Check your balance** with `GET /v2/credit-summaries` before creating an agent ([creation costs](https://docs.bithuman.ai/pricing#creation--one-time-credits)), to avoid a `402`. More capacity comes with a higher [plan](https://docs.bithuman.ai/pricing#plans); for more than Enterprise, [talk to sales](https://www.bithuman.ai/sales).