# bitHuman — Get started (choose your path, your API secret, performance, FAQ, glossary) > Realtime talking avatars from one portrait. Index: https://docs.bithuman.ai/llms.txt · every page is also served as markdown at .md · OpenAPI: https://docs.bithuman.ai/api/openapi.yaml · other sections: https://docs.bithuman.ai/llms/platforms.txt · https://docs.bithuman.ai/llms/deploy.txt · https://docs.bithuman.ai/llms/models.txt · https://docs.bithuman.ai/llms/build.txt · https://docs.bithuman.ai/llms/api.txt ## Instructions for AI agents - Always send `model` ("essence-2" or "expression-2") when creating an agent; poll until `status` is `ready` or `failed`. - API and SDK use requires the Creator plan or higher. Never tell a user they can build on a free plan. - In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET` and pass a minted token (https://docs.bithuman.ai/platforms/livekit.md). - Say where things happen: the avatar renders (device, browser, server, cloud); the conversation runs (your stack, the CLI's local conversation brain, or bitHuman's servers). - Quote speed only from https://docs.bithuman.ai/performance.json (× real time, with the device). Quote prices only from GET https://api.bithuman.ai/v1/pricing. - Do not claim offline on phones, Mac or the browser, a conversation brain on phones, or any certification. Where each model runs, the key facts and one command per path: https://docs.bithuman.ai/llms.txt ## Contents - Quickstart — https://docs.bithuman.ai/start - Your API secret — https://docs.bithuman.ai/start/api-secret - Performance — https://docs.bithuman.ai/performance Linked, not inlined (read the .md twin): - How we measure — https://docs.bithuman.ai/performance/method.md - FAQ — https://docs.bithuman.ai/resources/faq.md - Glossary — https://docs.bithuman.ai/resources/glossary.md - Examples: https://docs.bithuman.ai/examples.md · changelog: https://docs.bithuman.ai/changelog.md · API references: https://docs.bithuman.ai/platforms/cli/reference.md, https://docs.bithuman.ai/platforms/python/reference.md, https://docs.bithuman.ai/platforms/swift/reference.md, https://docs.bithuman.ai/platforms/android/reference.md --- # Quickstart URL: https://docs.bithuman.ai/start ## Choose your platform | You want to… | Use | Needs | First command | |---|---|---|---| | Put an avatar on a website | Web embed | nothing | `` | | Call it from any backend | REST API | API secret | `curl -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET"` | | Render or stream from Python | Python | API secret | `pip install "bithuman[expression-2]"` | | Run it from a terminal | CLI (macOS arm64, Linux x86_64 / arm64) | sign-in | `curl -fsSL https://install.bithuman.ai \| sh` | | Ship an iPhone, iPad or Mac app | iOS & iPadOS (Swift package) | Xcode 26+, API secret | `.package(url: "https://github.com/bithuman-product/homebrew-bithuman.git", from: "2.19.0")` | | Ship an Android app | Android | arm64 device, API secret | `implementation("ai.bithuman:expression2-android:0.5.2")` | | Add a face to a LiveKit voice agent | LiveKit | API secret | `pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv` | | Drive it from Claude or Cursor | MCP server | sign-in | `claude mcp add bithuman -- bithuman mcp` | | Run fully offline (kiosk, trade show, ATM) | Fully offline | Business or Enterprise plan; Essence 1 on Linux | — ([contact sales](https://www.bithuman.ai/sales)) | Fully offline: Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. ## Run it ### Put an avatar on a website: Web embed ```html ``` Expected: A live avatar in your page that listens and answers. Allow the microphone when the browser asks. ### Call it from any backend: REST API ```bash export BITHUMAN_API_SECRET="" curl -s -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET" # → {"valid":true} ``` Expected: Your API secret works. The API quickstart continues with speech, an agent and a talking video. ### Render or stream from Python: Python ```bash python3 -m venv .venv && source .venv/bin/activate pip install "bithuman[expression-2]" export BITHUMAN_API_SECRET="" curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2" curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav python -c 'import bithuman with bithuman.open("wise-pup.imx") as a: print(sum(1 for _ in a.render("speech.wav")), "frames")' # → 300 frames ``` Expected: 300 frames of 416×720 video: the 15 seconds of sample speech. ### Run it from a terminal: CLI (macOS arm64, Linux x86_64 / arm64) ```bash # macOS: brew install ffmpeg · Debian/Ubuntu: sudo apt install -y ffmpeg curl -fsSL https://install.bithuman.ai | sh export PATH="$HOME/.local/bin:$PATH" bithuman login curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav bithuman render wise-pup speech.wav -o out.mp4 # → out.mp4: 15 s of a talking avatar ``` Expected: out.mp4, 15 seconds of the avatar saying the sample. `bithuman run wise-pup` opens a live conversation instead. --- # Your API secret URL: https://docs.bithuman.ai/start/api-secret > The one credential for every bitHuman surface: where to get it, where each platform reads it, and what a shipped app holds. One API secret works on every surface: the REST API, the CLI, Python, Apple, Android and LiveKit. Credits pay for session time, talking or idle, by the exact second ([pricing](https://docs.bithuman.ai/pricing)). ## Get one Create an API secret under [Developer → API Secrets](https://www.bithuman.ai/developer/api-keys), then export it: ```bash export BITHUMAN_API_SECRET="" curl -s -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET" # → {"valid":true} ``` ## Where each platform reads it | Platform | Environment | In code | |---|---|---| | REST API | — | header `api-secret` | | CLI | `BITHUMAN_API_SECRET` | `bithuman login` stores a credential for you | | Python | `BITHUMAN_API_SECRET` (read by `bithuman.open`) | `api_secret=` on `AsyncBithuman.create()` only; `bithuman.open()` reads the environment | | Apple | `BITHUMAN_API_SECRET` | `Essence2Credential.set` / `Expression2Credential.set` | | Android | — | `Essence2Credential.set(secret)` / `Expression2Credential.set(secret)` before `fetch()` and `create()`; fetch the secret from your backend in a shipped app | | LiveKit worker | `BITHUMAN_MASTER_SECRET`, never `BITHUMAN_API_SECRET` | a short-lived token minted from it, never the secret ([LiveKit](https://docs.bithuman.ai/platforms/livekit#authenticate)) | | Web embed | — | none for a public agent; an [embed token](https://docs.bithuman.ai/api/embedding) for a private one | `BITHUMAN_API_KEY` is a deprecated alias of `BITHUMAN_API_SECRET`; rename it. It stops being read in CLI 3.0 and bithuman 4.0 (no earlier than 2026-12-26). ## Keep it safe - Keep the secret in the environment or a secrets manager, never in source control or on a command line. - Browsers and LiveKit rooms get short-lived tokens: an [embed token](https://docs.bithuman.ai/api/embedding) or a runtime token minted with [`POST /v1/runtime-tokens/mint`](https://docs.bithuman.ai/platforms/livekit#authenticate). - If a secret leaks, create a new one and delete the old one in the console. ## What a shipped app holds The Swift package and the Android SDK authenticate with your API secret, so every copy of an app you distribute carries it. Treat that secret as exposed: whoever extracts it can call the API as your account, spend your credits and create more secrets. - Fetch the secret from your backend when the app starts. Never compile it into a build you ship. - Give each app its own secret, so you can rotate one without touching the others ([API secrets](https://www.bithuman.ai/developer/api-keys)). - Watch your [balance](https://docs.bithuman.ai/pricing#check-your-balance), and rotate the secret at once if usage looks wrong. ## Next - [REST authentication](https://docs.bithuman.ai/api/authentication): the header, validation and error codes. - [Pick your platform](https://docs.bithuman.ai/start#choose-your-platform). --- # Performance URL: https://docs.bithuman.ai/performance > How fast Essence 2 and Expression 2 render on iPhone, Android, a WebGPU browser, a Mac, a Linux PC with no GPU and the bitHuman cloud, in times real time, with the raw data as performance.json. **× real time** is seconds of avatar video rendered per second, rounded down to one decimal: at 1.0× or more, an avatar holds a live conversation. Every configuration we publish renders faster than real time. The tables below list every published configuration. | Runs on | Hardware | Essence 2 fps | Essence 2 × real time | Expression 2 fps | Expression 2 × real time | |---|---|---|---|---|---| | Cloud API · GPU | NVIDIA RTX 4090 | 104 | **4.1×** real time | 340 | **17.0×** real time | | Cloud API · Apple silicon | Apple M4 Max | 71 | **2.8×** real time | 111 | **5.5×** real time | | Cloud API · CPU | x86 server CPU | 28 | **1.1×** real time | 27 | **1.3×** real time | | macOS · CLI | Apple M4 | 106 | **4.2×** real time | 168 | **8.4×** real time | | macOS · Python | Apple M4 | 174 | **6.9×** real time | 169 | **8.4×** real time | | macOS · Swift package | Apple M4 | 120 | **4.8×** real time | 177 | **8.8×** real time | | Linux · CLI | Intel Core i7-13700F (x86_64) | 50 | **2.0×** real time | 44 | **2.2×** real time | | Linux · Python | Intel Core i7-13700F (x86_64) | 49 | **1.9×** real time | 47 | **2.3×** real time | | iPhone · Swift package | iPhone 15 | 54 | **2.1×** real time | 111 | **5.5×** real time | | Android | Samsung Galaxy S25+ | 52 | **2.0×** real time | 48 | **2.4×** real time | | Web browser (WebGPU) | Chrome on Apple M4 | 43 | **1.7×** real time | 39 | **1.9×** real time | Each figure is one avatar session rendering as fast as the hardware allows. On the cloud API the service picks the tier for each session; the three Cloud API rows show each tier. The sections below go platform by platform, on-device first. [How we measure](https://docs.bithuman.ai/performance/method) has the method, the releases measured, memory, the time to a finished video and the raw data. ## Mobile First a short burst on each phone, then one session held for 10 minutes. | Runs on | Hardware | Essence 2 fps | Essence 2 × real time | Expression 2 fps | Expression 2 × real time | |---|---|---|---|---|---| | iPhone · Swift package | iPhone 15 | 54 | **2.1×** real time | 111 | **5.5×** real time | | Android | Samsung Galaxy S25+ | 52 | **2.0×** real time | 48 | **2.4×** real time | Measured in September 2026 on Swift package 2.17.3, Swift package 2.18.0, essence2-android 0.7.0 and expression2-android 0.4.10. - The first table is one render of a speech clip on a cool phone, as fast as the phone allows. - Some phone figures use a shorter speech clip than the other platforms; each cell's clip is in [performance.json](https://docs.bithuman.ai/performance.json). - Only an iPhone 15 and a Samsung Galaxy S25+ are measured. Other phones render at other rates. Setup for each SDK: [Apple](https://docs.bithuman.ai/platforms/ios), [Android](https://docs.bithuman.ai/platforms/android). ## Held for 10 minutes | Runs on | Hardware | Essence 2 fps | Essence 2 × real time | Expression 2 fps | Expression 2 × real time | |---|---|---|---|---|---| | iPhone · Swift package | iPhone 15 | 33 | **1.3×** real time | 103 | **5.1×** real time | | Android | Samsung Galaxy S25+ | 37 | **1.4×** real time | 44 | **2.2×** real time | | Web browser (WebGPU) | Chrome on Apple M4 | 54 | **2.1×** real time | 41 | **2.0×** real time | One session held open for ten minutes from a cool start on Swift package 2.15.0, essence2-android 0.8.1, expression2-android 0.5.2 and web viewer, rendering as fast as the device allows. Each number is the median 30-second stretch of the slowest of that row's sessions (four on iPhone · Swift package, three on Android, three on Web browser (WebGPU) Essence 2, four on Web browser (WebGPU) Expression 2); the slowest single stretch was lower (iPhone · Swift package Essence 2 31 fps; iPhone · Swift package Expression 2 100 fps; Android Essence 2 32 fps; Android Expression 2 37 fps; Web browser (WebGPU) Essence 2 50 fps; Web browser (WebGPU) Expression 2 41 fps). A phone warms up over a long conversation and slows its processor to stay cool, so a kiosk or any screen that renders all day should plan on this number rather than the short-burst rate. Memory over the ten minutes (the probe's own reading at the start and the end of the held window): Android Essence 2 memory (PSS) 0.9 GB to 0.9 GB (+0 MB); Android Expression 2 memory (PSS) 0.7 GB to 0.8 GB (+34 MB); Web browser (WebGPU) Essence 2 Chrome tab memory footprint 3.1 GB to 3.0 GB (-55 MB). The Web browser row is the engine's render throughput with WebGPU in Chrome on an Apple M4, measured in a visible (headed) browser window. It is not the frame rate a visitor sees on the page, where the voice and the display share the browser with the engine. ## Web These figures are for the avatar rendering in the visitor's own tab (`render=local`). By default the avatar renders on the [cloud API](https://docs.bithuman.ai/performance#cloud) and streams to the page. | Runs on | Hardware | Essence 2 fps | Essence 2 × real time | Expression 2 fps | Expression 2 × real time | |---|---|---|---|---|---| | Web browser (WebGPU) | Chrome on Apple M4 | 43 | **1.7×** real time | 39 | **1.9×** real time | Measured in September 2026 on the hosted web viewer. - Each figure is the engine's render throughput with WebGPU in Chrome on an Apple M4, measured in an automated browser. - It is not the frame rate a visitor sees on the page, where the voice and the display share the browser with the engine. - Only Chrome on an Apple M4 is measured. Other browsers, other GPUs and devices without WebGPU are not. Setup and the WebGPU check: [Web](https://docs.bithuman.ai/platforms/web). ## Desktop Pick the row for the product you use: the CLI, Python and the Swift package render at different rates on the same machine. | Runs on | Hardware | Essence 2 fps | Essence 2 × real time | Expression 2 fps | Expression 2 × real time | |---|---|---|---|---|---| | macOS · CLI | Apple M4 | 106 | **4.2×** real time | 168 | **8.4×** real time | | macOS · Python | Apple M4 | 174 | **6.9×** real time | 169 | **8.4×** real time | | macOS · Swift package | Apple M4 | 120 | **4.8×** real time | 177 | **8.8×** real time | | Linux · CLI | Intel Core i7-13700F (x86_64) | 50 | **2.0×** real time | 44 | **2.2×** real time | | Linux · Python | Intel Core i7-13700F (x86_64) | 49 | **1.9×** real time | 47 | **2.3×** real time | Measured in September 2026 on CLI 2.8.1, bithuman 2.11.12, bithuman 2.11.13 and Swift package 2.15.0. - Each figure is one render of a reference speech clip, as fast as the machine allows, with nothing else running. - The macOS CLI Essence 2 figure was measured with `BITHUMAN_THREADS=8`; by default the CLI uses one thread per CPU it may use, up to 16. The Linux CLI uses default settings. - Only these two machines are measured: an Apple M4 Mac and an Intel Core i7-13700F desktop. Other processors render at other rates. Memory per render is on [How we measure](https://docs.bithuman.ai/performance/method#memory). Setup for each product: [CLI](https://docs.bithuman.ai/platforms/cli), [Python](https://docs.bithuman.ai/platforms/python), [Apple](https://docs.bithuman.ai/platforms/ios). ## Cloud By default the service picks the tier for each session; each row is one tier. To benchmark one tier you can pin it ([below](#pin-a-tier-for-a-benchmark)); in production, let the service choose. | Runs on | Hardware | Essence 2 fps | Essence 2 × real time | Expression 2 fps | Expression 2 × real time | |---|---|---|---|---|---| | Cloud API · GPU | NVIDIA RTX 4090 | 104 | **4.1×** real time | 340 | **17.0×** real time | | Cloud API · Apple silicon | Apple M4 Max | 71 | **2.8×** real time | 111 | **5.5×** real time | | Cloud API · CPU | x86 server CPU | 28 | **1.1×** real time | 27 | **1.3×** real time | Measured in September 2026 on the cloud API. - Each figure is how fast one finished video is delivered, including encoding the video file, on a server with no other sessions. - With other sessions on the same server, a session can render more slowly than shown. - A live conversation plays at the model's own rate, 25 fps for Essence 2 and 20 fps for Expression 2. Speed above that makes a video file finish sooner; it does not put more frames on screen. ### Pin a tier for a benchmark To measure one tier, append `?model=` with a tier slug to the viewer or embed URL: ```text https://www.bithuman.ai/embed/A23WJF0199?model=expression-2-apple ``` | Model | Tier slugs | |---|---| | `essence-2` | `essence-2-gpu` · `essence-2-apple` · `essence-2-cpu` | | `expression-2` | `expression-2-gpu` · `expression-2-apple` · `expression-2-cpu` | - **A recognized slug pins the session** to that tier: if the tier is unavailable, the session fails rather than playing elsewhere. - **An unrecognized slug is ignored** and the session plays as usual. If a pin seems to have no effect, check the spelling. - **To be told about a typo**, set the embed token's `model` field instead: an unknown value is refused with a `400` listing the accepted names when you [mint the token](https://docs.bithuman.ai/api/embedding#production-mint-a-token). In production, omit `?model=` and let the service choose.