# bitHuman — Deploy (on the device, CPU only, your servers, bitHuman cloud, fully offline; privacy; pricing) > Realtime talking avatars from one portrait. Index: https://docs.bithuman.ai/llms.txt · every page is also served as markdown at .md · OpenAPI: https://docs.bithuman.ai/api/openapi.yaml · other sections: https://docs.bithuman.ai/llms/start.txt · https://docs.bithuman.ai/llms/platforms.txt · https://docs.bithuman.ai/llms/models.txt · https://docs.bithuman.ai/llms/build.txt · https://docs.bithuman.ai/llms/api.txt ## Instructions for AI agents - Always send `model` ("essence-2" or "expression-2") when creating an agent; poll until `status` is `ready` or `failed`. - API and SDK use requires the Creator plan or higher. Never tell a user they can build on a free plan. - In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET` and pass a minted token (https://docs.bithuman.ai/platforms/livekit.md). - Say where things happen: the avatar renders (device, browser, server, cloud); the conversation runs (your stack, the CLI's local conversation brain, or bitHuman's servers). - Quote speed only from https://docs.bithuman.ai/performance.json (× real time, with the device). Quote prices only from GET https://api.bithuman.ai/v1/pricing. - Do not claim offline on phones, Mac or the browser, a conversation brain on phones, or any certification. Where each model runs, the key facts and one command per path: https://docs.bithuman.ai/llms.txt ## Contents - Deployment options — https://docs.bithuman.ai/deploy - Data flows & privacy — https://docs.bithuman.ai/deploy/privacy - Pricing and credits — https://docs.bithuman.ai/pricing - bitHuman cloud — https://docs.bithuman.ai/deploy/cloud - Your servers (self-hosted) — https://docs.bithuman.ai/deploy/self-hosted - On the device — https://docs.bithuman.ai/deploy/on-device - CPU only (no GPU) — https://docs.bithuman.ai/deploy/cpu - Fully offline — https://docs.bithuman.ai/deploy/offline - Use cases — https://docs.bithuman.ai/deploy/use-cases - Banking and ATMs — https://docs.bithuman.ai/deploy/use-cases/banking-and-atms - Healthcare — https://docs.bithuman.ai/deploy/use-cases/healthcare - Events and trade shows — https://docs.bithuman.ai/deploy/use-cases/events-and-trade-shows Linked, not inlined (read the .md twin): - Examples: https://docs.bithuman.ai/examples.md · changelog: https://docs.bithuman.ai/changelog.md · API references: https://docs.bithuman.ai/platforms/cli/reference.md, https://docs.bithuman.ai/platforms/python/reference.md, https://docs.bithuman.ai/platforms/swift/reference.md, https://docs.bithuman.ai/platforms/android/reference.md --- # Deployment options URL: https://docs.bithuman.ai/deploy > The avatar renders in one of five places. Pick the one that matches where your users are and what may leave their device. Every mode uses the same avatar and the same API secret. What changes is where the avatar renders, where the conversation runs, and what reaches bitHuman. ## Compare the modes | Compare | [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud) | [Your servers](https://docs.bithuman.ai/deploy/self-hosted) | [On the device](https://docs.bithuman.ai/deploy/on-device) | [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu) | [Fully offline](https://docs.bithuman.ai/deploy/offline) | |---|---|---|---|---|---| | **The avatar renders** | on bitHuman's servers, in the US | on your own Mac or Linux machines | on the iPhone, iPad, Mac or Android phone in front of the user, or in a WebGPU browser tab | on a standard Linux PC's CPU, with no GPU | on your Linux PCs and terminals | | **The conversation runs** | on bitHuman's voice service, or with your own provider keys | your choice: the CLI's local conversation brain, your own services, or bitHuman's | your app's choice; with the web embed, on bitHuman's servers | your choice: the CLI's local conversation brain, your own services, or bitHuman's | agreed with sales for your site | | **What reaches bitHuman** | the session's audio and conversation, to run it | a credential check, the avatar download, and usage reports with no audio, video or text | with your own voice and language services, usage metering only, never audio, video or conversation text | a credential check, the avatar download, and usage reports with no audio, video or text | usage is metered on the machine; no reconnection is required | | **Network** | required for the whole session | to start; rendering continues through a drop of up to 5 minutes | to start; rendering continues through a drop of up to 5 minutes | to start; rendering continues through a drop of up to 5 minutes | off the internet; creating the avatar happens online first | | **Price** | 4 credits per minute (Essence 2, Expression 2) | 2 credits per minute (Essence 2, Expression 2) | 2 credits per minute (Essence 2, Expression 2) | 2 credits per minute (Essence 2, Expression 2) | from 100,000 credits; arranged through sales | | **Products** | web embed, REST API, LiveKit plugin, cloud avatar in your room | CLI, Python SDK, LiveKit plugin (`model_path`) | Swift package, Android SDK, Flutter plugin, web embed with `render=local` | CLI, Python SDK on Linux x86_64 or arm64 | Linux PCs and terminals | | **Models** | Essence 2, Expression 2, Essence 1, Expression 1 | Essence 2, Expression 2, Essence 1 | Essence 2, Expression 2 | Essence 2, Expression 2, Essence 1 | Essence 1 (Linux x86_64 and ARM64); Essence 2 and Expression 2 later | | **Plan** | Creator plan or higher | Creator plan or higher | Creator plan or higher | Creator plan or higher | Business & Enterprise | Every mode bills active session time, talking or idle, to the second ([pricing](https://docs.bithuman.ai/pricing)). ## bitHuman cloud bitHuman renders the avatar in the US and streams it to your page, app or LiveKit room. [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud) ## Your servers (self-hosted) The CLI, Python or the LiveKit plugin on your own Mac or Linux machines. When the avatar renders on your hardware, its audio and video stay there. [Your servers](https://docs.bithuman.ai/deploy/self-hosted) ## On the device Essence 2 and Expression 2 render on iPhone, iPad, Mac and Android, or in a WebGPU browser tab. [On the device](https://docs.bithuman.ai/deploy/on-device) ## CPU only (no GPU) Both models run live on a standard Linux PC with no GPU. [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu) ## Fully offline Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. [Fully offline](https://docs.bithuman.ai/deploy/offline) ## Choosing a mode - **Building for the web, or want the fewest moving parts?** The [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud), starting with the [web embed](https://docs.bithuman.ai/platforms/web). - **Must audio and video stay on your network?** [Your servers](https://docs.bithuman.ai/deploy/self-hosted) or [on the device](https://docs.bithuman.ai/deploy/on-device), with the [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) or your own models. What reaches bitHuman in each: [Data flows & privacy](https://docs.bithuman.ai/deploy/privacy). - **Want the lowest per-minute rate?** Render on the device or on your servers ([rates](https://docs.bithuman.ai/pricing)). - **No reliable internet at the site?** [Fully offline](https://docs.bithuman.ai/deploy/offline), on the Business and Enterprise plans. - **Planning for a branch, a clinic or a show floor?** Start from [Use cases](https://docs.bithuman.ai/deploy/use-cases). ## Try it Talk to a sample avatar streamed from the bitHuman cloud: the same web embed a customer puts on a site. --- # Data flows & privacy URL: https://docs.bithuman.ai/deploy/privacy > What reaches bitHuman in each deployment mode: where the avatar renders, where the conversation runs, what usage reports carry, and what bitHuman stores. Two things decide what leaves your hardware: where the avatar **renders**, and where the conversation **runs**. This page answers both for each mode. ## By mode | Question | bitHuman cloud | Your servers | On the device | |---|---|---|---| | **The avatar renders** | on bitHuman's servers, in the US | on your Mac or Linux machines; its audio and video stay there | in your app on the device, or in the browser tab (`render=local`) | | **The conversation runs** | on bitHuman's voice service, or with the providers whose keys you connect | where you choose: the CLI's local conversation brain, your own services, or bitHuman's | where your app chooses; with the web embed, on bitHuman's servers | | **Usage reports** | none: bitHuman runs the session | usage only, no audio, video, images or conversation text | the same as your servers | | **Transcripts at bitHuman** | stored with the agent | none stored | none stored | | **Your API secret** | on your server; browsers get scoped embed tokens | on your machine | held by your app on the device | [Fully offline](https://docs.bithuman.ai/deploy/offline) runs on Linux PCs and terminals with no required reconnection; it is arranged through sales. ## What leaves your hardware Pick a mode to see where each kind of data goes. Rows marked **Reaches bitHuman** are what crosses from your hardware to bitHuman. | Data | bitHuman cloud | Your servers | On the device | CPU only (no GPU) | Fully offline | Web embed, render=local | |---|---|---|---|---|---|---| | Portrait | Reaches bitHuman: uploaded once to the bitHuman cloud, where the avatar is created | Reaches bitHuman: uploaded once to the bitHuman cloud, where the avatar is created | Reaches bitHuman: uploaded once to the bitHuman cloud, where the avatar is created | Reaches bitHuman: uploaded once to the bitHuman cloud, where the avatar is created | Reaches bitHuman: uploaded once to the bitHuman cloud, where the avatar is created | Reaches bitHuman: uploaded once to the bitHuman cloud, where the avatar is created | | Avatar model file | At bitHuman: stays in the bitHuman cloud, where the avatar renders, in the US | Comes from bitHuman: downloads once, then renders on your machine | Comes from bitHuman: downloads once; on Android, usage reporting is then the only traffic | Comes from bitHuman: downloads once, then renders on your machine | Stays on your hardware: runs on the machine; creating the avatar happens online first | Comes from bitHuman: the avatar's web bundle downloads to the browser | | Live audio | Reaches bitHuman: reaches bitHuman to run the session, encrypted in transit | Your choice of services: goes where the conversation runs: nowhere with the CLI's local conversation brain, or to your services or bitHuman's | Your choice of services: goes to the voice and language services your app uses; with your own, bitHuman receives none | Your choice of services: goes where the conversation runs: nowhere with the CLI's local conversation brain, or to your services or bitHuman's | Stays on your hardware: stays on the machine, off the internet | Reaches bitHuman: reaches bitHuman: the conversation runs on bitHuman's servers | | Avatar video | Comes from bitHuman: rendered by bitHuman and streamed to your viewer | Stays on your hardware: stays on your machine | Stays on your hardware: renders on the device; bitHuman never receives it | Stays on your hardware: renders on the PC's CPU and stays there | Stays on your hardware: stays on the machine | Stays on your hardware: renders in the visitor's tab with WebGPU | | Transcripts | At bitHuman: kept with the agent; deleting the agent deletes them | Stays on your hardware: none stored at bitHuman | Stays on your hardware: none stored at bitHuman | Stays on your hardware: none stored at bitHuman | Stays on your hardware: stay on the machine | At bitHuman: the conversation runs on bitHuman's servers | | Knowledge and persona | At bitHuman: your persona lives with the managed agent; deleting the agent deletes its records | Your choice of services: where the conversation runs; any OpenAI-compatible model works, including one in your own network | Your choice of services: stays with your app and the services it uses | Your choice of services: where the conversation runs; any OpenAI-compatible model works, including one in your own network | Stays on your hardware: agreed with sales for your site | At bitHuman: your persona lives with the managed agent | | Provider keys | Reaches bitHuman: encrypted at rest when you connect them | Your choice of services: stay with the services you run | Your choice of services: stay with your app and the services it uses | Your choice of services: stay with the services you run | Stays on your hardware: agreed with sales for your site | Reaches bitHuman: encrypted at rest when you connect them | | API secret | Stays on your hardware: on your server; browsers get scoped embed tokens | Stays on your hardware: on your machine; checked when a session starts | Stays on your hardware: held by your app; checked when a session starts | Stays on your hardware: on your machine; checked when a session starts | Stays on your hardware: arranged through sales, for Business & Enterprise | Stays on your hardware: never in the browser: the embed uses scoped tokens | | Usage reports | At bitHuman: bitHuman meters the session it runs: active session time, talking or idle | Reaches bitHuman: reach bitHuman, with no audio, video, images or conversation text | Reaches bitHuman: usage metering only, never audio, video or conversation text | Reaches bitHuman: reach bitHuman, with no audio, video, images or conversation text | Stays on your hardware: metered on the machine; no required reconnection | At bitHuman: bitHuman meters the session: active session time, talking or idle | ## Rendering on your hardware When the avatar renders on your hardware, its audio and video stay there. When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. On Android, after the one-time model download, the only network traffic is usage reporting. With the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab (`render=local`). ## The conversation - **The local conversation brain:** with the CLI's [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) (`BITHUMAN_LOCAL=1`), speech recognition, the language model and the voice run on the machine; audio, transcripts and generated speech never leave it. The session still reports usage online. - **Your own model:** a managed agent works with any OpenAI-compatible language model endpoint, including one in your own network ([Providers](https://docs.bithuman.ai/api/providers)). - **Your own stack:** the Swift and Android SDKs take any 16 kHz mono speech your pipeline produces and return frames. ## Usage reports Self-hosted and on-device sessions check your credential when they start and report usage while they run. Usage reports contain no audio, video, images or conversation text. Self-hosted and on-device sessions store no transcript at bitHuman. ## Creating an avatar Avatar creation from a portrait happens in the bitHuman cloud; the finished avatar model then runs on your devices. The portrait is uploaded for that step, in every mode. ## Security and access - **In transit:** encrypted with HTTPS/TLS; WebRTC media uses DTLS-SRTP. Provider keys you connect are encrypted at rest. - **Access:** organization roles (owner, admin, member), an audit-log API, API-secret rotation with immediate revocation, and scoped runtime and embed tokens so browsers never hold your secret ([Organizations](https://docs.bithuman.ai/api/organizations), [API secrets](https://docs.bithuman.ai/api/api-keys)). - **Region:** avatars in the bitHuman cloud render in the US. ## Retention and deletion Deleting an agent deletes its records, including transcripts, and its model files ([Agents](https://docs.bithuman.ai/api/agents)). ## Your content bitHuman's [privacy policy](https://www.bithuman.ai/legal/privacy) says: "We do not use your content to train our AI models unless you explicitly opt in." and "We do not sell your personal information." For the EU AI Act, see [our reading of Article 50](https://docs.bithuman.ai/legal/eu-ai-act). --- # Pricing and credits URL: https://docs.bithuman.ai/pricing > Credits pay for active session time, talking or idle, by the exact second. Rates per model and platform, creation costs, plans, offline licensing, and how to check your balance. Credits pay for the time an avatar session is running, talking or idle, billed by the exact second. Every platform (cloud, self-hosted and on-device) bills the same way, against your [API secret](https://docs.bithuman.ai/start/api-secret). This page is the one source for every price; other pages link here. ## Serving — credits per live minute The table and the billing rule under it are generated from [`GET /v1/pricing`](https://docs.bithuman.ai/api/billing#get-the-pricing-schedule) (`data.realtime`). | Model | Cloud | Self-hosted and on-device | |---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) (`essence-2`) | 4 credits/min | 2 credits/min | | [Expression 2](https://docs.bithuman.ai/models/expression-2) (`expression-2`) | 4 credits/min | 2 credits/min | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) (`essence-1`) | 2 credits/min | 1 credit/min | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) (`expression-1`) | 4 credits/min | — | A managed conversational agent bills ONE all-inclusive rate: it covers the avatar, whether it renders in the bitHuman cloud or in the viewer's browser. | Surface | Rate | |---|---| | Managed agent — voice chat (all-inclusive) | 10 credits/min | | Managed agent — camera on (vision chat; replaces the chat rate) | 30 credits/min | An avatar-only session (your own agent through the plugin or the API) that renders in the viewer's browser bills the model's **self-hosted** rate above. Inside a managed agent's chat the all-inclusive rate covers it. How live sessions are billed: active session time, talking or idle: exact seconds x rate / 60, rounded down per session with the remainder carried to your next session; no minimum. A session bills while it is **running**, whether the avatar is talking or idle, to the exact second: seconds × rate ÷ 60, rounded down per session, with the fraction carried to your next session. A stopped or disconnected session accrues nothing. File rendering (`bithuman render`, or `render()` in the Python SDK) bills the duration of the video it writes, at the self-hosted rate. Expression 1 (`expression-1`) runs in the bitHuman cloud only. *Diagram: A session and what it bills.* A realtime session bills active session time, talking or idle, to the second, from the moment it starts until it ends. Online self-hosted and on-device sessions check your credential when they start and keep rendering through a network drop of up to 5 minutes. ## Creation — one-time credits | Action | Credits | |---|---| | Create an agent: `essence-1` or `expression-1` | 250 | | Create an agent: `essence-2` | 500 | | Create an agent: `expression-2` | 2000 | | Create an agent: `auto` | the routed model's rate (500 or 2000) | | [Add a model](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent) to an agent | the same per-model rates; 0 for Expression 1 | | Generate gestures (dynamics) | 250 | A failed creation is refunded automatically. [`GET /v1/pricing`](https://docs.bithuman.ai/api/billing#get-the-pricing-schedule) returns this schedule as JSON. ## Talking video — per minute of output [Talking-video renders](https://docs.bithuman.ai/api/video) bill per minute of finished output, rounded up, minimum one minute. A failed render is refunded. | Model | Per minute of output | |---|---| | `essence-2` | 4 credits | | `expression-2` | 4 credits | | `essence-1` | 2 credits | | `expression-1` | 4 credits | ## Plans From **2026-10-12** (00:00 UTC), API and SDK use requires the Creator plan or higher. Free accounts cannot create agents or buy credit top-ups. A Free account with top-up credits bought before 2026-09-27 keeps API and SDK access until those credits are spent. The exact responses are under [`PLAN_REQUIRED`](https://docs.bithuman.ai/api/errors#authentication). | Plan | Monthly | Yearly | Credits / month | Agents | Concurrent cloud sessions | |---|---|---|---|---|---| | **Creator** | $20 | $204 | 1,800 | 7 | 3 | | **Pro** | $99 | $1,010 | 10,000 | 40 | 10 | | **Business** | $299 | $2,990 | 50,000 | 200 | 50 | | **Enterprise** | $999 | $9,990 | 250,000 | unlimited | 200 | | **Custom** | [Contact sales](https://www.bithuman.ai/sales) | — | by agreement | by agreement | by agreement | Annual plans bill twelve months of credits up front. - **Agents:** a creation over your plan's limit returns `403 AGENT_LIMIT_REACHED`. Existing agents keep working. - **Concurrent sessions** limit live cloud sessions; a session over the limit is refused with `403 CONCURRENCY_LIMIT_REACHED` ([rate limits](https://docs.bithuman.ai/api/rate-limits)). Self-hosted and on-device sessions are limited only by credits. - **Creation costs credits:** a creation you cannot pay for returns [`402 INSUFFICIENT_BALANCE`](https://docs.bithuman.ai/api/errors) and creates nothing. Essence 2 Max is available on the Enterprise plan only. [Contact sales](https://www.bithuman.ai/sales) to enable it. ## Estimate a month Choose where the avatar renders, then how long sessions run. The estimate counts active session time, talking or idle, at the rates above. One avatar session running 60 minutes a day for 30 days, billed as active session time, talking or idle: | Mode | Rate | Credits a month | At the top-up rate | Smallest plan that covers it | |---|---|---|---|---| | On the device or your servers | 2 credits per minute | 3,600 | $36.00 | Pro | | bitHuman cloud avatar | 4 credits per minute | 7,200 | $72.00 | Pro | | Managed voice chat, all-inclusive | 10 credits per minute | 18,000 | $180.00 | Business | ## Budget an app What an Essence 2 or Expression 2 avatar costs inside an iPhone, Android, Mac or web app, per minute of active session time. Creating your own avatar is a one-time cost ([Creation](#creation--one-time-credits)); the [companion app](https://docs.bithuman.ai/build/companion-app) recipe shows how to close the avatar when the app leaves the screen. | Mode | On the device or your servers | bitHuman cloud avatar | Managed voice chat, all-inclusive | |---|---|---|---| | Credits per minute of active session time | 2 | 4 | 10 | | At the top-up rate ($1 = 100 credits) | about $0.02 a minute | about $0.04 a minute | about $0.10 a minute | | At the Business plan's credit price | about $0.012 a minute | about $0.024 a minute | about $0.060 a minute | | Minutes in a Creator month (1,800 credits) | 900 | 450 | 180 | | Speech, language model and voice | your own services | your own services | included | | Sessions at once | limited by credits | [per plan](https://docs.bithuman.ai/pricing#plans) | [per plan](https://docs.bithuman.ai/pricing#plans) | - **Idle is billed.** A session bills for as long as it runs, talking or idle. End it when the user leaves, and show a still frame when nobody is talking. - **Worked example.** One user whose sessions add up to 20 minutes a day on the device uses about 1,200 credits a month ($12.00 at the top-up rate). ## Offline licensing Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. - **Models:** Essence 1 on Linux (x86_64 and ARM64), available now: buy a pack in the console, then run `python -m bithuman pack redeem` (bitHuman 2.11.16 or later) once on the machine ([how](https://docs.bithuman.ai/deploy/offline#first-command)). Essence 2 and Expression 2 offline come later. - **Credit-based:** from 100,000 credits, metered on the machine at the self-hosted rate, with no required reconnection. - **Creation is online:** you create the avatar from a portrait in the bitHuman cloud; the finished avatar model then runs on your machines. - **Not for phones:** the Swift package and the Android SDK stay online. - **Not file rendering:** `bithuman render` writes a video file and signs in online; it needs no offline license. [Contact sales](https://www.bithuman.ai/sales) to arrange an offline license. Where it runs and what it covers: [Fully offline](https://docs.bithuman.ai/deploy/offline). ## Top-up credits On the Creator plan or higher, top up any time at **$1 = 100 credits**. Top-up credits never expire and are spent after plan credits. ## Connectivity | Situation | What happens | |---|---| | No API secret, or a rejected one, at the start | the session does not start | | No network when a session starts | the session does not start; retry when connected | | The network drops after the session started | the session continues for 5 minutes, then pauses until the connection returns; usage is reported when it does | | Credits run out | the session stops at the next usage report; top up to continue | ## Check your balance ```bash curl https://api.bithuman.ai/v2/credit-summaries -H "api-secret: $BITHUMAN_API_SECRET" ``` ```json { "success": true, "data": { "user_id": "00000000-0000-0000-0000-000000000000", "balance": 5240, "plan_credits": 240, "topup_credits": 5000, "is_enterprise": false, "minutes_estimate": { "essence_2_cloud": 1310, "essence_2_self_hosted": 2620, "expression_2_cloud": 1310, "expression_2_self_hosted": 2620, "essence_1_cloud": 2620, "essence_1_self_hosted": 5240, "expression_1_cloud": 1310, "voice_chat": 524, "camera_chat": 174, "essence_cloud": 2620, "essence_self_hosted": 5240, "expression_cloud": 1310 } } } ``` Each `_cloud` and `_self_hosted` value is the balance divided by that rate. Ignore `expression_1_self_hosted` and `expression_self_hosted`: Expression 1 has no self-hosted mode. The unversioned `essence_*` and `expression_*` keys are the first-generation models; for Essence 2 read `essence_2_*`. ## What is not billed - Stopped or disconnected sessions. - API secrets, SDK installs and model downloads. - Failed creations and renders (refunded) and failed authentication. ## Next - [Billing API](https://docs.bithuman.ai/api/billing) · [Rate limits](https://docs.bithuman.ai/api/rate-limits) · [Create your own avatar](https://docs.bithuman.ai/build/create-avatar) --- # bitHuman cloud URL: https://docs.bithuman.ai/deploy/cloud > bitHuman renders the avatar on its servers and streams it to your page, app or LiveKit room: the web embed, the REST API or the LiveKit plugin. ## What it is The fewest moving parts: bitHuman renders the avatar in the cloud and streams its video and audio to a web page, an app or a LiveKit room. You provision no GPU and install nothing. The live demo above is this mode. ## Where it renders | Question | bitHuman cloud | |---|---| | Where the avatar renders | on bitHuman's servers, in the US | | Where the conversation runs | on bitHuman's voice service, or with your own provider keys | | What reaches bitHuman | the session's audio and conversation, to run it | | Network | required for the whole session | *Diagram: bitHuman cloud.* In the bitHuman cloud the avatar renders on bitHuman's servers, in the US, and the conversation runs on bitHuman's voice service or with the provider keys you connect. The browser or app sends the microphone and shows the video. Your API secret stays on your server; browsers get scoped embed tokens. A managed agent's conversation runs on bitHuman's voice service with your persona, or with the voice and language providers whose keys you connect ([Voices](https://docs.bithuman.ai/build/voices), [Providers](https://docs.bithuman.ai/api/providers)). Traffic is encrypted in transit: HTTPS, and WebRTC media over DTLS-SRTP. ## Models available here | Model | bitHuman cloud | How | |---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) | Yes | web embed, REST API, LiveKit | | [Expression 2](https://docs.bithuman.ai/models/expression-2) | Yes | web embed, REST API, LiveKit | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) | Yes | web embed, REST API, LiveKit | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) | Yes | web embed, REST API, LiveKit | ## Speed The bitHuman cloud serves from several kinds of hardware; each is measured: | Configuration | Hardware | Essence 2 | Expression 2 | Measured | |---|---|---|---|---| | Cloud API · GPU | NVIDIA RTX 4090 | 4.1× real time | 17.0× real time | cloud API, 2026-09-27, 2026-09-23 | | Cloud API · Apple silicon | Apple M4 Max | 2.8× real time | 5.5× real time | cloud API, 2026-09-26, 2026-09-23 | | Cloud API · CPU (CPU only (no GPU)) | x86 server CPU | 1.1× real time | 1.3× real time | cloud API, 2026-09-25, 2026-09-24 | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance#cloud)). ## Price 4 credits per minute of active session time for Essence 2 and Expression 2, about $0.04 a minute at the top-up rate of $1 = 100 credits. A managed agent's voice chat bills 10 credits per minute, all-inclusive: the avatar is part of it. Realtime usage bills active session time, talking or idle, to the second. Every rate: [Pricing and credits](https://docs.bithuman.ai/pricing). ## Limits bitHuman cloud sessions are limited per plan: Creator 3, Pro 10, Business 50, Enterprise 200 concurrent sessions. On-device and self-hosted sessions are limited by credits ([plans](https://docs.bithuman.ai/pricing#plans)). A session over your plan's limit is refused with `403 CONCURRENCY_LIMIT_REACHED` ([Rate limits](https://docs.bithuman.ai/api/rate-limits)). The network is needed for the whole session. ## First command ```html tab="Web embed" ``` ```bash tab="REST API" curl -s -X POST https://api.bithuman.ai/v1/validate -H "api-secret: $BITHUMAN_API_SECRET" # → {"valid":true} ``` ```bash tab="LiveKit" pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv # then pass avatar_id= to bithuman.AvatarSession: /platforms/livekit ``` Next steps: [Web](https://docs.bithuman.ai/platforms/web), [REST API](https://docs.bithuman.ai/platforms/rest), [LiveKit](https://docs.bithuman.ai/platforms/livekit), or [a cloud avatar in your own room without the plugin](https://docs.bithuman.ai/api/cloud-avatar). ## Choosing between modes - **Keep audio and video on your own machines:** [Your servers](https://docs.bithuman.ai/deploy/self-hosted). - **Render inside your app on the phone, Mac or browser:** [On the device](https://docs.bithuman.ai/deploy/on-device). - **A Linux PC with no GPU:** [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu). - **No internet at the site:** [Fully offline](https://docs.bithuman.ai/deploy/offline). - **All five side by side:** [Deployment options](https://docs.bithuman.ai/deploy). --- # Your servers (self-hosted) URL: https://docs.bithuman.ai/deploy/self-hosted > Run the CLI, the Python SDK or the LiveKit plugin on your own Mac or Linux machines. When the avatar renders on your hardware, its audio and video stay there. ## What it is The avatar renders on machines you run: a Mac with Apple silicon, or a Linux PC or server on x86_64 or arm64, including one with no GPU. You choose where the conversation runs. There is no license to buy for online self-hosting: it needs the Creator plan or higher and bills credits at the self-hosted rate. | You want | Use | Models | |---|---|---| | A talking avatar or an MP4, no code | [CLI](https://docs.bithuman.ai/platforms/cli) | Essence 2 and Expression 2 (`run`, `render`); Essence 1 (`run`) | | Frames or MP4 clips from your own code | [Python SDK](https://docs.bithuman.ai/platforms/python) | Essence 2, Expression 2, Essence 1 | | A voice agent in your own LiveKit rooms, rendered on your machine | [LiveKit plugin](https://docs.bithuman.ai/platforms/livekit) with `model_path=` ([guide](https://docs.bithuman.ai/build/voice-agent)) | Essence 2, Expression 2 | ## Where it renders | Question | Your servers | |---|---| | Where the avatar renders | on your own Mac or Linux machines | | Where the conversation runs | your choice: the CLI's local conversation brain, your own services, or bitHuman's | | What reaches bitHuman | a credential check, the avatar download, and usage reports with no audio, video or text | | Network | to start; rendering continues through a drop of up to 5 minutes | *Diagram: Your servers (self-hosted).* Self-hosted, the avatar renders on your own Mac or Linux machine and its audio and video stay there. The conversation runs where you choose: the CLI's local conversation brain, your own services, or bitHuman's. For the rendering, bitHuman receives a credential check, the avatar download and usage reports with no audio, video, images or conversation text. When the avatar renders on your hardware, its audio and video stay there. For the conversation, the CLI's [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) keeps speech recognition, the language model and the voice on the machine, or you bring any OpenAI-compatible language model, including one in your own network ([Providers](https://docs.bithuman.ai/api/providers)). Self-hosted sessions store no transcript at bitHuman. ## Models available here | Model | Your servers | How | |---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) | Yes | CLI, Python, LiveKit plugin | | [Expression 2](https://docs.bithuman.ai/models/expression-2) | Yes | CLI, Python, LiveKit plugin | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) | Yes | CLI (`run`), Python | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) | — | | ## Speed | Configuration | Hardware | Essence 2 | Expression 2 | Measured | |---|---|---|---|---| | Linux · CLI (CPU only (no GPU)) | Intel Core i7-13700F (x86_64) | 2.0× real time | 2.2× real time | CLI 2.8.1, 2026-09-27 | | Linux · Python (CPU only (no GPU)) | Intel Core i7-13700F (x86_64) | 1.9× real time | 2.3× real time | bithuman 2.11.13, 2026-09-26 | | macOS · CLI | Apple M4 | 4.2× real time | 8.4× real time | CLI 2.8.1, 2026-09-27 | | macOS · Python | Apple M4 | 6.9× real time | 8.4× real time | bithuman 2.11.12, 2026-09-25 | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance#desktop)). ## Price 2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: [Pricing and credits](https://docs.bithuman.ai/pricing). Rendering an MP4 (`bithuman render`, or `render()` in Python) bills the length of the video it writes, at the same rate. ## Limits - **Credential:** rendering needs a credential. Sign in with `bithuman login`, or set `BITHUMAN_API_SECRET` ([Your API secret](https://docs.bithuman.ai/start/api-secret)). - **Network:** a session checks your credential when it starts and keeps rendering through a network drop of up to 5 minutes. Usage reports carry no audio, video, images or conversation text. - **Sessions:** self-hosted sessions are limited by credits. - **Operating systems:** macOS on Apple silicon; Linux on x86_64 or arm64. On Windows, use WSL2. - **Off the internet:** see [Fully offline](https://docs.bithuman.ai/deploy/offline). ## First command ```bash tab="CLI" curl -fsSL https://install.bithuman.ai | sh bithuman login # in CI, export BITHUMAN_API_SECRET instead curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav bithuman render wise-pup speech.wav -o out.mp4 ``` ```bash tab="Python" python3 -m venv .venv && source .venv/bin/activate pip install "bithuman[expression-2]" export BITHUMAN_API_SECRET="" curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2" ``` ```bash tab="LiveKit" pip install "livekit-agents[openai,silero]" livekit-plugins-bithuman python-dotenv # pass model_path="wise-pup.imx" to bithuman.AvatarSession: /build/voice-agent ``` `out.mp4` is the `wise-pup` sample avatar speaking the 15-second sample. The CLI's `ffmpeg` and live-session setup is on [CLI](https://docs.bithuman.ai/platforms/cli#before-you-start). ## Choosing between modes - **A Linux PC with no GPU:** [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu). - **Inside an app on the phone, Mac or browser:** [On the device](https://docs.bithuman.ai/deploy/on-device). - **Nothing to run yourself:** [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud). - **No internet at the site:** [Fully offline](https://docs.bithuman.ai/deploy/offline). - **All five side by side:** [Deployment options](https://docs.bithuman.ai/deploy). --- # On the device URL: https://docs.bithuman.ai/deploy/on-device > Essence 2 and Expression 2 render on the device in front of the user: iPhone, iPad and Mac, Android phones, or a WebGPU browser tab. ## What it is The avatar renders inside your app on the device in front of the user: iPhone, iPad and Mac with the [Swift package](https://docs.bithuman.ai/platforms/ios), Android phones with the [Android SDK](https://docs.bithuman.ai/platforms/android) or the [Flutter plugin](https://docs.bithuman.ai/platforms/flutter), or the visitor's browser tab with [WebGPU](https://docs.bithuman.ai/platforms/web). There is no render server to run. The mobile SDKs take any 16 kHz mono speech your pipeline produces and return frames, so any speech-recognition, language-model and voice stack works. ## Where it renders | Question | On the device | |---|---| | Where the avatar renders | on the iPhone, iPad, Mac or Android phone in front of the user, or in a WebGPU browser tab | | Where the conversation runs | your app's choice; with the web embed, on bitHuman's servers | | What reaches bitHuman | with your own voice and language services, usage metering only, never audio, video or conversation text | | Network | to start; rendering continues through a drop of up to 5 minutes | *Diagram: On the device.* On the device, your app renders the avatar on the iPhone, iPad, Mac or Android phone and brings its own voice and language services. The avatar model downloads once. When you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. - **In your app:** when the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. - **On Android:** after the one-time model download, the only network traffic is usage reporting. - **In the browser:** with the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab (`render=local`). ## Models available here | Model | [iPhone and iPad](https://docs.bithuman.ai/platforms/ios) | [Mac](https://docs.bithuman.ai/platforms/macos) | [Android](https://docs.bithuman.ai/platforms/android) | [Browser (WebGPU)](https://docs.bithuman.ai/platforms/web) | |---|---|---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) | Yes | Yes | Yes | Yes | | [Expression 2](https://docs.bithuman.ai/models/expression-2) | Yes | Yes | Yes | Yes | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) | — | Yes | — | Yes | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) | — | — | — | — | Essence 1 and Expression 1 are not available on phones or in the Swift package. ## Speed Measured on the device, including 10-minute held runs: | Configuration | Hardware | Essence 2 | Expression 2 | Measured | |---|---|---|---|---| | iPhone · Swift package | iPhone 15 | 2.1× real time | 5.5× real time | Swift package 2.17.3 and Swift package 2.18.0, 2026-09-27 | | iPhone · Swift package (held 10 min) | iPhone 15 | 1.3× real time | 5.1× real time | Swift package 2.15.0, 2026-09-25 | | Android | Samsung Galaxy S25+ | 2.0× real time | 2.4× real time | essence2-android 0.7.0 and expression2-android 0.4.10, 2026-09-25, 2026-09-23 | | Android (held 10 min) | Samsung Galaxy S25+ | 1.4× real time | 2.2× real time | essence2-android 0.8.1 and expression2-android 0.5.2, 2026-09-27 | | Web browser (WebGPU) | Chrome on Apple M4 | 1.7× real time | 1.9× real time | web viewer, 2026-09-27 | | macOS · Swift package | Apple M4 | 4.8× real time | 8.8× real time | Swift package 2.15.0, 2026-09-24 | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance#mobile)). ## Price 2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: [Pricing and credits](https://docs.bithuman.ai/pricing). ## Limits - **Network:** a session checks your credential when it starts and keeps rendering through a network drop of up to 5 minutes. - **Devices:** iPhone, iPad and Android need a physical device, not a simulator or an emulator. Essence 2 on Apple needs iOS 26 or macOS 26. - **Sessions:** on-device sessions are limited by credits, not by a session cap. - **First run:** each avatar downloads once (about 160–370 MB, by model and platform), then stays on the device. ## First command ```swift tab="iOS & iPadOS" // Package.swift (or Xcode → Add Package Dependencies) .package(url: "https://github.com/bithuman-product/homebrew-bithuman.git", from: "2.19.0") ``` ```kotlin tab="Android" // app/build.gradle.kts implementation("ai.bithuman:expression2-android:0.5.2") ``` ```bash tab="Mac" git clone https://github.com/bithuman-product/bithuman-examples.git cd bithuman-examples/swift/macos-expression2 && ./setup.sh BITHUMAN_API_SECRET="" swift run -c release MacOSExpression2 ``` ```html tab="Web" ``` The whole first frame for each: [iOS & iPadOS](https://docs.bithuman.ai/platforms/ios#first-frame) · [Android](https://docs.bithuman.ai/platforms/android#first-frame) · [macOS](https://docs.bithuman.ai/platforms/macos#first-frame) · [Web](https://docs.bithuman.ai/platforms/web#render-in-the-visitors-tab-webgpu). ## Choosing between modes - **The fewest moving parts, any device:** [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud). - **Your own Mac or Linux machines:** [Your servers](https://docs.bithuman.ai/deploy/self-hosted). - **A Linux PC with no GPU:** [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu). - **No internet at the site:** [Fully offline](https://docs.bithuman.ai/deploy/offline); not for phones, Mac or the browser. - **All five side by side:** [Deployment options](https://docs.bithuman.ai/deploy). --- # CPU only (no GPU) URL: https://docs.bithuman.ai/deploy/cpu > Both models run live on a standard Linux PC with no GPU. ## What it is Essence 2 and Expression 2 render live on the processor of an ordinary Linux PC, with no graphics card: with the [CLI](https://docs.bithuman.ai/platforms/cli) or the [Python SDK](https://docs.bithuman.ai/platforms/python), on Linux x86_64 or arm64. It suits screens that run all day where a GPU is not practical: kiosks, lobby screens, and servers without GPUs. ## Where it renders | Question | CPU only (no GPU) | |---|---| | Where the avatar renders | on a standard Linux PC's CPU, with no GPU | | Where the conversation runs | your choice: the CLI's local conversation brain, your own services, or bitHuman's | | What reaches bitHuman | a credential check, the avatar download, and usage reports with no audio, video or text | | Network | to start; rendering continues through a drop of up to 5 minutes | *Diagram: CPU only (no GPU).* On a standard Linux PC with no GPU, both models render live on the CPU, and the avatar's audio and video stay on the PC. The conversation runs where you choose: the CLI's local conversation brain, your own services, or bitHuman's. For the rendering, bitHuman receives a credential check, the avatar download and usage reports with no audio, video, images or conversation text. When the avatar renders on your hardware, its audio and video stay there. With the CLI's [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain), speech recognition, the language model and the voice run on the machine too; the session still reports usage online. ## Models available here | Model | Linux, no GPU | How | |---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) | Yes | CLI, Python | | [Expression 2](https://docs.bithuman.ai/models/expression-2) | Yes | CLI, Python | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) | Yes | CLI (`run`), Python | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) | — | | ## Speed Measured on a desktop CPU with no GPU: | Configuration | Hardware | Essence 2 | Expression 2 | Measured | |---|---|---|---|---| | Linux · CLI (CPU only (no GPU)) | Intel Core i7-13700F (x86_64) | 2.0× real time | 2.2× real time | CLI 2.8.1, 2026-09-27 | | Linux · Python (CPU only (no GPU)) | Intel Core i7-13700F (x86_64) | 1.9× real time | 2.3× real time | bithuman 2.11.13, 2026-09-26 | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance#desktop)). ## Price 2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: [Pricing and credits](https://docs.bithuman.ai/pricing). ## Limits - **Network:** a session checks your credential when it starts and keeps rendering through a network drop of up to 5 minutes. Usage reports carry no audio, video, images or conversation text. - **Operating system:** Linux on x86_64 or arm64. On Windows, use WSL2 or the [web embed](https://docs.bithuman.ai/platforms/web); Intel Macs are not supported. - **Sessions:** self-hosted sessions are limited by credits. ## First command ```bash tab="CLI" curl -fsSL https://install.bithuman.ai | sh bithuman login curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav bithuman render wise-pup speech.wav -o out.mp4 ``` ```bash tab="Python" python3 -m venv .venv && source .venv/bin/activate pip install "bithuman[expression-2]" export BITHUMAN_API_SECRET="" curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2" curl -fsSLo speech.wav https://docs.bithuman.ai/samples/speech.wav python -c 'import bithuman with bithuman.open("wise-pup.imx") as a: print(sum(1 for _ in a.render("speech.wav")), "frames")' # → 300 frames ``` `bithuman run wise-pup` opens a live conversation instead of a file ([CLI](https://docs.bithuman.ai/platforms/cli#first-frame)). ## Choosing between modes - **Off the internet, on Linux PCs and terminals:** [Fully offline](https://docs.bithuman.ai/deploy/offline), for Business and Enterprise. - **On your own Macs, or Linux machines you already run:** [Your servers](https://docs.bithuman.ai/deploy/self-hosted). - **Inside an app on the phone or in the browser:** [On the device](https://docs.bithuman.ai/deploy/on-device). - **Nothing to run yourself:** [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud). - **All five side by side:** [Deployment options](https://docs.bithuman.ai/deploy). --- # Fully offline URL: https://docs.bithuman.ai/deploy/offline > Realtime avatars that run completely locally, off the internet, on Linux PCs and terminals: for Business and Enterprise clients, arranged through sales. ## What it is Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. - **Plans:** Business & Enterprise. - **Billing:** credit-based, from 100,000 credits, metered on the machine. - **Connectivity:** no required reconnection. [Contact sales](https://www.bithuman.ai/sales) to arrange an offline license. ## Where it renders | Question | Fully offline | |---|---| | Where the avatar renders | on your Linux PCs and terminals | | Where the conversation runs | agreed with sales for your site | | What reaches bitHuman | usage is metered on the machine; no reconnection is required | | Network | off the internet; creating the avatar happens online first | *Diagram: Fully offline.* Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. The avatar renders on the machine and usage is metered there, with no required reconnection. Models: Essence 1, Essence 2 and Expression 2. Creating the avatar from a portrait happens in the bitHuman cloud; the finished avatar model then runs on your machines. ## Models available here | Model | Fully offline | How | |---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) | — | Coming later | | [Expression 2](https://docs.bithuman.ai/models/expression-2) | — | Coming later | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) | Yes | Linux x86_64 and ARM64, bitHuman 2.11.16 or later; Business & Enterprise | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) | — | | Essence 1 runs fully offline on Linux (x86_64 and ARM64) today. Essence 2 and Expression 2 offline come later. Expression 1 runs in the bitHuman cloud only. ## Speed An offline machine is a Linux PC. How fast each model renders on a Linux PC with no GPU is on [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu#speed). ## Price From 100,000 credits, credit-based and metered on the machine. Business & Enterprise; arranged through sales. ## Limits - **Linux PCs and terminals** only. Phones, Macs and browsers stay online: the Swift package, the Android SDK and the web embed check your credential when a session starts. - **Creation is online:** you create the avatar from a portrait in the bitHuman cloud before it runs offline. - **Not file rendering:** `bithuman render` and Python's `bithuman.offline` write a video file and sign in online, on any plan that can render. ## First command Available today: **Essence 1 on Linux x86_64 and Linux ARM64**. Essence 2 and Expression 2 come later. 1. **Buy a pack** for the [offline license](https://docs.bithuman.ai/deploy/offline) in the console (**Developer → Offline licenses**), choosing the model and the platform it will run on. A pack is at least 100,000 credits at the self-hosted rate. You can cancel it for a full refund until a machine redeems it. Through the API: `POST /v1/offline/entitlements` with `"platform": "linux-x86_64"` or `"linux-aarch64"`. 2. **Redeem it once, on the machine that will run it**, while it is online, with bitHuman 2.11.16 or later and your account's API secret: ```bash pip install -U "bithuman>=2.11.16" export BITHUMAN_API_SECRET=... python -m bithuman pack redeem ``` This binds the pack to this machine and installs it. If the install step fails, `python -m bithuman pack redeem --file ` retries it from the copy kept in `~/.bithuman/packs/`, with no connection and no second charge. 3. **Run offline.** The machine now renders Essence 1 avatars with no network and no API secret until the pack's credits are spent. Credits are metered on the machine, at the self-hosted rate for active session time. It never has to reconnect. ## Choosing between modes - **A Linux PC with a connection:** [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu), on the Creator plan or higher. - **Your own machines, online:** [Your servers](https://docs.bithuman.ai/deploy/self-hosted). - **Inside an app on the phone or in the browser:** [On the device](https://docs.bithuman.ai/deploy/on-device). - **Nothing to run yourself:** [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud). - **All five side by side:** [Deployment options](https://docs.bithuman.ai/deploy). --- # Use cases URL: https://docs.bithuman.ai/deploy/use-cases > Which deployment mode each use case starts from: banking and ATMs, healthcare, events and trade shows, kiosks, apps and AI companions, and websites. Every use case runs the same avatar. What changes is where it renders, where the conversation runs, and what reaches bitHuman. | Use case | Start from | Guide | |---|---|---| | Branch screens, teller terminals and ATMs | a Linux PC or terminal with no GPU, or an Android terminal; fully offline where there is no internet | [Banking and ATMs](https://docs.bithuman.ai/deploy/use-cases/banking-and-atms) | | Clinics and hospitals | your servers or the device, with the conversation kept in your environment | [Healthcare](https://docs.bithuman.ai/deploy/use-cases/healthcare) | | Booths, show floors and museums | a Linux PC with no GPU and the local conversation brain | [Events and trade shows](https://docs.bithuman.ai/deploy/use-cases/events-and-trade-shows) | | Lobbies, retail and information screens | a Linux PC with no GPU, full screen in Chrome | [Kiosk on a Linux PC](https://docs.bithuman.ai/build/kiosk) | | iPhone, iPad, Mac and Android apps, including AI companions | on the device, with your own voice and language services | [Companion app](https://docs.bithuman.ai/build/companion-app) | | Websites | the bitHuman cloud, through the web embed or a widget | [Website widget](https://docs.bithuman.ai/build/website-widget) | When the avatar renders on your hardware, its audio and video stay there. When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. With the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab. Every mode, side by side: [Data flows & privacy](https://docs.bithuman.ai/deploy/privacy). Healthcare and financial-services deployments are set up under an enterprise agreement and review. [Contact sales](https://www.bithuman.ai/sales) to start one. --- # Banking and ATMs URL: https://docs.bithuman.ai/deploy/use-cases/banking-and-atms > Run an always-on avatar on a branch screen, teller terminal or ATM: the avatar handles the conversation, card handling and transactions stay in your systems. Hardware, data flows, operations and cost. An avatar on a branch screen, a teller terminal or an ATM greets customers, answers questions and walks them through a task. The avatar handles the conversation; card handling and transactions stay in your systems. ## Where it runs | Terminal | How the avatar renders | |---|---| | A Linux PC or terminal, x86_64 or arm64, with no GPU | the CLI or the Python SDK render it on the CPU ([CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu)) | | An Android terminal (arm64) | the Android SDK renders it inside your app ([Android](https://docs.bithuman.ai/platforms/android)) | | Screens fed from your own servers | the CLI, the Python SDK or the LiveKit plugin on your Mac or Linux machines ([Your servers](https://docs.bithuman.ai/deploy/self-hosted)) | | Windows-based terminals | [talk to us](https://www.bithuman.ai/sales) | How fast each model renders on a standard desktop CPU: [Performance](https://docs.bithuman.ai/performance). ## What reaches bitHuman | Question | CPU only (no GPU) | |---|---| | Where the avatar renders | on a standard Linux PC's CPU, with no GPU | | Where the conversation runs | your choice: the CLI's local conversation brain, your own services, or bitHuman's | | What reaches bitHuman | a credential check, the avatar download, and usage reports with no audio, video or text | | Network | to start; rendering continues through a drop of up to 5 minutes | Usage reports contain no audio, video, images or conversation text, and self-hosted sessions store no transcript at bitHuman. Every mode: [Data flows & privacy](https://docs.bithuman.ai/deploy/privacy). ## The conversation - **On the terminal:** the CLI's [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) (`BITHUMAN_LOCAL=1`) runs speech recognition, the language model and the voice on the machine. Audio, transcripts and generated speech never leave it; the session still reports usage online. - **Your own model:** any OpenAI-compatible endpoint works, including one inside your own network ([Providers](https://docs.bithuman.ai/api/providers)). - **Your own systems:** account data and transactions stay in the services you already run. The avatar speaks what your conversation layer gives it. ## No internet at the site A site with no internet runs [Fully offline](https://docs.bithuman.ai/deploy/offline), on the Business and Enterprise plans. The models it covers: | Model | Fully offline | How | |---|---|---| | [Essence 2](https://docs.bithuman.ai/models/essence-2) | — | Coming later | | [Expression 2](https://docs.bithuman.ai/models/expression-2) | — | Coming later | | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) | Yes | Linux x86_64 and ARM64, bitHuman 2.11.16 or later; Business & Enterprise | | [Expression 1](https://docs.bithuman.ai/models/first-generation#expression-1) | — | | Creating the avatar from a portrait happens in the bitHuman cloud; the finished avatar model then runs on your machines. ## Run it all day - **One API secret per terminal**, so you can revoke one without touching the rest ([API secrets](https://www.bithuman.ai/developer/api-keys)). - **Download the avatar before opening:** `bithuman pull $AGENT_CODE`, then start it at boot from a service that restarts it if it exits, and show it full screen ([Kiosk on a Linux PC](https://docs.bithuman.ai/build/kiosk)). - **Network drops:** a session checks your API secret when it starts and keeps rendering through a network drop of up to 5 minutes. - **Session length:** a self-hosted session runs for up to 7 days, then ends with `403 SESSION_DURATION_LIMIT`; start a new one ([Rate limits](https://docs.bithuman.ai/api/rate-limits#session-concurrency)). ## What it costs 2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: [Pricing and credits](https://docs.bithuman.ai/pricing). A session bills while it runs, talking or idle, so close it when the branch closes. ## Agreements Healthcare and financial-services deployments are set up under an enterprise agreement and review. [Contact sales](https://www.bithuman.ai/sales) to start one. --- # Healthcare URL: https://docs.bithuman.ai/deploy/use-cases/healthcare > Configure bitHuman so patient audio, transcripts and video stay in your environment: rendering on your hardware, a local or private language model, and no transcript stored at bitHuman. Check-in desks, wayfinding, visitor information and patient-education screens: an avatar that answers questions in a clinic or a hospital. This page shows where each kind of data goes, so your privacy and security teams can assess a deployment. bitHuman does not certify your deployment; your organization makes its own assessment. ## Keep patient data in your environment 1. **Render on your hardware:** your servers, a Linux PC with no GPU, or your app on the device. When the avatar renders on your hardware, its audio and video stay there. 2. **Keep the conversation there too:** the CLI's [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) (`BITHUMAN_LOCAL=1`) runs speech recognition, the language model and the voice on the machine, and audio, transcripts and generated speech never leave it. Or use any OpenAI-compatible model inside your own network ([Providers](https://docs.bithuman.ai/api/providers)). In an app, the Swift package and the Android SDK take any 16 kHz mono speech your own services produce. 3. **What reaches bitHuman then:** a credential check when a session starts, the avatar download, and usage reports. Usage reports contain no audio, video, images or conversation text. Self-hosted and on-device sessions store no transcript at bitHuman. | Question | Your servers | |---|---| | Where the avatar renders | on your own Mac or Linux machines | | Where the conversation runs | your choice: the CLI's local conversation brain, your own services, or bitHuman's | | What reaches bitHuman | a credential check, the avatar download, and usage reports with no audio, video or text | | Network | to start; rendering continues through a drop of up to 5 minutes | When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. ## Creating the avatar Creating the avatar from a portrait happens in the bitHuman cloud; the finished avatar model then runs on your devices. Use a portrait you have the rights to use. ## In the bitHuman cloud instead With the web embed or a cloud avatar, the session's audio and conversation reach bitHuman to run it: - cloud avatars render in the US; - transcripts are stored with the agent, and deleting the agent deletes its records, including transcripts, and its model files; - traffic is encrypted in transit (HTTPS/TLS; WebRTC media uses DTLS-SRTP), and provider keys you connect are encrypted at rest. ## Access and accounts Organization roles (owner, admin, member), an audit-log API, API-secret rotation with immediate revocation, and scoped runtime and embed tokens so browsers never hold your secret ([Data flows & privacy](https://docs.bithuman.ai/deploy/privacy#security-and-access)). ## Agreements Healthcare and financial-services deployments are set up under an enterprise agreement and review. [Contact sales](https://www.bithuman.ai/sales) to start one. --- # Events and trade shows URL: https://docs.bithuman.ai/deploy/use-cases/events-and-trade-shows > Run an avatar on a show floor with unreliable Wi-Fi: a Linux PC with no GPU, the local conversation brain, a setup checklist, and fully offline for halls with no internet. A booth avatar that greets visitors, answers questions about your product and hands them to your staff. It renders on a machine at the booth, so its video does not travel over the hall's network. ## Where it runs A Linux PC with no GPU (x86_64 or arm64) or a Mac with Apple silicon, running the CLI. Both models run live on a standard Linux PC with no GPU; how fast each one renders is on [Performance](https://docs.bithuman.ai/performance). The step-by-step setup is [Kiosk on a Linux PC](https://docs.bithuman.ai/build/kiosk). ## When the Wi-Fi is weak - **The avatar renders at the booth.** When the avatar renders on your hardware, its audio and video stay there. - **Keep the conversation at the booth too:** the CLI's [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) (`BITHUMAN_LOCAL=1`) runs speech recognition, the language model and the voice on the machine. The session still reports usage online. - **Network drops:** a session checks your API secret when it starts and keeps rendering through a network drop of up to 5 minutes. - **No internet in the hall at all:** [Fully offline](https://docs.bithuman.ai/deploy/offline), on the Business and Enterprise plans. ## Setup checklist 1. **Before the show:** create the avatar and write its [persona](https://docs.bithuman.ai/build/persona). Creating the avatar from a portrait happens in the bitHuman cloud. 2. **On a good connection:** on the booth machine, run `bithuman pull $AGENT_CODE`, then `bithuman run` once with the conversation brain you will use, so the avatar and the brain are on the machine. 3. **Its own API secret:** give the booth a secret of its own, so you can revoke it after the show ([API secrets](https://www.bithuman.ai/developer/api-keys)). 4. **At the booth:** test the microphone with the hall's noise, and show the avatar full screen in Chrome's kiosk mode ([Kiosk on a Linux PC](https://docs.bithuman.ai/build/kiosk#show-it-full-screen)). 5. **When the hall closes:** stop the avatar. A session bills while it runs, talking or idle. ## What it costs 2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: [Pricing and credits](https://docs.bithuman.ai/pricing).