# Deployment options

URL: https://docs.bithuman.ai/deploy

> The avatar renders in one of five places. Pick the one that matches where your users are and what may leave their device.

Every mode uses the same avatar and the same API secret. What changes is where the avatar renders, where the conversation runs, and what reaches bitHuman.

## Compare the modes

| Compare | [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud) | [Your servers](https://docs.bithuman.ai/deploy/self-hosted) | [On the device](https://docs.bithuman.ai/deploy/on-device) | [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu) | [Fully offline](https://docs.bithuman.ai/deploy/offline) |
|---|---|---|---|---|---|
| **The avatar renders** | on bitHuman's servers, in the US | on your own Mac or Linux machines | on the iPhone, iPad, Mac or Android phone in front of the user, or in a WebGPU browser tab | on a standard Linux PC's CPU, with no GPU | on your Linux PCs and terminals |
| **The conversation runs** | on bitHuman's voice service, or with your own provider keys | your choice: the CLI's local conversation brain, your own services, or bitHuman's | your app's choice; with the web embed, on bitHuman's servers | your choice: the CLI's local conversation brain, your own services, or bitHuman's | agreed with sales for your site |
| **What reaches bitHuman** | the session's audio and conversation, to run it | a credential check, the avatar download, and usage reports with no audio, video or text | with your own voice and language services, usage metering only, never audio, video or conversation text | a credential check, the avatar download, and usage reports with no audio, video or text | usage is metered on the machine; no reconnection is required |
| **Network** | required for the whole session | to start; rendering continues through a drop of up to 5 minutes | to start; rendering continues through a drop of up to 5 minutes | to start; rendering continues through a drop of up to 5 minutes | off the internet; creating the avatar happens online first |
| **Price** | 4 credits per minute (Essence 2, Expression 2) | 2 credits per minute (Essence 2, Expression 2) | 2 credits per minute (Essence 2, Expression 2) | 2 credits per minute (Essence 2, Expression 2) | from 100,000 credits; arranged through sales |
| **Products** | web embed, REST API, LiveKit plugin, cloud avatar in your room | CLI, Python SDK, LiveKit plugin (`model_path`) | Swift package, Android SDK, Flutter plugin, web embed with `render=local` | CLI, Python SDK on Linux x86_64 or arm64 | Linux PCs and terminals |
| **Models** | Essence 2, Expression 2, Essence 1, Expression 1 | Essence 2, Expression 2, Essence 1 | Essence 2, Expression 2 | Essence 2, Expression 2, Essence 1 | Essence 1 (Linux x86_64 and ARM64); Essence 2 and Expression 2 later |
| **Plan** | Creator plan or higher | Creator plan or higher | Creator plan or higher | Creator plan or higher | Business & Enterprise |

Every mode bills active session time, talking or idle, to the second ([pricing](https://docs.bithuman.ai/pricing)).

## bitHuman cloud

bitHuman renders the avatar in the US and streams it to your page, app or LiveKit room. [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud)

## Your servers (self-hosted)

The CLI, Python or the LiveKit plugin on your own Mac or Linux machines. When the avatar renders on your hardware, its audio and video stay there. [Your servers](https://docs.bithuman.ai/deploy/self-hosted)

## On the device

Essence 2 and Expression 2 render on iPhone, iPad, Mac and Android, or in a WebGPU browser tab. [On the device](https://docs.bithuman.ai/deploy/on-device)

## CPU only (no GPU)

Both models run live on a standard Linux PC with no GPU. [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu)

## Fully offline

Offline license is only available to Business and Enterprise clients who want to run realtime avatars completely locally, off the internet — e.g. kiosks, trade shows, ATM machines, embedded screens. Linux PCs and terminals; arranged through sales. [Fully offline](https://docs.bithuman.ai/deploy/offline)

## Choosing a mode

- **Building for the web, or want the fewest moving parts?** The [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud), starting with the [web embed](https://docs.bithuman.ai/platforms/web).
- **Must audio and video stay on your network?** [Your servers](https://docs.bithuman.ai/deploy/self-hosted) or [on the device](https://docs.bithuman.ai/deploy/on-device), with the [local conversation brain](https://docs.bithuman.ai/platforms/cli/local-brain) or your own models. What reaches bitHuman in each: [Data flows & privacy](https://docs.bithuman.ai/deploy/privacy).
- **Want the lowest per-minute rate?** Render on the device or on your servers ([rates](https://docs.bithuman.ai/pricing)).
- **No reliable internet at the site?** [Fully offline](https://docs.bithuman.ai/deploy/offline), on the Business and Enterprise plans.
- **Planning for a branch, a clinic or a show floor?** Start from [Use cases](https://docs.bithuman.ai/deploy/use-cases).

## Try it

Talk to a sample avatar streamed from the bitHuman cloud: the same web embed a customer puts on a site.

## Pages in this section

## Overview

- [Deployment options](https://docs.bithuman.ai/deploy.md): The avatar renders in one of five places. Pick the one that matches where your users are and what may leave their device.
- [Data flows & privacy](https://docs.bithuman.ai/deploy/privacy.md): What reaches bitHuman in each deployment mode: where the avatar renders, where the conversation runs, what usage reports carry, and what bitHuman stores.
- [Pricing and credits](https://docs.bithuman.ai/pricing.md): Credits pay for active session time, talking or idle, by the exact second. Rates per model and platform, creation costs, plans, offline licensing, and how to check your balance.

## Modes

- [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud.md): bitHuman renders the avatar on its servers and streams it to your page, app or LiveKit room: the web embed, the REST API or the LiveKit plugin.
- [Your servers (self-hosted)](https://docs.bithuman.ai/deploy/self-hosted.md): Run the CLI, the Python SDK or the LiveKit plugin on your own Mac or Linux machines. When the avatar renders on your hardware, its audio and video stay there.
- [On the device](https://docs.bithuman.ai/deploy/on-device.md): Essence 2 and Expression 2 render on the device in front of the user: iPhone, iPad and Mac, Android phones, or a WebGPU browser tab.
- [CPU only (no GPU)](https://docs.bithuman.ai/deploy/cpu.md): Both models run live on a standard Linux PC with no GPU.
- [Fully offline](https://docs.bithuman.ai/deploy/offline.md): Realtime avatars that run completely locally, off the internet, on Linux PCs and terminals: for Business and Enterprise clients, arranged through sales.

## Use cases

- [Use cases](https://docs.bithuman.ai/deploy/use-cases.md): Which deployment mode each use case starts from: banking and ATMs, healthcare, events and trade shows, kiosks, apps and AI companions, and websites.
- [Banking and ATMs](https://docs.bithuman.ai/deploy/use-cases/banking-and-atms.md): Run an always-on avatar on a branch screen, teller terminal or ATM: the avatar handles the conversation, card handling and transactions stay in your systems. Hardware, data flows, operations and cost.
- [Healthcare](https://docs.bithuman.ai/deploy/use-cases/healthcare.md): Configure bitHuman so patient audio, transcripts and video stay in your environment: rendering on your hardware, a local or private language model, and no transcript stored at bitHuman.
- [Events and trade shows](https://docs.bithuman.ai/deploy/use-cases/events-and-trade-shows.md): Run an avatar on a show floor with unreliable Wi-Fi: a Linux PC with no GPU, the local conversation brain, a setup checklist, and fully offline for halls with no internet.
