# bitHuman — Models (Essence 2, Expression 2, Essence 2 Max, the first generation, how it works) > Realtime talking avatars from one portrait. Index: https://docs.bithuman.ai/llms.txt · every page is also served as markdown at .md · OpenAPI: https://docs.bithuman.ai/api/openapi.yaml · other sections: https://docs.bithuman.ai/llms/start.txt · https://docs.bithuman.ai/llms/platforms.txt · https://docs.bithuman.ai/llms/deploy.txt · https://docs.bithuman.ai/llms/build.txt · https://docs.bithuman.ai/llms/api.txt ## Instructions for AI agents - Always send `model` ("essence-2" or "expression-2") when creating an agent; poll until `status` is `ready` or `failed`. - API and SDK use requires the Creator plan or higher. Never tell a user they can build on a free plan. - In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET` and pass a minted token (https://docs.bithuman.ai/platforms/livekit.md). - Say where things happen: the avatar renders (device, browser, server, cloud); the conversation runs (your stack, the CLI's local conversation brain, or bitHuman's servers). - Quote speed only from https://docs.bithuman.ai/performance.json (× real time, with the device). Quote prices only from GET https://api.bithuman.ai/v1/pricing. - Do not claim offline on phones, Mac or the browser, a conversation brain on phones, or any certification. Where each model runs, the key facts and one command per path: https://docs.bithuman.ai/llms.txt ## Contents - Compare models — https://docs.bithuman.ai/models - Essence 2 — https://docs.bithuman.ai/models/essence-2 - Expression 2 — https://docs.bithuman.ai/models/expression-2 - First generation — https://docs.bithuman.ai/models/first-generation - How it works — https://docs.bithuman.ai/models/how-it-works - The avatar file — https://docs.bithuman.ai/models/avatar-file Linked, not inlined (read the .md twin): - Examples: https://docs.bithuman.ai/examples.md · changelog: https://docs.bithuman.ai/changelog.md · API references: https://docs.bithuman.ai/platforms/cli/reference.md, https://docs.bithuman.ai/platforms/python/reference.md, https://docs.bithuman.ai/platforms/swift/reference.md, https://docs.bithuman.ai/platforms/android/reference.md --- # Compare models URL: https://docs.bithuman.ai/models > Essence 2 renders a photoreal person and Expression 2 any character, each from one portrait. Where each model renders, which to pick, and how an avatar is created. Every model reads the same [`.imx` avatar file](https://docs.bithuman.ai/models/avatar-file) and has the same shape: [push audio in, take lip-synced frames out](https://docs.bithuman.ai/models/how-it-works#audio-in-frames-out). The same agent works on every platform that runs its model. ## The models - [Essence 2](https://docs.bithuman.ai/models/essence-2): a photoreal person from one portrait (Renders on the device, bitHuman cloud) - [Expression 2](https://docs.bithuman.ai/models/expression-2): any character from one portrait (Renders on the device, bitHuman cloud) Essence 2 Max is available on the Enterprise plan only. [Contact sales](https://www.bithuman.ai/sales) to enable it. Essence 1 and Expression 1 are the [first generation](https://docs.bithuman.ai/models/first-generation). They stay supported, and nothing changes for agents that use them. ## Which should I choose? - **A photorealistic person:** Essence 2. - **A stylized or non-human character, or a whole generated scene:** Expression 2. - **Not sure:** create with `model: "auto"`. A photorealistic person routes to Essence 2, anything else to Expression 2. - **On a phone, a Mac or in a browser:** Essence 2 or Expression 2. - **Maintaining a first-generation agent:** keep it. Essence 1 runs on your own CPU; Expression 1 runs in the bitHuman cloud. ## Where each model runs Each place links to the page that sets it up. | Where | Essence 2 | Expression 2 | Essence 1 | Expression 1 | |---|---|---|---|---| | [iPhone and iPad](https://docs.bithuman.ai/platforms/ios) | Yes | Yes | — | — | | [Mac](https://docs.bithuman.ai/platforms/macos) | Yes | Yes | Yes | — | | [Android](https://docs.bithuman.ai/platforms/android) | Yes | Yes | — | — | | [Linux, no GPU](https://docs.bithuman.ai/deploy/cpu) | Yes | Yes | Yes | — | | [Browser (WebGPU)](https://docs.bithuman.ai/platforms/web) | Yes | Yes | Yes | — | | [Your servers](https://docs.bithuman.ai/deploy/self-hosted) | Yes | Yes | Yes | — | | [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud) | Yes | Yes | Yes | Yes | | [Fully offline](https://docs.bithuman.ai/deploy/offline) | — | — | Yes | — | Fully offline is for Business and Enterprise clients, arranged through sales ([Fully offline](https://docs.bithuman.ai/deploy/offline)). Rendering on your own hardware, every place but the bitHuman cloud, bills at the self-hosted rate ([pricing](https://docs.bithuman.ai/pricing)). How fast each model renders on each device: [Performance](https://docs.bithuman.ai/performance). ## How creation works You create an agent once, with [`POST /v1/agent/generate`](https://docs.bithuman.ai/api/agents#generate-an-agent) or in the bitHuman app, and serve it anywhere its model runs. - **The input is one portrait image.** Essence 2 generates its identity video from it; Expression 2 trains straight from the photo. An uploaded image is treated as a reference and regenerated to a standard framing. - **Creation happens in the bitHuman cloud;** the finished avatar model then runs on your devices. - **Both second-generation models train on create.** Allow about 2 to 2.5 hours, and poll [`GET /v1/agent/status/{agent_id}`](https://docs.bithuman.ai/api/agents#poll-status) until the status is `ready` or `failed` (`success` is not terminal). - **Essence 2 needs a photorealistic human subject.** A stylized input is refused with [`422 MODEL_SUBJECT_MISMATCH`](https://docs.bithuman.ai/api/errors#model-errors) before anything is billed; `auto` routes it to Expression 2 instead. - **Always send `model`.** An omitted `model` creates an Expression 1 agent; send `essence-2`, `expression-2` or `auto`. - **An existing agent can gain a model** with [`POST /v1/agent/{code}/models`](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent). What creation costs is on [pricing](https://docs.bithuman.ai/pricing#creation--one-time-credits); request fields and failure modes are on the [Agents API](https://docs.bithuman.ai/api/agents). ## Naming & migration This is the one place the historical names are documented. Every other page uses the four product names. Deprecating a word does not rename a wire format, so some legacy names are still strings you read or type: | Legacy name you may meet | Where | What it means | Do you type it? | |---|---|---|---| | `essence`, `expression` | older `?model=` links and request bodies | Essence 1, Expression 1 | No — write `essence-1` / `expression-1` | | `essence2-light` | the `Engine:` line from `bithuman open` — a [legacy engine value](https://docs.bithuman.ai/models/avatar-file#the-engine-value-is-a-legacy-name) | Essence 2 | No — read the `Family:` line | | `essence-2-light` | the retired tier name (the old Light tier) | Essence 2 | No — a request naming it gets a `400`; write `essence-2` | | `elevate`, `essence-2-quality` | retired names of the premium tier, now Essence 2 Max (Enterprise plan only) | a separate tier, not Essence 2 | No — a request naming them gets a `400` | | `embody` | a retired request spelling | Expression 2 | No — a request naming it gets a `400` naming `expression-2` | | `.lebundle.imx`, `.avatar` | older file extensions | an Essence 2 or Expression 2 model file | Only if you already have one; it opens as-is | | `[embody]` | the legacy prefix on log lines of the Apple `Expression2` engine | Expression 2 | Grep your logs for it | | `BITHUMAN_EMBODY_DIR`, `EMBODY_DEBUG_FAIL_PREDICT` | legacy variables the Apple `Expression2` engine still reads beside their `EXPRESSION2_` twins | Expression 2 | No — set `BITHUMAN_EXPRESSION2_DIR` | | `libelevate`, `libelevate-android` | legacy library names | Essence 2 | No — the Android coordinate is `ai.bithuman:essence2-android` | | `libelevate-web` | the legacy path of the in-browser runtime | Essence 2 in a browser | No — embed with `https://www.bithuman.ai/embed/` | | `bithuman.tessera_offline`, `OfflineTesseraRenderer`, `TesseraOfflineError` | legacy Python module and class names, still importable | the Essence 2 MP4 route | No — write `bithuman.offline`, `OfflineRenderer`, `render_offline`, `OfflineRenderError` | | `BITHUMAN_TESSERA_DIRECTOR` and the other `BITHUMAN_TESSERA_*` variables | legacy environment variables, still read | Essence 2 engine settings | No — the defaults are the fast path | | `bithuman[tessera]`, `bithuman[offline]` | legacy pip extras, removed from the wheel in 2.11.6 | nothing — pip warns and installs the base wheel | No — `pip install bithuman` | Saved links keep working: `essence-2-light-gpu` / `essence-2-light-cpu` still pin their tiers, links carrying `essence-2-light` or `essence-2-light-ane` route to the Essence 2 default chain, and the older `essence-2-ane` / `expression-2-ane` spellings of the Apple tier stay accepted. A link carrying the retired `?model=essence-2-quality` falls back to the agent's stored model. One more naming point: the cloud's Apple tier is called **Apple**, not "ANE". It is the whole Apple silicon target, not one accelerator inside it. --- # Essence 2 URL: https://docs.bithuman.ai/models/essence-2 > Essence 2 renders a photoreal person from one portrait: the identity's own footage, lip-synced live, on the device or in the bitHuman cloud. ## What it is **Essence 2** (`essence-2`) renders a photoreal person from one portrait, up to 1080p: the identity's own canvas, 1080×1920 portrait for a standard identity. From your portrait the platform generates a 10-second identity video; the model then animates lip-sync and expression over it live, with a sharp mouth and teeth taken from that video. ## When to choose it - **A photorealistic person** — start here. - **Always-on displays** — kiosks, lobby screens and 24/7 assistants. - **On your own hardware** — every SDK platform runs it. For a stylized character, or a scene generated from one photo, choose [Expression 2](https://docs.bithuman.ai/models/expression-2). The side-by-side is on [Models](https://docs.bithuman.ai/models). ## Where it runs | Where | Essence 2 | How | |---|---|---| | [iPhone and iPad](https://docs.bithuman.ai/platforms/ios) | Yes | Swift package, `Essence2Kit` (iOS 26) | | [Mac](https://docs.bithuman.ai/platforms/macos) | Yes | Swift package (macOS 26, M3 or newer), CLI, Python | | [Android](https://docs.bithuman.ai/platforms/android) | Yes | `essence2-android` | | [Linux, no GPU](https://docs.bithuman.ai/deploy/cpu) | Yes | CLI, Python | | [Browser (WebGPU)](https://docs.bithuman.ai/platforms/web) | Yes | `render=local`, for identities with a browser build | | [Your servers](https://docs.bithuman.ai/deploy/self-hosted) | Yes | CLI, Python, LiveKit plugin | | [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud) | Yes | web embed, REST API, LiveKit | | [Fully offline](https://docs.bithuman.ai/deploy/offline) | — | Coming later | A complete app for iPhone and iPad is the [iOS Essence 2 example](https://docs.bithuman.ai/examples/ios-essence-2); for Android, the [Android Essence 2 example](https://docs.bithuman.ai/examples/android-essence-2). The file you download is `.imx`, from [`GET /v1/agent/{code}/model/download?model=essence-2`](https://docs.bithuman.ai/api/agents#download-an-agents-model) or `bithuman pull --model essence-2`. How fast it renders on each device is on [performance](https://docs.bithuman.ai/performance). ## How creation works Create the agent with [`POST /v1/agent/generate`](https://docs.bithuman.ai/api/agents#generate-an-agent) and `model: "essence-2"`, or add `essence-2` to an existing agent with [`POST /v1/agent/{code}/models`](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent). - **The input is a portrait image** of a photorealistic human. A stylized or non-human input is refused with [`422 MODEL_SUBJECT_MISMATCH`](https://docs.bithuman.ai/api/errors#model-errors) before anything is billed; `model: "auto"` routes it to Expression 2 instead. - **The platform generates the identity video** from the image, then trains the identity. Poll [`GET /v1/agent/status/{agent_id}`](https://docs.bithuman.ai/api/agents#poll-status) until `ready`; allow **about 2 to 2.5 hours**. - **`ready` serves before it downloads.** The downloadable file is published a little later; until then the download endpoint answers a retryable `404 MODEL_ARTIFACT_NOT_READY`. The creation cost is on [pricing](https://docs.bithuman.ai/pricing). ## Serving tiers In the bitHuman cloud, the service picks the hardware for each session. To benchmark one tier, see [pin a tier for a benchmark](https://docs.bithuman.ai/performance#pin-a-tier-for-a-benchmark); in production, let the service choose. ## Idle and speaking behavior The identity video plays continuously and loops **forward-only**: at its last frame it wraps to the first, and it never plays in reverse. While idle it is pure playback of your footage; while talking, the animated face is rendered over the same frames. A running session bills talking and idle time alike ([pricing](https://docs.bithuman.ai/pricing)). ## Limits and expectations - **Output plays at 25 frames a second** everywhere it runs. How fast a platform renders is on [performance](https://docs.bithuman.ai/performance). - **The downloadable file is about 140–160 MB**, varying per identity — read `Content-Length` rather than assuming a size. - **The identity is fixed at creation.** To change the face, create a new agent. - **The first session on a new agent** can take longer to connect while the identity is provisioned; later sessions reuse it. - **Before training completes**, a launch that requests this model is refused with [`409 MODEL_NOT_GENERATED`](https://docs.bithuman.ai/api/errors#model-errors). Once ready, `essence-2` appears in the agent's `supported_models`. ## Next steps - [Models](https://docs.bithuman.ai/models) — the four models side by side - [Agents API](https://docs.bithuman.ai/api/agents) — create, poll, download - [Embed widget](https://docs.bithuman.ai/api/embedding) — a live session in minutes - [Video API](https://docs.bithuman.ai/api/video) — render an MP4 with `model: "essence-2"` - [Session behavior & troubleshooting](https://docs.bithuman.ai/resources/troubleshooting) --- # Expression 2 URL: https://docs.bithuman.ai/models/expression-2 > Expression 2 renders any character from one portrait: the whole scene generated live from the audio, on the device or in the bitHuman cloud. ## What it is **Expression 2** (`expression-2`) generates the whole avatar scene live from the audio — expressions, mouth and head movement are synthesized each session, not replayed from a base video. It animates the **entire 416x720 frame** with no face detector or cropping step, so it works for any character: cartoons, animals, creatures, robots, and people. At creation the platform trains a **small model of your specific identity** from one photo. That per-identity model is what serves your sessions, and it is why creation takes a couple of hours. ## When to choose it - **Your character is not a photorealistic human** — this is the model for it, and where `model: "auto"` routes such inputs. - **You want motion generated from the audio itself**, not patched onto a base video. - **You only have a photo** — one image is enough. For a photorealistic person animated from their own footage, compare [Essence 2](https://docs.bithuman.ai/models/essence-2). The side-by-side is on [Models](https://docs.bithuman.ai/models). ## Where it runs | Where | Expression 2 | How | |---|---|---| | [iPhone and iPad](https://docs.bithuman.ai/platforms/ios) | Yes | Swift package, `Expression2` | | [Mac](https://docs.bithuman.ai/platforms/macos) | Yes | Swift package (macOS 13), CLI, Python | | [Android](https://docs.bithuman.ai/platforms/android) | Yes | `expression2-android` | | [Linux, no GPU](https://docs.bithuman.ai/deploy/cpu) | Yes | CLI, Python | | [Browser (WebGPU)](https://docs.bithuman.ai/platforms/web) | Yes | `render=local` | | [Your servers](https://docs.bithuman.ai/deploy/self-hosted) | Yes | CLI, Python, LiveKit plugin | | [bitHuman cloud](https://docs.bithuman.ai/deploy/cloud) | Yes | web embed, REST API, LiveKit | | [Fully offline](https://docs.bithuman.ai/deploy/offline) | — | Coming later | Complete apps: [iOS Expression 2](https://docs.bithuman.ai/examples/ios-expression-2), [macOS Expression 2](https://docs.bithuman.ai/examples/macos-expression-2) and [Android Expression 2](https://docs.bithuman.ai/examples/android-expression-2). The file you download from [`GET /v1/agent/{code}/model/download?model=expression-2`](https://docs.bithuman.ai/api/agents#download-an-agents-model) or `bithuman pull ` is labelled `.imx`; `.avatar` is the legacy extension for the same container. How fast it renders on each device is on [performance](https://docs.bithuman.ai/performance). ## How creation works Create the agent with [`POST /v1/agent/generate`](https://docs.bithuman.ai/api/agents#generate-an-agent) and `model: "expression-2"`, or add `expression-2` to an existing agent with [`POST /v1/agent/{code}/models`](https://docs.bithuman.ai/api/agents#add-a-model-to-an-existing-agent). - **The input is a portrait image**, of any subject. Without one, the platform generates a portrait from your prompt first. It also generates the agent's 10-second idle clip and prepares a voice. - **Training takes about 2 to 2.5 hours;** an identity that needs more work gets more, so up to 4 hours is normal. Poll [`GET /v1/agent/status/{agent_id}`](https://docs.bithuman.ai/api/agents#poll-status) until `ready` or `failed`, or wait for the completion email. - **A run that fails is refunded;** a completed creation is not, so a second `generate` is a second charge ([failure modes](https://docs.bithuman.ai/api/agents#errors)). The creation cost is on [pricing](https://docs.bithuman.ai/pricing). ## Serving tiers Every published configuration, including a desktop CPU with no GPU, renders faster than real time ([performance](https://docs.bithuman.ai/performance)). In the bitHuman cloud, the service picks the hardware for each session; to benchmark one tier, see [pin a tier for a benchmark](https://docs.bithuman.ai/performance#pin-a-tier-for-a-benchmark). ## Idle and speaking behavior During silences the avatar plays its **10-second idle clip**, generated from the identity at creation, looping forward-only without a seam. When speech starts, the engine hands off to generated frames with a per-identity color match, so the two stay visually continuous; idle resumes only after sustained silence, not in pauses inside a sentence. A running session bills talking and idle time alike ([pricing](https://docs.bithuman.ai/pricing)). **Speech onset.** The engine renders in fixed audio chunks; the moving idle clip covers the start of each reply. ## Limits and expectations - **Output is the full 416×720 scene**, playing at 20 frames a second. - **A clear, frontal, well-lit photo** gives the best result. The identity is fixed at creation — to change the face, create a new agent. - **The first session on a new agent** can take longer to connect while its model is provisioned; later sessions reuse it. - **Before training completes**, a launch that requests this model is refused with [`409 MODEL_NOT_GENERATED`](https://docs.bithuman.ai/api/errors#model-errors). Once ready, `expression-2` appears in the agent's `supported_models`. ## Next steps - [Models](https://docs.bithuman.ai/models) — the four models side by side - [Agents API](https://docs.bithuman.ai/api/agents) — create, poll, download - [Embed widget](https://docs.bithuman.ai/api/embedding) — a live session in minutes - [Video API](https://docs.bithuman.ai/api/video) — render an MP4 with `model: "expression-2"` - [Session behavior & troubleshooting](https://docs.bithuman.ai/resources/troubleshooting) --- # First generation URL: https://docs.bithuman.ai/models/first-generation > Essence 1 and Expression 1, bitHuman's first-generation avatar models: what each one is, where it runs, and what its file contains. Both are maintained; new work starts on Essence 2 or Expression 2. Both first-generation models are maintained, not deprecated. For new work, start on [Essence 2](https://docs.bithuman.ai/models/essence-2) (a photoreal person) or [Expression 2](https://docs.bithuman.ai/models/expression-2) (any character). | | Essence 1 | Expression 1 | |---|---|---| | **Renders** | a pre-built identity from an `.imx` file | facial motion generated from a portrait at runtime | | **Where it runs** | the bitHuman cloud; Python and the CLI on macOS and Linux; the browser (`render=local`) | the bitHuman cloud only | | **`model` value** | `essence-1` | `expression-1` | ## Essence 1 **Essence 1** (`essence-1`) reads a pre-built identity out of an [`.imx` file](https://docs.bithuman.ai/models/avatar-file), plays its base motion, and patches the mouth in real time to match 16 kHz mono audio, at 25 fps. It runs on a CPU, with no GPU or accelerator, and supports custom [gestures](https://docs.bithuman.ai/build/gestures). `?model=essence-1` serves it. ### Where Essence 1 runs In the bitHuman cloud, and on your own hardware through: - **Python**: `pip install bithuman` opens an Essence 1 `.imx` with no extra; see [Python](https://docs.bithuman.ai/platforms/python). - **The CLI**: `bithuman run` on macOS (Apple silicon) and Linux x86_64 and arm64; see [CLI](https://docs.bithuman.ai/platforms/cli). `render` does not take Essence 1: use the Python SDK or the [Talking video API](https://docs.bithuman.ai/api/video) for a file. - **The browser**: [`?render=local`](https://docs.bithuman.ai/platforms/web#integrate-into-your-app). Essence 1 is not in the Android SDK or the Swift package. On phones, use Essence 2 or Expression 2, or run Essence 1 from the cloud API, or from Python or the CLI on a desktop. The full matrix is on [Compare models](https://docs.bithuman.ai/models#where-each-model-runs). ### The Essence 1 file One file, `.imx`: the identity, and optionally baked-in idle and gesture clips. Download it with `bithuman pull --model essence-1`, or: ```bash curl -L -o ".imx" -H "api-secret: $BITHUMAN_API_SECRET" \ "https://api.bithuman.ai/v1/agent//model/download?model=essence-1" ``` ## Expression 1 **Expression 1** (`expression-1`) animates a face from a **portrait image** at runtime: you give it audio, and it generates the facial motion to match, with no per-identity build step. `?model=expression-1` serves it, and it is what `/v1/agent/generate` creates when the request names no `model`. It is a different engine from [Expression 2](https://docs.bithuman.ai/models/expression-2), not an earlier version of it. ### Where Expression 1 runs In the bitHuman cloud only, on cloud GPUs. There is no CPU, Apple, Android or browser build. For an expressive model on a Mac, a phone or in a browser, use Expression 2. Cloud output is 512×512. Expression 1 can also animate a photo with no agent: see [Cloud avatar in your room](https://docs.bithuman.ai/api/cloud-avatar#a-photo-instead-of-an-agent). ### The Expression 1 file Usually there is nothing to download: Expression 1 renders from the agent's portrait, so the download endpoint answers [`400 MODEL_NOT_DOWNLOADABLE`](https://docs.bithuman.ai/api/errors#model-errors) for most agents. A few older agents have a downloadable `.imx`, which the endpoint serves. The engine weights are not part of any download. ## Pricing Rates for both models are on [Pricing and credits](https://docs.bithuman.ai/pricing). --- # How it works URL: https://docs.bithuman.ai/models/how-it-works > How bitHuman is built: one portable engine, thin language SDKs on top, and your app on top of that. Push 16 kHz audio in, drain lip-synced frames out, the same way on every platform. *Diagram: The engine: speech in, frames out.* Your app pushes 16 kHz mono speech into the bitHuman engine and pulls lip-synced frames out. The same engine sits inside the Swift package, the Android SDK, the Python SDK, the CLI and the web embed. ## The three layers bitHuman is one portable engine with thin language bindings on top, and your app on top of that. Every layer reads the same [model file](https://docs.bithuman.ai/models/avatar-file) and produces the same lip-synced frames, on an iPhone, a Mac, a Linux PC, in a browser or in the bitHuman cloud. | Layer | What it is | |---|---| | **Apps and tools** | the bitHuman CLI, your own app, LiveKit for WebRTC transport | | **Language SDKs** | Python, Swift and Kotlin: thin bindings over the same engine; the browser through the web embed | | **The bitHuman engine** | the avatar renderer, inside every SDK, so there is nothing separate to install: macOS, iOS, Android, Linux and the browser | You integrate at the SDK layer. The engine is built into each SDK, so your app needs the bitHuman dependency and nothing else. To pick a platform, start at [Platforms](https://docs.bithuman.ai/platforms); which model runs where is on [Models](https://docs.bithuman.ai/models#where-each-model-runs). ## What stays true across every surface - **One model file, every surface.** The same audio drives the same lip-sync on every SDK; pixels can differ slightly between hardware backends. - **A stable public API.** Deprecated options keep working with a warning until the next major, and majors call out breaks explicitly. - **Surfaces mix.** The Swift package in your iOS app with the Python package on your backend is supported; keep each one current — [Downloads](https://docs.bithuman.ai/downloads#current-versions) lists the current versions. - **One credential.** The same key drives every surface; how it is exchanged and billed is on [Authentication](https://docs.bithuman.ai/api/authentication) and [pricing](https://docs.bithuman.ai/pricing). ## Audio in, frames out Every SDK has the same shape — audio in, video out: 1. **Push** audio as it arrives — a microphone, TTS, a WebRTC track. 2. **Drain** lip-synced frames at the model's own rate — 25 fps for Essence 2 and Essence 1, 20 fps for Expression 2. The engine buffers between the two, so your audio source and your render loop never have to run in lockstep. ### In Python `render()` takes audio a chunk at a time and yields frames as it goes, so a stream and a file are the same program: ```python import numpy as np import bithuman def chunks(path="speech16k.raw", ms=40): """16 kHz mono int16 audio, delivered a chunk at a time — a mic, TTS or a socket.""" pcm = np.fromfile(path, dtype=np.int16) step = 16000 * ms // 1000 for i in range(0, len(pcm), step): yield pcm[i:i + step] frames = 0 with bithuman.open("wise-pup.imx") as avatar: for image in avatar.render(chunks()): # frames come out as audio goes in frames += 1 # image: (height, width, 3) uint8, RGB print(frames, image.shape) ``` Install, the model download and the credential are on the [Python SDK](https://docs.bithuman.ai/platforms/python) page. `speech16k.raw` is any speech converted with `ffmpeg -i speech.wav -ac 1 -ar 16000 -f s16le speech16k.raw`. ### Audio format | Property | Value | |---|---| | Encoding | 16-bit signed PCM (`int16`), or `float32` in [-1, 1] | | Channels | mono | | Sample rate | 16 kHz for decoded samples; a file path in any format ffmpeg reads is converted for you | | Chunk size | anything; 10–40 ms is typical | ### Frame format Frames arrive at the model's own rate, whatever the chunk size: 25 fps for Essence 2 (up to 1080p: the identity's own canvas, 1080×1920 portrait for a standard identity) and 20 fps for Expression 2 (416x720). Python yields RGB `uint8` arrays; the Swift package and the Android SDK hand you their platform's image types. ### In the other SDKs - **Apple** — `feed()` PCM, then `pull()` frames. See the [Swift package](https://docs.bithuman.ai/platforms/ios). - **Android** — `feed()` PCM, then `pull()` into a reused buffer: a `Bitmap` for Expression 2, an RGBA `ByteBuffer` for Essence 2. See the [Android SDK](https://docs.bithuman.ai/platforms/android). - **CLI** — `bithuman render` takes an audio file; `bithuman run` streams a live conversation. See the [CLI](https://docs.bithuman.ai/platforms/cli). --- # The avatar file URL: https://docs.bithuman.ai/models/avatar-file > The self-contained .imx file every bitHuman avatar ships in — one container for Essence 1, Essence 2 and Expression 2 identities — where it comes from, how it's addressed by agent code, and how to inspect it. ## What an `.imx` is An `.imx` file is the container a bitHuman avatar ships in: one self-contained file of identity weights, textures and a manifest (model version, ABI, license) that an [engine](https://docs.bithuman.ai/models/how-it-works) reads to animate one specific face. Every model that renders on your own hardware uses it — a first-generation [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) identity, an [Essence 2](https://docs.bithuman.ai/models/essence-2) identity, and an [Expression 2](https://docs.bithuman.ai/models/expression-2) identity. Every download is named `.imx`; older Expression 2 files may carry the legacy `.avatar` extension, which opens the same way. The same file opens on every on-device runtime — [Python](https://docs.bithuman.ai/platforms/python), [Swift](https://docs.bithuman.ai/platforms/ios) and the [CLI](https://docs.bithuman.ai/platforms/cli) — and `bithuman open` tells you which model a file you were given holds. ## Where `.imx` files come from | Source | How | |---|---| | **Showcase** | `bithuman pull ` — pre-built avatars from [bithuman.ai → Explore](https://www.bithuman.ai/explore), which opens on Essence 2 and Expression 2 agents. | | **Dashboard** | Upload a portrait + voice samples in [bithuman.ai → Studio](https://www.bithuman.ai). | | **API** | [`POST /v1/agent/generate`](https://docs.bithuman.ai/api/reference) returns an `agent_code` whose `.imx` you can download. | See [Building avatars](https://docs.bithuman.ai/build/create-avatar) for the full creation flow and media tips. ## Agent codes The `.imx` is keyed by an **agent code** (e.g. `A23WJF0199`). The **cloud runtime and REST API** resolve an agent by its code — you don't ship a file. The **on-device SDKs open a local `.imx`** — the file you downloaded for that code — and the key comes from `BITHUMAN_API_SECRET` in the environment, checked at the first frame: ```python import bithuman with bithuman.open("A23WJF0199.imx") as avatar: # the local file — required on-device for image in avatar.render("speech.wav"): # (height, width, 3) uint8, RGB ... ``` To get the file for a local run, download it by code or slug — `bithuman pull ` on macOS or Linux, or [`GET /v1/agent/{code}/model/download`](https://docs.bithuman.ai/api/agents#download-an-agents-model) — see [Caching for offline use](#caching-for-offline-use). > **Note** Use `agent_code`, never the deprecated `figure_id` — the old identifier returns a 400. ## Caching for offline use You can also pull the file down and pass it by path. A showcase slug needs no account — `bithuman pull` downloads it anonymously: ```bash bithuman pull sofia-ramirez # → ~/.cache/bithuman/showcase/sofia-ramirez.imx ``` `sofia-ramirez` (agent code `A52DHS2219`) is an Essence 2 sample identity from the showcase, about 148 MB. `bithuman list` prints every showcase slug; a slug that is not in that list is refused with `slug '' not found in manifest`. `bithuman pull `, `bithuman list` and `bithuman open` need no credential for a sample avatar. Pulling your own agent by code, and playing any model with `bithuman run` or `bithuman render`, need `bithuman login` or `BITHUMAN_API_SECRET`; session time bills at the [published rates](https://docs.bithuman.ai/pricing). Cache locations by surface: | Surface | Cache location | |---|---| | CLI | pulls in `~/.cache/bithuman/showcase/` (samples) and `~/.cache/bithuman/agents/` (your agents); unpacked copies in `~/.cache/bithuman/bundles/` | | Python | unpacked copies in `~/.cache/bithuman/avatars/`; engine files in `~/.bithuman/deps/` | | Swift (Expression on Mac/iPad) | `~/.cache/bithuman/expression/` | Downloads are integrity-verified and cached. Later launches skip the download. ## One container, one file per model Each model produces its own per-identity file in that container, downloaded with [`GET /v1/agent/{code}/model/download`](https://docs.bithuman.ai/api/agents#download-an-agents-model) (or `bithuman pull `, with `--model` when the agent has more than one): | Model | Artifact | What it is | |---|---|---| | [`essence-1`](https://docs.bithuman.ai/models/first-generation#essence-1) | `.imx` | The first-generation identity — a pre-rendered base whose mouth is patched to the audio. Opens in the [Python SDK](https://docs.bithuman.ai/platforms/python) and the [CLI](https://docs.bithuman.ai/platforms/cli)'s `run`. | | [`essence-2`](https://docs.bithuman.ai/models/essence-2) | `.imx` | The Essence 2 bundle; size is per identity, so read `Content-Length`. Licensed weights; renders locally in the [CLI](https://docs.bithuman.ai/platforms/cli#platform-notes), the [Python SDK](https://docs.bithuman.ai/platforms/python), the [Android library](https://docs.bithuman.ai/platforms/android) and the Swift [`Essence2` product](https://docs.bithuman.ai/platforms/ios) — the first local play checks the license with the cloud, so it needs your sign-in. | | [`expression-2`](https://docs.bithuman.ai/models/expression-2) | `.imx` (older downloads: `.avatar`): the same container under two names (a few early identities use an older format; `bithuman open` tells you which) | Renders locally in the [CLI](https://docs.bithuman.ai/platforms/cli), [Python](https://docs.bithuman.ai/platforms/python), [Apple](https://docs.bithuman.ai/platforms/ios) and [Android](https://docs.bithuman.ai/platforms/android), or on the cloud. | Older releases saved Essence 2 files as `.lebundle.imx`, a legacy extension. Such a file keeps working and `bithuman open` reads it; today's downloads are named `.imx`. The model is [`essence-2`](https://docs.bithuman.ai/models/essence-2). ## Inspecting an `.imx` `bithuman open ` prints the container format, the model family (`Family: essence-2 (Essence 2)`) and the files inside; `--json` adds the manifest. It reads the file on your own disk, so it needs no account and no network: ```bash bithuman open ~/.cache/bithuman/showcase/sofia-ramirez.imx ``` ### The `engine` value is a legacy name `bithuman open` reports an **`engine`** read from the container header (also `engine` in [`--json`](https://docs.bithuman.ai/platforms/cli/reference#json-output)), and the Python runtime quotes the same string verbatim in load errors — for example `backend loader for engine='essence2-light'`. **These engine ids are legacy names kept for compatibility.** They are the literal strings readers parse, spelled here exactly as you will see them: | `engine` in the header | The model you actually have | |---|---| | `essence1` | [Essence 1](https://docs.bithuman.ai/models/first-generation#essence-1) — also the value an older container with no header resolves to | | `essence2-light` | **[Essence 2](https://docs.bithuman.ai/models/essence-2)** — request it as `essence-2` | | `essence2-quality` | Essence 2 Max (Enterprise plan only) — not a model you can request on other plans; treat the file as **[Essence 2](https://docs.bithuman.ai/models/essence-2)** | | `expression2` | **[Expression 2](https://docs.bithuman.ai/models/expression-2)** — request it as `expression-2` | So a current Essence 2 bundle reports `engine: essence2-light`. The model is **Essence 2**, requested as `essence-2`: the engine id names the *loader family*, not the product, so the value is expected, not a mismatch. > **Warning** Never send an engine id to the API. `model` takes `essence-2`, `expression-2`, `auto`, `essence-1` or `expression-1`; any other value returns [`400 VALIDATION_ERROR`](https://docs.bithuman.ai/api/agents#errors). ## File-format stability The `.imx` format is **forward-compatible within a major version**. The first open unpacks the file into that cache, using about its size again on disk; later opens reuse it. Your file is never rewritten. ## Where to go next - [Building avatars](https://docs.bithuman.ai/build/create-avatar) — design likeness, voice, and personality. - [Audio streaming](https://docs.bithuman.ai/models/how-it-works#audio-in-frames-out) — drive the `.imx` with audio. - [CLI reference](https://docs.bithuman.ai/platforms/cli) — `bithuman open`, `pull`, `list`, and more.