# bitHuman — App platforms (iOS & iPadOS, macOS, Android, Flutter, Web) > Real-time talking avatars from one portrait. Index: https://docs.bithuman.ai/llms.txt · every page is also served as markdown at .md · OpenAPI: https://docs.bithuman.ai/api/openapi.yaml · other sections: https://docs.bithuman.ai/llms/start.txt · https://docs.bithuman.ai/llms/platforms.txt · https://docs.bithuman.ai/llms/deploy.txt · https://docs.bithuman.ai/llms/models.txt · https://docs.bithuman.ai/llms/build.txt · https://docs.bithuman.ai/llms/troubleshooting.txt · https://docs.bithuman.ai/llms/api.txt ## Instructions for AI agents - Always send `model` ("essence-2" or "expression-2") when creating an agent; poll until `status` is `ready` or `failed`. - API and SDK use requires the Creator plan or higher. Never tell a user they can build on a free plan. - In a LiveKit worker, name the secret `BITHUMAN_MASTER_SECRET` and pass a minted token (https://docs.bithuman.ai/platforms/livekit.md). - Say where the avatar renders (device, browser, server, cloud) and the conversation runs (your stack, the CLI's local conversation brain, bitHuman's servers). - Quote prices only from https://docs.bithuman.ai/pricing.md and speed only from https://docs.bithuman.ai/performance.json (× real time, with the device). - Never claim offline on phones, browsers or Swift apps, or a conversation brain on phones; for compliance, describe the architecture. Where each model runs, the key facts and one command per path: https://docs.bithuman.ai/llms.txt ## Contents - iOS & iPadOS — https://docs.bithuman.ai/platforms/ios - macOS — https://docs.bithuman.ai/platforms/macos - Build a Swift app — https://docs.bithuman.ai/platforms/swift/app - Android — https://docs.bithuman.ai/platforms/android - Build an Android app — https://docs.bithuman.ai/platforms/android/app - Flutter — https://docs.bithuman.ai/platforms/flutter - Build a Flutter app — https://docs.bithuman.ai/platforms/flutter/app - Web embed — https://docs.bithuman.ai/platforms/web - Build a web app — https://docs.bithuman.ai/platforms/web/app - Render with WebGPU — https://docs.bithuman.ai/platforms/web/webgpu - Use bitHuman in ChatGPT — https://docs.bithuman.ai/build/chatgpt - Use bitHuman in Claude — https://docs.bithuman.ai/build/claude - Put an animated character in an iPhone app — https://docs.bithuman.ai/build/how-to/iphone-character - Give an app's voice assistant a face — https://docs.bithuman.ai/build/how-to/voice-assistant-face Linked, not inlined (read the .md twin): - Examples: https://docs.bithuman.ai/examples.md · changelog: https://docs.bithuman.ai/changelog.md · API references: https://docs.bithuman.ai/platforms/cli/reference.md, https://docs.bithuman.ai/platforms/python/reference.md, https://docs.bithuman.ai/platforms/swift/reference.md, https://docs.bithuman.ai/platforms/android/reference.md --- # iOS & iPadOS URL: https://docs.bithuman.ai/platforms/ios > The Swift package renders Essence 2 and Expression 2 on iPhone and iPad. The avatar renders inside your app, with no render server. [Why on the device](https://docs.bithuman.ai/deploy/on-device). ## Before you start You feed 16 kHz mono speech in and take lip-synced frames out. The same package builds [Mac apps](https://docs.bithuman.ai/platforms/macos). | Detail | Expression 2 | Essence 2 | |---|---|---| | **Renders** | [any character from one portrait](https://docs.bithuman.ai/models/expression-2) | [a photoreal person from one portrait](https://docs.bithuman.ai/models/essence-2) | | **Devices** | iPhone or iPad, iOS 16 or newer; measured on iPhone 15 only | iPhone, or an M-series iPad, iOS 26 or newer; measured on iPhone 15 only | | **Product** | `.product(name: "Expression2", package: "homebrew-bithuman")` | `.product(name: "Essence2Kit", package: "homebrew-bithuman")` (Swift), or `.product(name: "Essence2", package: "homebrew-bithuman")` (C) | | **Credential** | an [API secret](https://docs.bithuman.ai/start/api-secret), Creator plan or higher | an API secret, Creator plan or higher | | **First-run download** | about 370 MB (avatar and shared engine) | about 250 MB (avatar and engine resources) | | **Worked example** | [iOS Expression 2](https://docs.bithuman.ai/examples/ios-expression-2) | [iOS Essence 2](https://docs.bithuman.ai/examples/ios-essence-2) | - **Xcode 26 or newer** and an Apple Developer team. - **A paid plan** (Creator or higher): usage bills per second while the avatar runs ([pricing](https://docs.bithuman.ai/pricing)). - **A physical iPhone or iPad** for device builds. Essence 2 does not run in the Simulator: use a physical device. Expression 2 does run in the Simulator. - **Essence 1** is not available on phones or in the Swift package: use Essence 2 or Expression 2 on devices ([First generation](https://docs.bithuman.ai/models/first-generation)). ## Install In Xcode choose *File → Add Package Dependencies…* and paste `https://gitlab.com/bithuman/sdk/homebrew-bithuman`. In a `Package.swift`: ```swift .package(url: "https://gitlab.com/bithuman/sdk/homebrew-bithuman", from: "2.21.4") // then attach the product your target uses: // .product(name: "Expression2", package: "homebrew-bithuman") // .product(name: "Essence2Kit", package: "homebrew-bithuman") // .product(name: "Essence2", package: "homebrew-bithuman") ``` The products: | Product | Import | What it is | Deployment target | |---|---|---|---| | `Expression2` | `import Expression2` | the Expression 2 engine with a Swift API | iOS 16 · macOS 13 | | `Essence2Kit` | `import Essence2Kit` | the Essence 2 engine with a Swift API; it includes `Essence2` | iOS 26 · macOS 26 | | `Essence2` | `import Essence2` | the Essence 2 engine as a C library, for C, C++ and plugins | iOS 26 · macOS 26 | Every product ships `ios-arm64`, `ios-arm64-simulator` (arm64 only) and `macos-arm64`. The package lives in the `homebrew-bithuman` repository, so its package identity is `homebrew-bithuman`; the name is expected. The Simulator slices are arm64 only: for a `generic/platform=iOS Simulator` or other command-line build, set `EXCLUDED_ARCHS[sdk=iphonesimulator*] = x86_64` in your target's build settings. ## Authenticate The engines check an API secret when a session starts. Set `BITHUMAN_API_SECRET` in the scheme's environment, or pass it in code before you create an engine: `Expression2Credential.set(secret)` or `Essence2Credential.set(secret)` ([Your API secret](https://docs.bithuman.ai/start/api-secret)). The scheme's environment is for local builds. A shipped app fetches its secret from your backend and keeps it in the Keychain ([What a shipped app holds](https://docs.bithuman.ai/start/api-secret#what-a-shipped-app-holds)). ## Run your first avatar The quickest way to see it work is the example app: clone [bithuman-examples](https://gitlab.com/bithuman/sdk/bithuman-examples), open the [iOS Expression 2 example](https://docs.bithuman.ai/examples/ios-expression-2) in Xcode, set `BITHUMAN_API_SECRET` in the scheme and run it on your iPhone. The code below is the core of that app. Download the `wise-pup` sample avatar (agent code `A23WJF0199`), the shared Expression 2 engine and a 16 kHz speech clip. No account is needed for these downloads: ```bash curl -fL -o wise-pup.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2" curl -fLO "https://downloads.bithuman.ai/homebrew-bithuman/expression2-engine-mac-arm64-1.0.0/mac-arm64-1.0.0.engine" curl -fL -o speech16k.wav "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2&member=demo_speech_16k.wav" ``` Add them to your app, then render. The `mac` engine file is the right one for iPhone apps too. For Essence 2, download the avatar in the app with `Essence2Download` ([Swift: Integrate into your app](https://docs.bithuman.ai/platforms/swift/app#download-an-avatar-in-the-app)). ```swift tab="Expression 2" // excerpt: inside your app. samples is the clip as [Float], 16 kHz mono; // show(_:_:_:) draws B, G, R bytes; the URLs point at the files above. import Expression2 import AVFoundation Expression2Credential.set(ProcessInfo.processInfo.environment["BITHUMAN_API_SECRET"] ?? "") let engine = try Expression2Engine.create( avatarContainer: avatarURL, // wise-pup.imx sharedEngineContainer: sharedEngineURL, // mac-arm64-1.0.0.engine stagingDir: stagingURL) // any writable directory; keep it between launches let audio = AVAudioEngine(), player = AVAudioPlayerNode() // your app's audio output let format = AVAudioFormat(standardFormatWithSampleRate: 16000, channels: 1)! audio.attach(player); audio.connect(player, to: audio.mainMixerNode, format: format); try audio.start() let reply = AVAudioPCMBuffer(pcmFormat: format, frameCapacity: AVAudioFrameCount(samples.count))! reply.frameLength = reply.frameCapacity samples.withUnsafeBufferPointer { reply.floatChannelData![0].update(from: $0.baseAddress!, count: samples.count) } // seconds of the current reply your player has played (nil before it starts); frames() reads it // off the main actor, so the player goes in a Sendable box final class PlayedSeconds: @unchecked Sendable { let node: AVAudioPlayerNode init(_ node: AVAudioPlayerNode) { self.node = node } func callAsFunction() -> Double? { guard let t = node.lastRenderTime, let pt = node.playerTime(forNodeTime: t) else { return nil } return Double(pt.sampleTime) / pt.sampleRate } } let played = PlayedSeconds(player) engine.feed(samples) // [Float], 16 kHz mono engine.flushTail() // end of the reply for await frame in engine.frames(audioClock: { played() }) { show(frame.bgr, frame.width, frame.height) // B, G, R bytes if frame.audioTime == 0 { player.scheduleBuffer(reply, completionHandler: nil); player.play() } // the reply's first frame: start its audio if frame.endsReply { break } // the reply is over; idle frames follow } ``` ```swift tab="Essence 2" // excerpt: inside your app. samples is the clip as [Float], 16 kHz mono; // show(_:_:_:) draws B, G, R bytes. import Essence2Kit import AVFoundation Essence2Credential.set(ProcessInfo.processInfo.environment["BITHUMAN_API_SECRET"] ?? "") let imxURL = try await Essence2Download.identity(agentCode: "A52DHS2219") // sofia-ramirez let engine = try await Essence2Engine.create(identity: imxURL) // waits until the engine is ready let audio = AVAudioEngine(), player = AVAudioPlayerNode() // your app's audio output let format = AVAudioFormat(standardFormatWithSampleRate: 16000, channels: 1)! audio.attach(player); audio.connect(player, to: audio.mainMixerNode, format: format); try audio.start() let reply = AVAudioPCMBuffer(pcmFormat: format, frameCapacity: AVAudioFrameCount(samples.count))! reply.frameLength = reply.frameCapacity samples.withUnsafeBufferPointer { reply.floatChannelData![0].update(from: $0.baseAddress!, count: samples.count) } engine.feed(samples) // [Float], 16 kHz mono engine.flushTail() // that is the whole reply for await frame in engine.frames(following: player) { show(frame.bgr, frame.width, frame.height) // B, G, R bytes, width * height * 3 if frame.audioTime == 0 { // the reply's first speech frame: player.stop(); player.scheduleBuffer(reply, completionHandler: nil); player.play() // start its audio now } if frame.endsReply { break } // the reply is over; idle frames follow } engine.shutdown() ```
Expected result - **Expression 2:** 416×720 frames, 20 a second: idle motion between replies, speech while a reply plays. Each speech frame is handed out when your player has played its audio. The first start prepares the engine for the device; later starts reuse the staging directory. - **Essence 2:** 25 frames a second at the avatar's own size (1080×1920 for `sofia-ramirez`). `frames(following: player)` keeps voice and lips together however long the reply is; a frame that would be shown late is skipped. The first `create` downloads the engine's three runtime files (about 112 MB), checks their sha256 and keeps them in Application Support. To ship them in your app instead, pass `resourcesDirectory:`.
The C interface for C, C++ and plugins (`Essence2`) is on the [Swift reference](https://docs.bithuman.ai/platforms/swift/reference#essence-2-c).
*Capture: The wise-pup avatar mid-sentence, as drawn by the ios-expression2 example app on an iPhone 15. Captured on iPhone 15 (iOS 26) · Swift package 2.14.2 · wise-pup (Expression 2) · 2026-09-23.* (https://docs.bithuman.ai/examples/ios-expression-2/poster.webp) ## Performance | Configuration | Hardware | Essence 2 | Expression 2 | |---|---|---|---| | iPhone · Swift package | iPhone 15 | 2.1× real time | 5.5× real time | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance)). --- # macOS URL: https://docs.bithuman.ai/platforms/macos > Render Essence 2 and Expression 2 on a Mac with Apple silicon. The iPhone and iPad Swift package, in a Mac app or from a terminal with `swift run`. ## Before you start No device, provisioning profile or entitlement is needed to try it from a terminal. [Why render on the device](https://docs.bithuman.ai/deploy/on-device). | You need | Expression 2 | Essence 2 | |---|---|---| | **A Mac** | Apple silicon, macOS 13 or newer | Apple silicon M3 or newer, macOS 26 or newer | | **Toolchain** | Xcode 26 or newer (to build; Xcode 26 needs macOS 15.6 or newer) | the same | | **Credential** | an [API secret](https://docs.bithuman.ai/start/api-secret) (Creator plan or higher) | the same | | **Disk** | about 800 MB for the example: downloads plus the folder the engine unpacks them into | about 250 MB of downloads | ## Install In Xcode choose *File → Add Package Dependencies…* and paste `https://gitlab.com/bithuman/sdk/homebrew-bithuman`. In a `Package.swift`: ```swift .package(url: "https://gitlab.com/bithuman/sdk/homebrew-bithuman", from: "2.21.4") // then attach the product your target uses: // .product(name: "Expression2", package: "homebrew-bithuman") // .product(name: "Essence2Kit", package: "homebrew-bithuman") // .product(name: "Essence2", package: "homebrew-bithuman") ``` The products: | Product | Import | What it is | Deployment target | |---|---|---|---| | `Expression2` | `import Expression2` | the Expression 2 engine with a Swift API | iOS 16 · macOS 13 | | `Essence2Kit` | `import Essence2Kit` | the Essence 2 engine with a Swift API; it includes `Essence2` | iOS 26 · macOS 26 | | `Essence2` | `import Essence2` | the Essence 2 engine as a C library, for C, C++ and plugins | iOS 26 · macOS 26 | Every product ships `ios-arm64`, `ios-arm64-simulator` (arm64 only) and `macos-arm64`. ## Authenticate The engines check an API secret when a session starts. Set `BITHUMAN_API_SECRET` in the scheme's environment, or pass it in code before you create an engine: `Expression2Credential.set(secret)` or `Essence2Credential.set(secret)` ([Your API secret](https://docs.bithuman.ai/start/api-secret)). The scheme's environment is for local builds. A shipped app fetches its secret from your backend and keeps it in the Keychain ([What a shipped app holds](https://docs.bithuman.ai/start/api-secret#what-a-shipped-app-holds)). A Mac app built from Xcode's App template turns on App Sandbox. Under *Signing & Capabilities → App Sandbox*, tick **Outgoing Connections (Client)**, or the engines cannot check your secret. ## Run your first avatar The [macOS Expression 2 example](https://docs.bithuman.ai/examples/macos-expression-2) is one Swift file. Clone it, fetch the sample avatar, and run it: ```bash git clone https://gitlab.com/bithuman/sdk/bithuman-examples.git cd bithuman-examples/swift/macos-expression2 ./setup.sh # the wise-pup avatar, the shared engine and a speech clip export BITHUMAN_API_SECRET="" swift run -c release MacOSExpression2 ```
Expected result ```text engine ready: 416x720, isReady=true audio: 325451 samples, 20.34 s generated 407 frames in … s -> out/first-frame.png ``` 407 frames for 20.34 seconds of audio: one frame per 50 ms of speech. The first run prepares the engine for your Mac; keep `Model/staged/` and later runs start faster.
The core of `Sources/main.swift`: ```swift // excerpt: swift/macos-expression2/Sources/main.swift let engine = try Expression2Engine.create( avatarContainer: model.appendingPathComponent("agent.imx"), sharedEngineContainer: model.appendingPathComponent("shared-engine.imx"), stagingDir: model.appendingPathComponent("staged")) // … let started = Date() engine.feed(samples) engine.flushTail() // … // Time to the LAST frame, not to the end of the idle drain below (that would add 5 s). var frames = 0, idleTicks = 0, lastAt = started while idleTicks < 100 { var got = false while let (frame, _) = engine.pull() { if frames == 0 { writePNG(frame, width: engine.width, height: engine.height, to: out.appendingPathComponent("first-frame.png")) } frames += 1 lastAt = Date() got = true } if got { idleTicks = 0 } else { idleTicks += 1; usleep(50_000) } } ``` ### Essence 2 on a Mac Essence 2 needs an M3 or newer Mac on macOS 26. The same `Essence2Kit` calls as on the iPhone render a file from a terminal tool: download the avatar in code, set `pacing = .unpaced`, and take the frames as fast as they render. ```swift // excerpt: a terminal tool (swift run) with the Essence2Kit product; // samples is a 16 kHz mono clip as [Float]. import Essence2Kit Essence2Credential.set(ProcessInfo.processInfo.environment["BITHUMAN_API_SECRET"] ?? "") let imxURL = try await Essence2Download.identity(agentCode: "A52DHS2219") // sofia-ramirez let engine = try await Essence2Engine.create(identity: imxURL) // waits until the engine is ready engine.pacing = .unpaced // a file render: frames as fast as they render engine.feed(samples) // [Float], 16 kHz mono engine.flushTail() // that is the whole reply var speech = 0 for await frame in engine.frames() { if frame.isSpeech { speech += 1 } // frame.bgr: B, G, R bytes, width * height * 3 if frame.endsReply { break } // the reply is over; idle frames follow } print("\(engine.width)x\(engine.height): \(speech) frames for \(Double(samples.count) / 16_000) s of audio") engine.shutdown() ```
Expected result With the [15-second sample clip](https://docs.bithuman.ai/samples/speech-16k.wav): ```text 1080x1920: 375 frames for 14.997375 s of audio ``` 25 frames a second at the avatar's own size. The first `create` downloads the avatar and the engine's runtime files, then keeps them in Caches and Application Support. The engine also prints its own diagnostic lines on stderr.
*Capture: The wise-pup avatar speaking, rendered by the macos-expression2 example. Captured on an Apple M4 iMac (macOS) · Swift package 2.14.2 · wise-pup (Expression 2) · 2026-09-23.* (https://docs.bithuman.ai/examples/macos-expression-2/clip.mp4) ## Performance | Configuration | Hardware | Essence 2 | Expression 2 | |---|---|---|---| | macOS · Swift package | Apple M4 iMac (Mac16,3, 10-core GPU) | 4.3× real time | 8.8× real time | | macOS · CLI | Apple M4 | 4.2× real time | 8.4× real time | | macOS · Python | Apple M4 | 6.9× real time | 8.4× real time | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance)). --- # Build a Swift app URL: https://docs.bithuman.ai/platforms/swift/app > Stream audio into a Swift avatar on iPhone, iPad or Mac. ## Integrate into your app | Job | Expression 2 | Essence 2 | |---|---|---| | Audio in | 16 kHz mono `[Float]`: `feed(chunk)` as it arrives | the same | | Show frames | `frames(audioClock:)` on your player's clock, or `pull()`, which returns frames as soon as they render | `frames(following:)`, or `pull()` paced to 25 a second | | End of a reply | `flushTail()`; the first frame after it has `endsReply`, and `events()` reports `.replyEnded` | the same | | Start the reply's audio | with its first frame (`audioTime == 0`, or `events()` `.replyStarted`) | with its first speech frame; `frames(following: player)` keeps the picture on it | | Idle between replies | `frames()` keeps returning idle frames (`isSpeech == false`), or `engine.idle` | `frames()` / `pull()` keep returning idle frames | | Interrupt the reply | `interrupt()` | `interrupt()` | | Check the session | `meteringRefusal` | `meteringRefusal`, `runtimeFailure` | | Quit | `shutdown()` | `shutdown()`, then `Essence2Engine.quiesceAll()` from `applicationWillTerminate` | For file rendering, set `engine.pacing = .unpaced`: Essence 2 then hands out frames as fast as it renders them. If your audio does not go through an `AVAudioPlayerNode`, pass your own clock: `frames(audioClock: { secondsOfThisReplyPlayed })`. Resample 24 kHz speech (OpenAI Realtime's) to 16 kHz, and close the avatar when the app leaves the screen: [Companion app](https://docs.bithuman.ai/build/companion-app#resample-speech-to-16-khz). ### Download an avatar in the app Your app can download an avatar file itself, with the secret you set in [Authenticate](https://docs.bithuman.ai/platforms/ios#authenticate): ```swift let avatarURL = try await Expression2Download.avatar(agentCode: "A23WJF0199") // Expression 2 let imxURL = try await Essence2Download.identity(agentCode: "A52DHS2219") // Essence 2 ``` Both return a local file to pass to `create`. They download the Apple build of the avatar, which is smaller than the full file, and refuse a file whose sha256 does not match. Files are kept in the app's Caches directory under their sha256, so a second call for the same avatar downloads nothing. Pass `directory:` to keep them somewhere else. The shared Expression 2 engine file is not an avatar; download it from the release as in [First frame](https://docs.bithuman.ai/platforms/ios#run-your-first-avatar). ### On a Mac The Swift API is the same on the Mac as on iPhone and iPad: feed 16 kHz mono audio, take frames on your player's clock, end and interrupt replies. The whole table is [above](#integrate-into-your-app); every entry point is on the [Swift reference](https://docs.bithuman.ai/platforms/swift/reference). On a Mac: - **Files:** add the `.imx` files and engine resources to the app bundle. A sandboxed app reads only its bundle and container. - **App Sandbox:** tick **Outgoing Connections (Client)** so the engine can check your secret, and **Audio Input** (`com.apple.security.device.audio-input`) if the app uses the microphone. - **Quitting:** call `Essence2Engine.quiesceAll()` from `applicationWillTerminate`. ## Complete example Two SwiftUI apps you can clone and run on an iPhone or iPad, each with a microphone button, idle motion and interruption: - [iOS Expression 2](https://docs.bithuman.ai/examples/ios-expression-2): the `wise-pup` sample avatar. - [iOS Essence 2](https://docs.bithuman.ai/examples/ios-essence-2): a photoreal Essence 2 avatar at full resolution. [macOS Expression 2](https://docs.bithuman.ai/examples/macos-expression-2) walks through the tool above: requirements, your own avatar and audio, and troubleshooting. For a window with a microphone button, the [iOS Expression 2 example](https://docs.bithuman.ai/examples/ios-expression-2) is the same engine in a SwiftUI app. ## Platform notes - **Your own MLX:** Essence 2 contains no MLX. Link your own `mlx-swift` (`MLX`, `MLXNN`) in the same target, also with `-ObjC` or `-all_load`; nothing to embed. - **Simulator:** simulator slices are arm64 only; pass `ARCHS=arm64`, or set `EXCLUDED_ARCHS[sdk=iphonesimulator*] = x86_64` in the target. Essence 2 does not run in the Simulator: use a physical device. Expression 2 does run in the Simulator. - **Privacy strings:** add `NSMicrophoneUsageDescription` to hear the user. - **Check the version you resolved.** SwiftPM keeps what `Package.resolved` holds, so run `swift package update` after you raise `from:`, then read it back: ```bash grep -A3 homebrew-bithuman Package.resolved # "version" must be the one on Downloads & versions ``` - **Also on a Mac:** the [CLI](https://docs.bithuman.ai/platforms/cli) renders an avatar or runs a live conversation with no code, and [Python](https://docs.bithuman.ai/platforms/python) renders frames from your own scripts. Both run on Apple silicon. - **Intel Macs** are not supported. ## Reference - [Swift reference](https://docs.bithuman.ai/platforms/swift/reference): every Swift and C entry point. - Examples: [iOS Expression 2](https://docs.bithuman.ai/examples/ios-expression-2) · [iOS Essence 2](https://docs.bithuman.ai/examples/ios-essence-2) · [macOS Expression 2](https://docs.bithuman.ai/examples/macos-expression-2). - Sample avatars: [Ready-made avatars](https://docs.bithuman.ai/examples/avatars). Your own agent's model: [`GET /v1/agent/{code}/model/download`](https://docs.bithuman.ai/api/agents#download-an-agents-model) with your API secret. - [Changelog](https://docs.bithuman.ai/changelog) and [Downloads & versions](https://docs.bithuman.ai/downloads). - [macOS Expression 2 example](https://docs.bithuman.ai/examples/macos-expression-2) and its [source on GitLab](https://gitlab.com/bithuman/sdk/bithuman-examples/-/tree/main/swift/macos-expression2). --- # Android URL: https://docs.bithuman.ai/platforms/android > Render Essence 2 and Expression 2 on Android phones, from bitHuman's Maven repository. After the one-time model download, the only network traffic is usage reporting. [Why on the device](https://docs.bithuman.ai/deploy/on-device). ## Before you start You feed 16 kHz mono speech in and pull picture frames out. Each model is one Gradle dependency, from bitHuman's Maven repository at `maven.bithuman.ai`. | Detail | Expression 2 | Essence 2 | |---|---|---| | **Renders** | [any character from one portrait](https://docs.bithuman.ai/models/expression-2) | [a photoreal person from one portrait](https://docs.bithuman.ai/models/essence-2) | | **Devices** | a physical `arm64-v8a` phone, `minSdk 26` | a physical `arm64-v8a` phone, `minSdk 29` | | **Dependency** | `implementation("ai.bithuman:expression2-android:0.6.2")` | `implementation("ai.bithuman:essence2-android:0.10.0")` | | **Credential** | an [API secret](https://docs.bithuman.ai/start/api-secret), Creator plan or higher | an API secret, Creator plan or higher | | **First-run download** | about 160 MB | 226–281 MB | | **Adds to your APK** | about 3 MB, plus a 70 MB accelerator runtime you can leave out | 12.1 MB | | **Worked example** | [Android Expression 2](https://docs.bithuman.ai/examples/android-expression-2) | [Android Essence 2](https://docs.bithuman.ai/examples/android-essence-2) | - **JDK 17, Gradle 8.11 or newer and Android Gradle Plugin 8.7 or newer.** - **A paid plan** (Creator or higher): usage bills per second while the avatar runs ([pricing](https://docs.bithuman.ai/pricing)). - **A physical arm64 phone.** Emulators cannot load the engines. - **Frame rate depends on the phone.** Essence 2 plays at 25 frames a second, and the phone has to render at least that fast to keep up with live speech. Recent flagship chips render well above that; older ones, such as the Snapdragon 8 Gen 2 in a Galaxy Z Flip5, can fall below real time, more so as the phone warms. The [performance page](https://docs.bithuman.ai/performance) has the measured rates. Test on the phones your app targets. - **Essence 1** is not available on phones: use Essence 2 or Expression 2 on devices ([First generation](https://docs.bithuman.ai/models/first-generation)). ## Install Add bitHuman's Maven repository for the `ai.bithuman` group, restrict the build to `arm64-v8a`, and turn on legacy packaging so the engines' native libraries are extracted to disk. Merge the lines into your existing app module, which keeps its own `namespace`, `compileSdk` and Java 17 targets. The API secret reaches your code through `BuildConfig`. ```kotlin // settings.gradle.kts dependencyResolutionManagement { repositories { google() mavenCentral() exclusiveContent { // ai.bithuman resolves from bitHuman's repository only forRepository { maven { url = uri("https://maven.bithuman.ai") } } filter { includeGroup("ai.bithuman") } } } } // app/build.gradle.kts // bithuman.apiSecret=… in local.properties (git-ignored), or BITHUMAN_API_SECRET in the environment val bithumanApiSecret: String = run { val props = java.util.Properties() // fully qualified, so these lines can go anywhere in the file val f = rootProject.file("local.properties") if (f.isFile) f.inputStream().use { props.load(it) } props.getProperty("bithuman.apiSecret") ?: System.getenv("BITHUMAN_API_SECRET") ?: "" } android { buildFeatures { buildConfig = true } defaultConfig { minSdk = 26 // 29 for Essence 2 ndk { abiFilters += "arm64-v8a" } buildConfigField("String", "BITHUMAN_API_SECRET", "\"$bithumanApiSecret\"") } packaging { jniLibs { useLegacyPackaging = true } } // required } dependencies { implementation("ai.bithuman:expression2-android:0.6.2") // or: implementation("ai.bithuman:essence2-android:0.10.0") } ``` Put the secret in `local.properties` as `bithuman.apiSecret=…` (Android Studio keeps that file out of git), or export `BITHUMAN_API_SECRET` before you build; the example apps read it the same way. An empty secret makes the session refuse to start. The build prints `Unable to strip the following libraries, packaging them as they are: libLiteRt.so, libQnn…`. It is expected when no NDK is installed. The `dependencyResolutionManagement` block works unchanged in a Groovy `settings.gradle`. Only `ai.bithuman` comes from `maven.bithuman.ai`; the engines' own dependencies still come from `google()` and `mavenCentral()`. `expression2-android` brings the Qualcomm accelerator runtime with it (`com.qualcomm.qti:qnn-litert-delegate:2.49.0` and `com.qualcomm.qti:qnn-runtime:2.49.0`). To keep the APK small and render on the CPU instead, exclude it: ```kotlin implementation("ai.bithuman:expression2-android:0.6.2") { exclude(group = "com.qualcomm.qti") } ``` ## Authenticate Pass your API secret ([create one](https://www.bithuman.ai/developer/api-keys)) in code before you download or create an avatar: `Expression2Credential.set(secret)` for Expression 2, `Essence2Credential.set(secret)` for Essence 2. That one call covers the download and the session. See [Your API secret](https://docs.bithuman.ai/start/api-secret). > **Warning:** a `buildConfigField` compiles the secret into the APK, where anyone with the file can read it. Use it for local builds only. A shipped app fetches its secret from your backend ([What a shipped app holds](https://docs.bithuman.ai/start/api-secret#what-a-shipped-app-holds)). ## Run your first avatar Expression 2, with the published `wise-pup` avatar (agent code `A23WJF0199`). Call `render` off the main thread. ```kotlin import ai.bithuman.expression2.Expression2Avatar import ai.bithuman.expression2.Expression2Credential import ai.bithuman.expression2.Expression2ModelStore import ai.bithuman.expression2.Expression2Options import android.content.Context import android.graphics.Bitmap /** [pcm16k] is 16 kHz mono float32 in [-1, 1]. */ fun render(context: Context, pcm16k: FloatArray, show: (Bitmap) -> Unit) { Expression2Credential.set(BuildConfig.BITHUMAN_API_SECRET) // before fetch() and create() val model = Expression2ModelStore(context).fetch("A23WJF0199") // ~160 MB, first run only Expression2Avatar.create(context, model, Expression2Options()).use { avatar -> val frame = avatar.newFrameBitmap() // 416 x 720, allocate once avatar.feed(pcm16k) avatar.flushTail() // end of the utterance while (true) { if (avatar.pull(frame) != null) { show(frame); continue } if (!avatar.hasPendingTail && avatar.queuedFrames == 0) break Thread.sleep(10) // null means "not ready yet" } } } ``` Expected: `show` receives 20 frames for each second of audio, and the avatar's lips follow the speech. The first `create()` after install prepares the accelerator once and takes noticeably longer than later launches, which reuse it. Create once, at app start, on a background thread. Essence 2 takes 16-bit little-endian PCM bytes, as a 16 kHz mono WAV stores them, and fills an RGBA `ByteBuffer` sized from the identity: ```kotlin import ai.bithuman.essence2.Essence2Avatar import ai.bithuman.essence2.Essence2Credential import ai.bithuman.essence2.Essence2ModelStore import android.content.Context import java.nio.ByteBuffer fun render(context: Context, pcm16le: ByteArray, show: (ByteBuffer, Int, Int) -> Unit) { Essence2Credential.set(BuildConfig.BITHUMAN_API_SECRET) // before fetch() and create() val identity = Essence2ModelStore(context).fetch("A52DHS2219") // sofia-ramirez; 226–281 MB, first run only Essence2Avatar.create(identity.dir).use { avatar -> val frame = avatar.newFrameBuffer() // width * height * 4, RGBA avatar.feed(pcm16le) avatar.endOfAudio() var quietMs = 0 while (quietMs < 5_000) { // 5 s with no frame at all = finished if (avatar.pull(frame)) { show(frame, avatar.width, avatar.height); quietMs = 0; continue } Thread.sleep(10) // false means "not ready yet"; the first frame can take seconds quietMs += 10 } avatar.checkRender() // throws Essence2RenderFailed if the engine stopped // a refused session throws Essence2MeteringRefused from pull() or idle() } } ``` Frame size belongs to the identity (portrait 1080×1920, landscape 1920×1080 or 1280×720). Read `avatar.width` and `avatar.height`; do not hard-code them.
*Capture: The sofia-ramirez avatar speaking in the essence2-hello app on a Galaxy S25+, at 1080×1920. Captured on Samsung Galaxy S25+ (Android) · essence2-android 0.5.14 · sofia-ramirez (Essence 2) · 2026-09-23.* (https://docs.bithuman.ai/examples/android-essence-2/clip.mp4) ## Performance | Configuration | Hardware | Essence 2 | Expression 2 | |---|---|---|---| | Android | Samsung Galaxy S25+ | 2.0× real time | 2.4× real time | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance)). --- # Build an Android app URL: https://docs.bithuman.ai/platforms/android/app > Stream audio into an Android avatar and show its frames in your app. ## Integrate into your app In a live conversation, keep one avatar open and stream into it. | Job | Expression 2 | Essence 2 | |---|---|---| | Audio in | 16 kHz mono `FloatArray`, −1 to 1 | 16 kHz mono 16-bit little-endian PCM `ByteArray` | | Stream audio as it arrives | `feed(chunk)` per chunk | `feed(chunk)` per chunk | | Show frames | `pull(bitmap)`, 20 a second | `pull(buffer)`, 25 a second | | End of a reply | `flushTail()` | `endOfAudio()` | | Idle between replies | `avatar.idleLoop?.next(bitmap)` | `idle(buffer)` | | Interrupt the reply | `resetState(true)` | `resetAudio()` | | Check the session | `Expression2Exception` from `create` or `pull` | `Essence2MeteringRefused` from `pull`/`idle`; `checkRender()` throws `Essence2RenderFailed` if the engine stopped | | Close it | `close()` | `close()` | To feed the microphone, declare `` in your manifest and ask for it at runtime (`ActivityResultContracts.RequestPermission`) before you start `AudioRecord`; without it the buffers are silent. The SDKs need only `INTERNET`, which their AARs merge in for you. Resample 24 kHz speech (OpenAI Realtime's) to 16 kHz, and close the avatar when the app leaves the screen: [Companion app](https://docs.bithuman.ai/build/companion-app#resample-speech-to-16-khz). After your API secret is accepted, a network loss does not stop the session for 5 minutes of rendered video. After that, render calls throw a retryable exception until the connection returns. Usage is reported to your account when it does. For a Flutter app, the [Flutter plugin](https://docs.bithuman.ai/platforms/flutter) wraps these engines. ## Complete example Two apps you can clone and run on a phone, each with idle motion between replies: - [Android Expression 2](https://docs.bithuman.ai/examples/android-expression-2): the `wise-pup` sample avatar. - [Android Essence 2](https://docs.bithuman.ai/examples/android-essence-2): the `sofia-ramirez` sample avatar at full resolution. ## Platform notes - **Release builds:** `isMinifyEnabled = true` needs nothing extra. Both AARs ship their own keep rules. - **Two models in one app:** `essence2-android` needs `minSdk 29`. Raise the app to 29, or put each model in its own module. - **Threads:** download and `create()` on a background thread. The first download is the size shown above, into app-private storage. - **Check the version you resolved:** Gradle keeps an exact version, so read it back when behaviour differs from this page. ```bash ./gradlew :app:dependencies --configuration releaseRuntimeClasspath | grep ai.bithuman ``` - **Private avatars:** an avatar you created downloads with the secret you set with `Expression2Credential.set` or `Essence2Credential.set`; there is nothing else to pass. ## Reference - [Android API reference](https://docs.bithuman.ai/platforms/android/reference): every public class in both AARs. - Examples: [Expression 2](https://docs.bithuman.ai/examples/android-expression-2) · [Essence 2](https://docs.bithuman.ai/examples/android-essence-2), complete apps you can clone. - [Flutter](https://docs.bithuman.ai/platforms/flutter): the Flutter plugin, built on these engines. - [Changelog](https://docs.bithuman.ai/changelog) and [Downloads & versions](https://docs.bithuman.ai/downloads). - FFmpeg in `essence2-android` is LGPL: [relink materials](https://docs.bithuman.ai/legal/android-ffmpeg-lgpl). --- # Flutter URL: https://docs.bithuman.ai/platforms/flutter > The Flutter plugin renders Essence 2 and Expression 2 on Android, iOS and macOS. One Flutter dependency gives your app an avatar widget. [Why on the device](https://docs.bithuman.ai/deploy/on-device). ## Before you start On Android the plugin runs the same engines as the [Android SDK](https://docs.bithuman.ai/platforms/android), so the avatar renders on the phone: your app's speech audio goes in, and a lip-synced picture comes out as a Flutter `Texture`. The plugin runs on Android phones (arm64), and on iPhone, iPad and Mac (Apple silicon) after one bootstrap step that fetches the engines ([Platform notes](https://docs.bithuman.ai/platforms/flutter/app#platform-notes)). A native iPhone or iPad app can use the [Swift package](https://docs.bithuman.ai/platforms/ios) instead. | You need | Notes | |---|---| | Dart 3.11.5 or newer (the Flutter release that ships it), with the Android toolchain | JDK 17 and the Android SDK; an older Dart fails `flutter pub get` | | A physical Android phone, arm64, Android 10 (API 29) or newer | emulators cannot load the engines; the plugin needs API 29 whichever model you use | | An [API secret](https://docs.bithuman.ai/start/api-secret) | the Creator plan or higher; usage bills per second while the avatar runs ([pricing](https://docs.bithuman.ai/pricing)) | ## Install The plugin is published as a tag in the public `homebrew-bithuman` repository on GitLab. Pin it in `pubspec.yaml`: ```yaml dependencies: bithuman: git: url: https://gitlab.com/bithuman/sdk/homebrew-bithuman.git path: packages/flutter-plugin ref: flutter-plugin-v2.6.40 ``` Then run `flutter pub get`. On Android, Gradle resolves `ai.bithuman:essence2-android` and `ai.bithuman:expression2-android` from the repository the plugin declares, so you add no repository yourself. In `android/app/build.gradle.kts`, set the plugin's floor and keep the engines' native libraries as they ship: ```kotlin android { defaultConfig { minSdk = 29 // the plugin's floor, whichever model you use ndk { abiFilters += "arm64-v8a" } // the engines ship arm64-v8a only } packaging { jniLibs { useLegacyPackaging = true } } // required } ``` ## Authenticate The engines check your API secret when an avatar loads: pass it to `BithumanAvatar.load(…, apiSecret:)`. For a local build, the example app asks for it on first launch and keeps it in the Android Keystore, or you pass `--dart-define=BITHUMAN_API_SECRET=…` when you build. A shipped app fetches the secret from your backend ([What a shipped app holds](https://docs.bithuman.ai/start/api-secret#what-a-shipped-app-holds)). ## Run your first avatar Build the example app for your phone, with the `wise-pup` sample avatar: ```bash git clone https://gitlab.com/bithuman/sdk/bithuman-examples.git cd bithuman-examples/app/avatar_chat flutter pub get flutter run --release --dart-define=AGENT_CODE=A23WJF0199 ```
Expected result The app installs on the connected phone, asks for your API secret once, downloads the avatar (about 160 MB, first run only) and shows it idling full screen. Speak, or type a line, and it answers with its lips in sync.
## Performance The plugin runs the Android SDK's own engines (the same native libraries), so the Android rows apply: | Configuration | Hardware | Essence 2 | Expression 2 | |---|---|---|---| | Android | Samsung Galaxy S25+ | 2.0× real time | 2.4× real time | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance)). --- # Build a Flutter app URL: https://docs.bithuman.ai/platforms/flutter/app > Stream audio into a Flutter avatar and show its frames in your app. ## Integrate into your app The avatar is a `Texture` in your widget tree. Pass `engine:` every time, `'expression2'` or `'essence2'`: leave it out and Android refuses the load (`unsupported`). On Android the first argument to `load` is the agent code; the plugin downloads the avatar and keeps it. ```dart import 'package:bithuman/bithuman.dart'; final avatar = await BithumanAvatar.load( 'A23WJF0199', // on Android, the agent code engine: 'expression2', // required: 'expression2' or 'essence2' apiSecret: secret, ); Texture(textureId: avatar.textureId); // the avatar in your layout await avatar.audioStart(enableMic: false); // the speaker; required on iOS and macOS await avatar.playSpeakerPCM(chunk); // Uint8List, 24 kHz mono PCM16: heard, and the lips follow await avatar.notifyTurnEnd(); // after the reply's last chunk await avatar.interrupt(); // cut the current reply await avatar.dispose(); // release the engine ``` | Job | Call | |---|---| | Show the avatar | `Texture(textureId: avatar.textureId)`; `frameWidth` and `frameHeight` give its size | | Start the speaker | `audioStart(enableMic: false)`, once, before the first chunk | | Play your speech | `playSpeakerPCM(Uint8List)`, 24 kHz mono PCM16, chunk by chunk; the same audio drives the lips | | End a reply | `notifyTurnEnd()` after the last chunk, so the last word is not clipped | | Know it is ready | `isReady`, or `await avatar.ready`; audio sent before it is dropped | | Interrupt the reply | `interrupt()` | | Stop | `audioStop()`, then `dispose()` | | Pick the model | `engine: 'expression2'` or `engine: 'essence2'` (required) | **A voice conversation instead of your own audio:** `BithumanRealtimeSession(apiKey: secret, avatar: avatar, model: 'gpt-realtime-mini')` from `package:bithuman/bithuman_realtime.dart` connects the avatar to bitHuman's [realtime relay](https://docs.bithuman.ai/api/realtime) with your API secret, billed 10 credits a minute with the avatar included. Show `spokenTranscriptStream` as captions and handle `errorStream` ([Flutter errors](https://docs.bithuman.ai/platforms/flutter/errors)). ## Complete example The [`avatar_chat` app](https://gitlab.com/bithuman/sdk/bithuman-examples/-/tree/main/app/avatar_chat) is a complete voice conversation with idle motion and interruption, in one layout for every platform. From plugin 2.6.20 its voice session connects through bitHuman's [realtime relay](https://docs.bithuman.ai/api/realtime) with your API secret; no token is minted. ## Platform notes - **Android build settings:** the plugin needs `minSdk 29` whichever model you use, `arm64-v8a` only, and `useLegacyPackaging = true` ([Install](https://docs.bithuman.ai/platforms/flutter#install)). - **Your own voice pipeline:** any speech your stack produces works: resample it to 24 kHz mono PCM16 and send it through `playSpeakerPCM`, as above. `pushAudio` is not the call for this: on iOS and macOS it moves the lips with no sound (on Android it plays from plugin 2.6.36). - **Sign-out:** call `BithumanAvatar.clearCredentials()` when an account signs out (plugin 2.6.36 or newer). The engines forget the secret, and on Android every load still running ends with `load_cancelled`. On iOS and macOS, set the Expression 2 agent folder again (`setExpression2AgentDir`) before the next load that needs one. - **Errors:** what `load` throws and why a voice session ends are on [Flutter errors](https://docs.bithuman.ai/platforms/flutter/errors). - **Captions:** with the plugin's realtime session, show `spokenTranscriptStream` rather than `botTranscriptStream`. It releases the agent's words as they are heard, so the caption keeps pace with the voice; each event holds the reply's caption so far, and a cut reply ends with only the words heard (plugin 2.6.27 or newer). - **iOS and macOS (Apple silicon) builds:** run the plugin's `scripts/bootstrap.sh` once (it downloads the published engines and checks their sha256; for a git dependency the plugin's folder is `packages/flutter-plugin` under `~/.pub-cache/git/homebrew-bithuman-…`) and raise the deployment targets (`platform :ios, …` in `ios/Podfile`, `platform :osx, …` in `macos/Podfile`, and the Runner targets to match). Each model has its own minimum: Expression 2 renders on iOS 18 and macOS 15 and later, so an app that renders it sets iOS 18.0 and macOS 15.0; Essence 2 needs iOS 26 and macOS 26. Below a model's minimum, `load` fails with `unsupported` and names it ([Flutter errors](https://docs.bithuman.ai/platforms/flutter/errors#loading-an-avatar)). The pod itself builds from iOS 16.4 and macOS 14.0. Nothing is installed with Homebrew: from plugin 2.6.39 the app links the plugin's own ONNX Runtime, so do not add another ONNX Runtime to it. - **Versions:** each plugin tag fixes the Android SDK versions it uses. The current tag and its line are on [Downloads & versions](https://docs.bithuman.ai/downloads). ## Reference - [Flutter errors](https://docs.bithuman.ai/platforms/flutter/errors): the codes `load` throws and the codes a voice session ends with. - [Android](https://docs.bithuman.ai/platforms/android): the engines under the plugin, their API and their settings. - [`avatar_chat` example](https://gitlab.com/bithuman/sdk/bithuman-examples/-/tree/main/app/avatar_chat) and the [plugin source](https://gitlab.com/bithuman/sdk/homebrew-bithuman/-/tree/main/packages/flutter-plugin). - [Changelog](https://docs.bithuman.ai/changelog) and [Downloads & versions](https://docs.bithuman.ai/downloads). --- # Web embed URL: https://docs.bithuman.ai/platforms/web > Put a live, talking avatar on any web page with one iframe. One URL in an ` ``` Expected: the avatar's picture with **Tap to talk**. Select it, allow the microphone, and the avatar greets you and answers when you speak. Keep the `*` in `allow`, or the microphone is blocked. To try it without a page, open [https://www.bithuman.ai/embed/A23WJF0199](https://www.bithuman.ai/embed/A23WJF0199). The `wise-pup` sample ends each session after 1 minute. For your site, embed your own agent ([create your own avatar](https://docs.bithuman.ai/build/create-avatar)); its visitors get 10-minute chats and up to 120 visitor-minutes a day in all ([visitor limits](https://docs.bithuman.ai/build/website-widget), adjustable in the Console).
*Capture: The sofia-ramirez avatar answering a spoken question in the web embed, in a plain HTML page. Captured on Chrome 154 on Linux · the web embed · sofia-ramirez (Essence 2) · 2026-09-27. The visitor's question is the microphone input, mixed into the recording.* (https://docs.bithuman.ai/examples/web-embed/clip.mp4) ## Performance Measured in the tab with WebGPU, in Chrome on an Apple M4. | Configuration | Hardware | Essence 2 | Expression 2 | |---|---|---|---| | Web browser (WebGPU) | Chrome on Apple M5 | 4.0× real time | 2.4× real time | × real time: seconds of video rendered per second; 1.0× or more holds a live conversation ([method](https://docs.bithuman.ai/performance)). --- # Build a web app URL: https://docs.bithuman.ai/platforms/web/app > Set the embed's URL parameters and fit the avatar into your site. ## Integrate into your app Add parameters to the URL: | Parameter | Values | Effect | |---|---|---| | `render` | `cloud` (default), `local` | Where the avatar renders: our servers, or the visitor's tab | | `rendering_mode` | `browser`, `avatar` | Long form of `render=local`; `avatar` renders in the tab with no conversation: an Essence 1 avatar lip-syncs the visitor's own microphone, and an Essence 2 or Expression 2 avatar plays its idle loop | | `greetingLang` | a language code, for example `es` | Language of the first greeting | | `greetingMsg` | text | The first thing the avatar says | | `model` | `essence-2`, `expression-2`, `essence-1`, `expression-1` | Pins the model for this session; it must be in the agent's `supported_models`. Without it, the agent's own model | A private agent also takes `token` ([Embedding](https://docs.bithuman.ai/api/embedding)). Other parameters are ignored. ### React and other frameworks There is no npm package: the embed is an iframe in any framework. In React: ```jsx export function Avatar({ code }) { return ``` ### Run it ```bash python3 -m http.server 8765 --bind 127.0.0.1 ``` Open `http://127.0.0.1:8765/`. Any free port works: if 8765 is taken, pick another and open that one. ### Expected output The avatar's picture with **Tap to talk**. Select it and allow the microphone: the avatar greets you within a few seconds. Speak, or type into the **Type or speak…** box, and it answers out loud with its lips in sync. The red button ends the session. ### How it works The iframe loads the hosted viewer for agent `A23WJF0199`. The viewer opens a real-time session: your microphone audio goes to the agent, and the agent's voice and video come back. `allow="microphone *"` lets the iframe ask for the microphone; without the `*` the browser blocks it. URL parameters are [above](#integrate-into-your-app); session events on [Embedding](https://docs.bithuman.ai/api/embedding). ### Make it your own - **Your own avatar:** replace `A23WJF0199` with your agent code. Anyone with the code can open it and sessions bill your account; turn off Anonymous Share in the agent's sharing settings to stop that. - **Push what it says:** from your backend, `POST /v1/agent/{code}/speak` makes a live avatar say a line ([Agents](https://docs.bithuman.ai/api/agents)). - **Size and layout:** any width and height work; keep roughly a 7:12 portrait shape for Expression 2 avatars. ## Platform notes - Expression 1 avatars render in the cloud only. For an agent whose own model is Expression 1, add `model=` with another model it supports to render in the tab. - The in-tab render works for Essence 1, Expression 2, and Essence 2 avatars that have a browser build. - `render=local` never falls back to the cloud: a device that can't render the avatar in real time shows a message and starts no session ([Render with WebGPU](https://docs.bithuman.ai/platforms/web/webgpu)). ## Reference - [Embedding](https://docs.bithuman.ai/api/embedding): embed tokens, sizing and private agents. - [LiveKit](https://docs.bithuman.ai/platforms/livekit): your own UI over a cloud-rendered avatar. - Examples that need your own LiveKit server and a Python agent (not web embeds): [Next.js front end for a LiveKit agent](https://gitlab.com/bithuman/sdk/bithuman-examples/-/tree/main/integrations/nextjs-ui) · [Gradio, a Python app that also needs an OpenAI key](https://gitlab.com/bithuman/sdk/bithuman-examples/-/tree/main/integrations/gradio-web). --- # Render with WebGPU URL: https://docs.bithuman.ai/platforms/web/webgpu > Render an avatar in the browser with a device check. For Essence 2 and Expression 2, add `render=local` to request rendering in the browser. The embed checks the device before starting a session. If the device can't render the avatar in real time, it shows a message without starting a cloud session. With the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab (`render=local`). `render=local` supports Essence 2 and Expression 2 avatars that have a browser build. Essence 2 requires a hardware WebGPU adapter. Expression 2 can also use a supported browser without WebGPU when the viewer is cross-origin isolated and passes the device check. An ordinary iframe without cross-origin isolation needs WebGPU for either model. Models without an in-browser renderer show "This avatar can only be rendered in the cloud" instead. - **Download:** when a new probe is needed, the embed loads the avatar's browser bundle. The browser can cache it for later visits. Tell visitors before downloading on a metered connection. - **No cloud fallback:** an unsupported browser or a device that measures too slow shows "This device can't render this avatar in real time", with a reason. No session starts. - **First visit:** after the visitor selects the start button, the tab shows "Checking this device…" and checks the avatar before starting a session. - **Later visits:** a remembered pass or device refusal lasts up to 30 days for that model family and device/browser combination. A change to the browser major version or reported graphics vendor/architecture triggers another check. Blocked browser storage means the check runs again. - **Connection problems:** a temporary check failure retries locally; it isn't remembered as a device refusal. A missing browser build shows an avatar message instead. - **Where the conversation runs:** with the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab (`render=local`). - **Private agents:** the embed token the iframe already uses covers it ([Embedding](https://docs.bithuman.ai/api/embedding)). *Diagram: The web embed with render=local.* With render=local the avatar renders in the visitor's browser tab with WebGPU. The conversation runs on bitHuman's servers, even when the avatar renders in the tab: the microphone audio goes to bitHuman and the voice reply comes back. Choose the render mode explicitly. A GPU check alone doesn't establish browser support or whether the avatar renders fast enough. Let the embed perform its device check: ```html ``` Expected: after selecting **Tap to talk**, the embed checks the device, then starts rendering in the tab or displays a refusal message. Use `render=cloud` when you want cloud rendering instead. Do not send `Cross-Origin-Embedder-Policy` from the page that holds the iframe: the embed does not send one itself, so the browser refuses to load it. --- # Use bitHuman in ChatGPT URL: https://docs.bithuman.ai/build/chatgpt > Add bitHuman characters to ChatGPT for talking clips and live voice chats. Add bitHuman to ChatGPT and its AI characters speak in your chat: ask one to say a line and a short video clip plays, or open a live voice conversation with it. You don't need a bitHuman account. The short version, with a QR code for your phone, is at [bithuman.ai/mcp](https://www.bithuman.ai/mcp?utm_source=docs&utm_medium=referral&utm_campaign=agentic-apps-launch). ## What you'll build A ChatGPT chat where you can: - see bitHuman's characters, with a picture of each - have one say your exact words (3 to 200 characters, about 15 seconds) as a video clip you can download or share - talk to one live, by voice, for up to a minute ## Steps ### Turn on developer mode Until bitHuman is listed in ChatGPT's app directory, you add it as your own app. This needs a Plus, Pro, Business or Enterprise plan. 1. Open your profile menu, then **Settings → Apps → Advanced settings**. 2. Turn on **Developer mode**. Expected: Settings → Apps shows a **Create** button. ### Add bitHuman 1. In **Settings → Apps**, choose **Create**. 2. Name: `bitHuman`. MCP server URL: `https://mcp.bithuman.ai/mcp`. Authentication: **No authentication**. 3. Tick **I trust this application**, then **Create**. ChatGPT moves these menus from time to time; if a label differs, look for Apps and Developer mode in Settings. Expected: bitHuman is listed under Settings → Apps. ### Try it Start a new chat, choose **+**, pick **bitHuman**, and ask: - "Show me the bitHuman characters." - "Make Gio say: Welcome to the team, everyone!" - "Have Wise Pup explain photosynthesis in one sentence." - "Let me talk to Pip live." Expected: A row of character cards appears in the chat; asking for a line plays a short clip of that character saying it. ## How it works ChatGPT writes the words; bitHuman voices them and animates the character on its own servers, then the clip plays in the chat. A live conversation opens in ChatGPT's large view; closing it ends the call. Each person can make 10 new clips a day, and asking for the same line again replays it. Clips are kept for 7 days. What is stored, and for how long: [bitHuman in ChatGPT and Claude](https://docs.bithuman.ai/legal/mcp-server). ## Make it your own The ChatGPT app speaks only as bitHuman's own characters. To use your own avatar or voice, [create an agent](https://docs.bithuman.ai/build/create-avatar) and build with the [CLI's MCP server](https://docs.bithuman.ai/build/mcp) or the [SDKs](https://docs.bithuman.ai/platforms). ## Troubleshooting | Symptom | Fix | |---|---| | No Create button under Apps | turn on Developer mode first (Plus, Pro, Business or Enterprise) | | bitHuman is missing from the **+** menu | open a new chat; check the app is listed in Settings → Apps | | New features or characters don't show after an update | ChatGPT keeps the old version of the app: delete the bitHuman app in Settings → Apps and add it again | | The live call stops | closing ChatGPT's large view ends it; a live call lasts up to a minute | | "Daily limit reached" | 10 new clips per person per day; a line you already made replays | ## Next - [Use bitHuman in Claude](https://docs.bithuman.ai/build/claude) - [Claude & Cursor (MCP)](https://docs.bithuman.ai/build/mcp), for your own avatars from the CLI - [What the connector stores](https://docs.bithuman.ai/legal/mcp-server) --- # Use bitHuman in Claude URL: https://docs.bithuman.ai/build/claude > Add bitHuman characters to Claude for talking clips and live voice chats. Add bitHuman to Claude as a connector and its AI characters speak in your chat: ask one to say a line and a short video clip plays, or open a live voice conversation with it. You don't need a bitHuman account. The short version, with a QR code for your phone, is at [bithuman.ai/mcp](https://www.bithuman.ai/mcp?app=claude&utm_source=docs&utm_medium=referral&utm_campaign=agentic-apps-launch). ## What you'll build A Claude chat (web, desktop or mobile) where you can: - see bitHuman's characters, with a picture of each - have one say your exact words (3 to 200 characters, about 15 seconds) as a video clip you can download or share - talk to one live, by voice, for up to a minute ## Steps ### Add the connector 1. On claude.ai, open **Settings → Connectors** and choose **Add custom connector**. 2. Name: `bitHuman`. URL: `https://mcp.bithuman.ai/mcp`. Leave the sign-in fields empty, then **Add**. The connector then works in Claude on the web, the desktop app and the mobile apps. Expected: bitHuman is listed under Settings → Connectors. ### Claude Code ```bash claude mcp add --transport http bithuman https://mcp.bithuman.ai/mcp ``` The same command with a copy button: [bithuman.ai/mcp, Claude Code tab](https://www.bithuman.ai/mcp?app=code&utm_source=docs&utm_medium=referral&utm_campaign=agentic-apps-launch). Expected: `claude mcp list` shows `bithuman` as connected. ### Try it In a new chat, ask: - "Show me the bitHuman characters." - "Make Gio say: Welcome to the team, everyone!" - "Have Wise Pup explain photosynthesis in one sentence." - "Let me talk to Pip live." Expected: A row of character cards appears in the chat; asking for a line plays a short clip of that character saying it. ## How it works Claude writes the words; bitHuman voices them and animates the character on its own servers, then the clip plays in the chat. A live conversation runs inside the chat, with **Open in a new tab** as a fallback. Each person can make 10 new clips a day, and asking for the same line again replays it. Clips are kept for 7 days. What is stored, and for how long: [bitHuman in ChatGPT and Claude](https://docs.bithuman.ai/legal/mcp-server). ## Make it your own The connector speaks only as bitHuman's own characters. To use your own avatar or voice, [create an agent](https://docs.bithuman.ai/build/create-avatar) and add the [CLI's MCP server](https://docs.bithuman.ai/build/mcp) to Claude Code or Claude Desktop instead. ## Troubleshooting | Symptom | Fix | |---|---| | Claude doesn't use bitHuman | name it in the prompt ("bitHuman characters"), and check the connector is on for the chat | | New features or characters don't show after an update | Claude keeps the old version: **Settings → Connectors → bitHuman → Disconnect**, then **Connect** | | The live call shows a black screen | choose **Open in a new tab** | | "Daily limit reached" | 10 new clips per person per day; a line you already made replays | ## Next - [Use bitHuman in ChatGPT](https://docs.bithuman.ai/build/chatgpt) - [Claude & Cursor (MCP)](https://docs.bithuman.ai/build/mcp), for your own avatars from the CLI - [What the connector stores](https://docs.bithuman.ai/legal/mcp-server) --- # Put an animated character in an iPhone app URL: https://docs.bithuman.ai/build/how-to/iphone-character > Feed speech to the Swift package; the character renders on the iPhone. ## What you'll build Add the Swift package's `Expression2` product to your app, open an Expression 2 avatar, and feed it 16 kHz mono speech: it hands back lip-synced frames that you draw. The character renders on the iPhone or iPad itself, with no render server. Expression 2 animates any character from one portrait, so the avatar can be a cartoon, an animal, a robot or a creature as well as a person. This page follows the [iOS Expression 2 example](https://docs.bithuman.ai/examples/ios-expression-2), which uses the `wise-pup` sample avatar. You need: - a Mac with Xcode 26 or newer, and an Apple Developer team; - an iPhone or iPad, or the iOS Simulator (Expression 2 also runs in the Simulator; judge speed on a device); - an [API secret](https://docs.bithuman.ai/start/api-secret) (Creator plan or higher). ## Steps ### Run the example first Clone the example and download the `wise-pup` avatar, the shared engine and a 16 kHz speech clip. The download is anonymous: ```bash git clone https://gitlab.com/bithuman/sdk/bithuman-examples.git cd bithuman-examples/swift/ios-expression2 ./setup.sh ``` Add `BITHUMAN_API_SECRET` under **Product → Scheme → Edit Scheme → Run → Environment Variables**, then open the project, pick your team under **Signing & Capabilities** and press **Run**: ```bash open IOSExpression2.xcodeproj ``` Expected: The avatar appears and idles. Tap **Speak**: it says the sample line with its lips in sync. The first launch prepares the engine on the device and takes a few seconds longer than later launches. ### Add the Swift package to your app In Xcode choose *File → Add Package Dependencies…*, paste `https://gitlab.com/bithuman/sdk/homebrew-bithuman` and attach the `Expression2` product to your target. The current version pin and the `Package.swift` line are on [iOS & iPadOS: Install](https://docs.bithuman.ai/platforms/ios#install). Expected: Your app builds with `import Expression2`. ### Set your API secret The engine checks an API secret when a session starts. Set it before you create the engine: ```swift Expression2Credential.set(ProcessInfo.processInfo.environment["BITHUMAN_API_SECRET"] ?? "") ``` Set `BITHUMAN_API_SECRET` in the scheme's environment while you build. Keep the secret out of the app bundle you ship ([What a shipped app holds](https://docs.bithuman.ai/start/api-secret#what-a-shipped-app-holds)). Expected: No refusal when the engine opens in the next step. ### Add the character's files Download the `wise-pup` sample avatar, the shared Expression 2 engine and a 16 kHz speech clip, and add them to your app. No account is needed for these downloads: ```bash curl -fL -o A23WJF0199.imx "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2" curl -fLO "https://downloads.bithuman.ai/homebrew-bithuman/expression2-engine-mac-arm64-1.0.0/mac-arm64-1.0.0.engine" curl -fL -o speech16k.wav "https://api.bithuman.ai/v1/agent/A23WJF0199/model/download?model=expression-2&member=demo_speech_16k.wav" ``` The `mac` engine file is the right one for iPhone apps too. Expected: Three files in your app: the avatar (`A23WJF0199.imx`), the shared engine and the speech clip. ### Feed it speech and draw the frames The example keeps the engine in an actor. `create` opens the two files as they are, `feed` takes 16 kHz mono float audio, `flushTail()` ends a reply, and `pull()` returns the next frame, or `nil` until a chunk of frames is ready: ```swift // excerpt: swift/ios-expression2/Sources/App.swift actor Renderer { private var engine: Expression2Engine? // … let e = try Expression2Engine.create(avatarContainer: avatar, sharedEngineContainer: sharedEngine, stagingDir: staging) engine = e // … func idleFrame() -> [UInt8]? { engine?.idle } func feed(_ samples: [Float]) { engine?.feed(samples) } func flushTail() { engine?.flushTail() } func reset() { engine?.resetState(clearFrames: true) } // … func pullOne() -> [UInt8]? { engine?.pull()?.frame } func queued() -> Int { engine?.queuedFrames ?? 0 } } ``` Feed the speech in chunks and take frames out as they appear, so the app feeds and drains at the same time: ```swift // excerpt: swift/ios-expression2/Sources/App.swift var i = 0 while i < pcm.count { let j = min(i + chunk, pcm.count) await renderer.feed(Array(pcm[i..`. Expected: Your character in place of `wise-pup`, with the same calls. ### Run it on a device Choose your iPhone or iPad as the run destination. Expression 2 also runs in the iOS Simulator, which is fine while you build; Essence 2, the photoreal model, needs a physical device. How fast each device renders is on [Performance](https://docs.bithuman.ai/performance). Expected: The character idles and speaks on the device. ## How it works The avatar renders inside your app, from 16 kHz mono speech to picture frames. Your app keeps its own speech recognition, language model and voice, and feeds the reply's audio to the engine. The engine contacts bitHuman only to check your API secret and report session time, and a session bills active session time, talking or idle ([pricing](https://docs.bithuman.ai/pricing)). ## Make it your own - **A live conversation:** pass the audio your voice stack produces into `feed`, in chunks, as it arrives: [Give an app's voice assistant a face](https://docs.bithuman.ai/build/how-to/voice-assistant-face). - **A companion screen:** interruptions, resampling and closing the avatar with the screen are on [Companion app](https://docs.bithuman.ai/build/companion-app). - **A photoreal person:** the [iOS Essence 2 example](https://docs.bithuman.ai/examples/ios-essence-2) renders an Essence 2 avatar on the device. - **A Mac app:** the same calls run on a Mac: [macOS example](https://docs.bithuman.ai/examples/macos-expression-2). ## Troubleshooting | Symptom | Fix | |---|---| | `refusing to serve: no API secret was found` | add `BITHUMAN_API_SECRET` to the scheme's environment variables, or call `Expression2Credential.set` before `create` | | `Sources/Model/agent.imx is missing` or a missing shared engine file | run `./setup.sh` again; it names the step that failed | | The view stays empty and nothing throws | keep polling `pull()` while you feed; it returns `nil` between chunks | | `unable to resolve module dependency: 'Expression2'` on a Simulator build of your own project | the simulator slices are arm64 only: set `EXCLUDED_ARCHS[sdk=iphonesimulator*] = x86_64` (the example project already does) | More on [Apple: Troubleshooting](https://docs.bithuman.ai/platforms/swift/troubleshooting#ios). --- # Give an app's voice assistant a face URL: https://docs.bithuman.ai/build/how-to/voice-assistant-face > Feed your assistant's speech to the SDK and draw the lip-synced frames. ## What you'll build The Swift and Android SDKs render: your app passes in 16 kHz mono speech from any voice stack and draws the frames, so the persona, the voice and the language model are yours to choose. Keep the speech-to-speech stack your assistant already uses, resample its reply audio to 16 kHz mono if it returns another rate, and feed it to the avatar. The avatar renders on the device, inside your app. You need: - an app with a voice assistant that produces reply audio (any speech recognition, language model and voice); - Xcode 26 and an iPhone or iPad, or Android Studio and a physical arm64 Android phone; - an [API secret](https://docs.bithuman.ai/start/api-secret) (Creator plan or higher). ## Steps The code in these steps is from [Companion app](https://docs.bithuman.ai/build/companion-app) and [Android: Run your first avatar](https://docs.bithuman.ai/platforms/android#run-your-first-avatar). ### Choose where the face renders | Where | Use | The conversation | |---|---|---| | In your app, on the device | the [Swift package](https://docs.bithuman.ai/platforms/ios) (iPhone, iPad, Mac) or the [Android SDK](https://docs.bithuman.ai/platforms/android) (arm64) | stays with your app and the services it uses | | In your own Python code, on a Mac or a Linux machine | the [Python SDK](https://docs.bithuman.ai/platforms/python): `render` takes 16 kHz mono audio | stays with your code and the services it uses | | In the bitHuman cloud or a browser tab | the [web embed](https://docs.bithuman.ai/platforms/web) or a [LiveKit](https://docs.bithuman.ai/platforms/livekit) agent with a bitHuman cloud avatar | runs on bitHuman's servers with the web embed | This page follows the first row. Android, and Essence 2 on iPhone and iPad, need a physical device, not an emulator or the Simulator. Expected: One SDK: the Swift package or the Android SDK. ### Pick the avatar Use the `wise-pup` sample (Expression 2, agent code `A23WJF0199`) for a character, or the `sofia-ramirez` sample (Essence 2, agent code `A52DHS2219`) for a photoreal person, while you build. Or [create your own](https://docs.bithuman.ai/build/create-avatar) from one portrait. Expected: An agent code. ### Add the SDK and set your API secret - **iPhone and iPad:** the Swift package, product `Expression2` or `Essence2Kit` ([install](https://docs.bithuman.ai/platforms/ios#install)). - **Android:** the Android SDK, `expression2-android` or `essence2-android` ([install](https://docs.bithuman.ai/platforms/android#install)). One call covers the avatar's download and the session. Set it before anything downloads: ```swift tab="Swift" Expression2Credential.set(ProcessInfo.processInfo.environment["BITHUMAN_API_SECRET"] ?? "") ``` ```kotlin tab="Kotlin" Expression2Credential.set(BuildConfig.BITHUMAN_API_SECRET) // before fetch() and create() ``` On Android, `BuildConfig.BITHUMAN_API_SECRET` comes from your app's `build.gradle.kts` (set `BITHUMAN_API_SECRET` in the environment before you build, or read `local.properties` as on [the Android page](https://docs.bithuman.ai/platforms/android#install)): ```kotlin android { buildFeatures { buildConfig = true } defaultConfig { buildConfigField("String", "BITHUMAN_API_SECRET", "\"${System.getenv("BITHUMAN_API_SECRET") ?: ""}\"") } } ``` That field compiles the secret into the APK, so use it for local builds only. Keep the secret out of the app bundle you ship: fetch it from your backend at runtime ([What a shipped app holds](https://docs.bithuman.ai/start/api-secret#what-a-shipped-app-holds)). Expected: No refusal when the avatar opens. Without a secret, opening it fails and says why. ### Resample speech to 16 kHz The engines take 16 kHz mono speech. A voice service that returns another rate needs one conversion first; a service that returns 24 kHz 16-bit PCM, for example, needs a 24 kHz to 16 kHz converter. Keep one converter for the whole reply and pass each chunk through it before you feed the avatar: ```swift tab="Swift" // Add to your app: 24 kHz 16-bit mono PCM in, the 16 kHz [Float] that feed(_:) takes out. import AVFoundation final class To16k { private let from = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 24_000, channels: 1, interleaved: false)! private let to = AVAudioFormat(commonFormat: .pcmFormatFloat32, sampleRate: 16_000, channels: 1, interleaved: false)! private lazy var converter = AVAudioConverter(from: from, to: to)! func convert(_ pcm24k: Data) -> [Float] { let n = AVAudioFrameCount(pcm24k.count / 2) guard n > 0, let input = AVAudioPCMBuffer(pcmFormat: from, frameCapacity: n), let output = AVAudioPCMBuffer(pcmFormat: to, frameCapacity: n) else { return [] } input.frameLength = n pcm24k.withUnsafeBytes { input.int16ChannelData![0].update(from: $0.bindMemory(to: Int16.self).baseAddress!, count: Int(n)) } var given = false _ = converter.convert(to: output, error: nil) { _, status in if given { status.pointee = .noDataNow; return nil } given = true status.pointee = .haveData return input } return Array(UnsafeBufferPointer(start: output.floatChannelData![0], count: Int(output.frameLength))) } } ``` ```kotlin tab="Kotlin" // Add to your app: 24 kHz 16-bit mono PCM in; floats() for Expression 2, pcm16() for Essence 2. import java.nio.ByteBuffer import java.nio.ByteOrder class To16k { private var rest = FloatArray(0) // input not used yet private var pos = 0.0 // read position in the input, in samples fun floats(pcm24k: ByteArray): FloatArray { val s = ByteBuffer.wrap(pcm24k).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer() val input = rest + FloatArray(s.remaining()) { s.get() / 32768f } val out = ArrayList() while (pos + 1 < input.size) { val i = pos.toInt() val f = (pos - i).toFloat() out.add(input[i] * (1 - f) + input[i + 1] * f) pos += 1.5 // 24 000 / 16 000 } rest = input.copyOfRange(pos.toInt(), input.size) pos -= pos.toInt() return out.toFloatArray() } fun pcm16(pcm24k: ByteArray): ByteArray { val f = floats(pcm24k) val b = ByteBuffer.allocate(f.size * 2).order(ByteOrder.LITTLE_ENDIAN) for (x in f) b.putShort((x.coerceIn(-1f, 1f) * 32767).toInt().toShort()) return b.array() } } ``` Expected: Each 24 kHz chunk comes out as two thirds as many 16 kHz samples, and the lips keep pace with the voice. Speech fed at the wrong rate makes the mouth run slow and long. ### Feed the reply and draw the frames Your speech recognition hears the user, your language model writes the reply, and your voice turns it into audio. Feed that audio as 16 kHz mono, say where the reply ends, and draw the frames the avatar hands back. On Android, with Expression 2 and the `wise-pup` sample (call `render` off the main thread): ```kotlin import ai.bithuman.expression2.Expression2Avatar import ai.bithuman.expression2.Expression2Credential import ai.bithuman.expression2.Expression2ModelStore import ai.bithuman.expression2.Expression2Options import android.content.Context import android.graphics.Bitmap /** [pcm16k] is 16 kHz mono float32 in [-1, 1]. */ fun render(context: Context, pcm16k: FloatArray, show: (Bitmap) -> Unit) { Expression2Credential.set(BuildConfig.BITHUMAN_API_SECRET) // before fetch() and create() val model = Expression2ModelStore(context).fetch("A23WJF0199") // ~160 MB, first run only Expression2Avatar.create(context, model, Expression2Options()).use { avatar -> val frame = avatar.newFrameBitmap() // 416 x 720, allocate once avatar.feed(pcm16k) avatar.flushTail() // end of the utterance while (true) { if (avatar.pull(frame) != null) { show(frame); continue } if (!avatar.hasPendingTail && avatar.queuedFrames == 0) break Thread.sleep(10) // null means "not ready yet" } } } ``` Pass `To16k().floats(chunk)` as `pcm16k`. In Swift, `feed(samples)` and `flushTail()` do the same, and the reply's audio starts on its first frame (`audioTime == 0`): the complete loop is in [iOS & iPadOS: First frame](https://docs.bithuman.ai/platforms/ios#run-your-first-avatar), and an Essence 2 version is in [Companion app: Speak a reply](https://docs.bithuman.ai/build/companion-app#speak-a-reply). Expected: The lips follow the reply's audio from its first word, and idle motion returns when the reply ends. ### Let the user interrupt When your speech recognition hears the user start talking over the avatar, stop your audio player and drop the rest of the reply: `interrupt()` in Swift; in Kotlin, `resetState(true)` for Expression 2 or `resetAudio()` for Essence 2 ([Android: Integrate into your app](https://docs.bithuman.ai/platforms/android/app#integrate-into-your-app)). Idle motion continues from the current frame. Expected: The mouth stops with the voice, and the next reply starts cleanly. ## How it works The avatar renders inside your app, from 16 kHz mono speech to picture frames; it does not listen, think or speak on its own. When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. A session bills active session time, talking or idle, for as long as the avatar is open ([pricing](https://docs.bithuman.ai/pricing)). ## Make it your own - **A full companion screen:** closing the avatar with the screen and an Essence 2 version of every step are on [Companion app](https://docs.bithuman.ai/build/companion-app). - **A character from your own art:** [Turn a drawing, mascot or pet photo into a talking character](https://docs.bithuman.ai/build/how-to/drawing-to-character). - **One codebase:** the [Flutter plugin](https://docs.bithuman.ai/platforms/flutter) wraps both engines. - **The conversation on bitHuman's servers instead:** the [web embed](https://docs.bithuman.ai/platforms/web) runs a managed agent with your persona; with the web embed, the conversation runs on bitHuman's servers, even when the avatar renders in the tab. ## Troubleshooting | Symptom | Fix | |---|---| | Opening the avatar fails with a metering refusal | Set the API secret before the download and the first `create`, on the Creator plan or higher. | | It works on a phone but not in the Simulator or an emulator | Expected: Essence 2 and the Android SDK need a physical device; Expression 2 also runs in the iOS Simulator. | | The mouth runs slow and the reply lasts too long | The speech is not 16 kHz: [resample it](#resample-speech-to-16-khz) before you feed it. | | The lips run ahead of the voice | Start the reply's audio with its first speech frame (`audioTime == 0` in Swift), not when you feed it. | | The reply is cut short | Mark the end of each reply once: `flushTail()`, or `endOfAudio()` for Essence 2 on Android. |