On the device

Essence 2 and Expression 2 render on the device in front of the user: iPhone, iPad and Mac, Android phones, or a WebGPU browser tab.

Creator plan or higher Renders on the device In the browser (WebGPU) Physical device

What it is

The avatar renders inside your app on the device in front of the user: iPhone, iPad and Mac with the Swift package, Android phones with the Android SDK or the Flutter plugin, or the visitor’s browser tab with WebGPU. There is no render server to run.

The mobile SDKs take any 16 kHz mono speech your pipeline produces and return frames, so any speech-recognition, language-model and voice stack works.

Where it renders

QuestionOn the device
Where the avatar renderson the iPhone, iPad, Mac or Android phone in front of the user, or in a WebGPU browser tab
Where the conversation runsyour app's choice; with the web embed, on bitHuman's servers
What reaches bitHumanwith your own voice and language services, usage metering only, never audio, video or conversation text
Networkto start; rendering continues through a drop of up to 5 minutes
IPHONE, IPAD, MAC OR ANDROIDYour appyour own voice and language servicesThe avatar renders on the deviceBITHUMANUsage meteringusage only: noaudio, video ortextthe avatar,once
On the device. On the device, your app renders the avatar on the iPhone, iPad, Mac or Android phone and brings its own voice and language services. The avatar model downloads once. When you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text.
  • In your app: when the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text.
  • On Android: after the one-time model download, the only network traffic is usage reporting.
  • In the browser: with the web embed, the conversation runs on bitHuman’s servers, even when the avatar renders in the tab (render=local).

Models available here

Essence 1 and Expression 1 are not available on phones or in the Swift package.

Speed

Measured on the device, including 10-minute held runs:

ConfigurationEssence 2Expression 2
iPhone 15 iPhone · Swift package
2.1× real timeiPhone 15 · Swift package 2.17.3 · measured 2026-09-27
5.5× real timeiPhone 15 · Swift package 2.18.0 · measured 2026-09-27
iPhone 15 iPhone · Swift package held 10 min
1.3× real timeiPhone 15 · Swift package 2.15.0 · held 10 min · measured 2026-09-25
5.1× real timeiPhone 15 · Swift package 2.15.0 · held 10 min · measured 2026-09-25
Samsung Galaxy S25+ Android
2.0× real timeSamsung Galaxy S25+ · essence2-android 0.7.0 · measured 2026-09-25
2.4× real timeSamsung Galaxy S25+ · expression2-android 0.4.10 · measured 2026-09-23
Samsung Galaxy S25+ Android held 10 min
1.4× real timeSamsung Galaxy S25+ · essence2-android 0.8.1 · held 10 min · measured 2026-09-27
2.2× real timeSamsung Galaxy S25+ · expression2-android 0.5.2 · held 10 min · measured 2026-09-27
Chrome on Apple M4 Web browser (WebGPU)
1.7× real timeChrome on Apple M4 · web viewer · measured 2026-09-27
1.9× real timeChrome on Apple M4 · web viewer · measured 2026-09-27
Apple M4 macOS · Swift package
4.8× real timeApple M4 · Swift package 2.15.0 · measured 2026-09-24
8.8× real timeApple M4 · Swift package 2.15.0 · measured 2026-09-24

Times real time: seconds of avatar video rendered per second. At 1.0× or more, an avatar holds a live conversation. Select a figure for its release and date. All configurations and how we measure.

Price

2 credits per minute of active session time for Essence 2 and Expression 2, about $0.02 a minute at the top-up rate of $1 = 100 credits. Realtime usage bills active session time, talking or idle, to the second. Every rate: Pricing and credits.

Limits

  • Network: a session checks your credential when it starts and keeps rendering through a network drop of up to 5 minutes.
  • Devices: iPhone, iPad and Android need a physical device, not a simulator or an emulator. Essence 2 on Apple needs iOS 26 or macOS 26.
  • Sessions: on-device sessions are limited by credits, not by a session cap.
  • First run: each avatar downloads once (about 160–370 MB, by model and platform), then stays on the device.

First command

iOS & iPadOS

// Package.swift (or Xcode → Add Package Dependencies)
.package(url: "https://github.com/bithuman-product/homebrew-bithuman.git", from: "2.18.0")

Android

// app/build.gradle.kts
implementation("ai.bithuman:expression2-android:0.5.2")

Mac

git clone https://github.com/bithuman-product/bithuman-examples.git
cd bithuman-examples/swift/macos-expression2 && ./setup.sh
BITHUMAN_API_SECRET="<your API secret>" swift run -c release MacOSExpression2

Web

<iframe src="https://www.bithuman.ai/embed/A23WJF0199?render=local" allow="microphone *"
        style="width:100%;height:600px;border:0"></iframe>

The whole first frame for each: iOS & iPadOS · Android · macOS · Web.

Choosing between modes