iOS Expression 2

A complete SwiftUI app that renders a talking Expression 2 avatar on an iPhone or iPad, on the device: clone it, fetch the avatar, set your API secret, run.

The wise-pup avatar mid-sentence, as drawn by the ios-expression2 example app on an iPhone 15
Captured on iPhone 15 (iOS 26) · Swift package 2.14.2 · wise-pup (Expression 2) · 2026-09-23.

A SwiftUI app: Speak plays a speech clip through the avatar, and Talk to it drives the avatar from the microphone, live. Everything renders on the device at 416×720, 20 frames a second; the engine contacts bitHuman only to check your API secret and report session time.

Requirements

You needNotes
A Mac with Xcode 26 or newer, and an Apple Developer teama device build is a signed build
A physical iPhone or iPad with Apple siliconthe Simulator cannot run the engine; no Apple entitlement is needed
The bitHuman CLIsetup.sh uses it once: brew install bithuman-product/bithuman/bithuman-cli
An API secretthe engine bills session time, talking or idle

Get the code

git clone https://github.com/bithuman-product/bithuman-examples.git
cd bithuman-examples/swift/ios-expression2
./setup.sh

setup.sh puts the wise-pup avatar, the shared engine files and a speech clip in Sources/Model/ (git-ignored). For your own avatar run BITHUMAN_API_SECRET=… ./setup.sh <AGENT_CODE>.

Set your API secret

In Xcode: Product → Scheme → Edit Scheme → Run → Environment Variables, add BITHUMAN_API_SECRET. The engine reads it at create.

Run it

open IOSExpression2.xcodeproj

Pick your team under Signing & Capabilities, choose your iPhone as the run destination, and press Run.

Expected output

The avatar appears and idles. Tap Speak: it says the sample line with its lips in sync. Xcode’s console prints engine ready, then the frames generated for the clip, and writes the first frame to the app’s Documents/first-frame.png. The first launch prepares the engine on the device and takes a few seconds longer than later launches.

How it works

Sources/App.swift is the whole app. Its Renderer actor wraps Expression2Engine from the Swift package’s Expression2 product:

  1. Expression2Engine.create(modelPath:sharedEngineDir:) opens the avatar;
  2. feed(samples) takes 16 kHz mono float audio as it arrives, and flushTail() ends a reply;
  3. pull() returns the next frame, or nil until a chunk of frames is ready, so the app feeds and drains at the same time;
  4. a display loop shows one frame per 50 ms of audio, on the audio clock.

The same three calls run on a Mac: macOS example. The API is on Apple and Apple API reference.

The code that matters

Sources/App.swift keeps the engine in an actor, then feeds the speech in 100 ms chunks and takes frames out as they appear:

// excerpt: swift/ios-expression2/Sources/App.swift
actor Renderer {
    private var engine: Expression2Engine?
        // …
        let e = try Expression2Engine.create(modelPath: dir, sharedEngineDir: sharedEngine)
        engine = e
    // …
    func idleFrame() -> [UInt8]? { engine?.idle }
    func feed(_ samples: [Float]) { engine?.feed(samples) }
    func flushTail() { engine?.flushTail() }
    func reset() { engine?.resetState(clearFrames: true) }
    // …
    func pullOne() -> [UInt8]? { engine?.pull()?.frame }
    func queued() -> Int { engine?.queuedFrames ?? 0 }
}
// excerpt: swift/ios-expression2/Sources/App.swift
var i = 0
while i < pcm.count {
    let j = min(i + chunk, pcm.count)
    await renderer.feed(Array(pcm[i..<j]))
    i = j
    // Generation is ASYNCHRONOUS — pull() returns nil until a chunk of
    // frames lands, so this usually takes nothing on the first passes and
    // then keeps up. It is not a busy-wait: it only removes what is ready.
    while let f = await renderer.pullOne() {
        if let cg = makeCGImage(f, w, h) { frames.append(cg) }
    }
}
await renderer.flushTail()

The complete file is on GitHub.

Make it your own

  • Your own avatar: create one with the Agents API ("model": "expression-2"), then BITHUMAN_API_SECRET=… ./setup.sh <AGENT_CODE>.
  • Your own voice pipeline: feed the audio your text-to-speech produces into feed, in chunks, as it arrives.
  • Ship it: don’t put the secret in the app bundle. Fetch it from your backend or the Keychain and call Expression2Credential.set(key) before create.
  • A photoreal person: the iOS Essence 2 example renders an Essence 2 avatar at full resolution.

Troubleshooting

SymptomFix
refusing to serve: no API secret was foundadd BITHUMAN_API_SECRET to the scheme’s environment variables
missing w2v_frontend_cpuAndNE.mlpackagerun ./setup.sh again; it stages the shared engine files
the bithuman CLI is not on PATH from setup.shbrew install bithuman-product/bithuman/bithuman-cli, then rerun
The view stays empty and nothing throwskeep polling pull() while you feed; it returns nil between chunks
It builds for the Simulator and crashes thererun on a physical device

More on Apple: Troubleshooting.

Next