Give an app's voice assistant a face
Feed your assistant's speech to the SDK and draw the lip-synced frames.
What you’ll build
The Swift and Android SDKs render: your app passes in 16 kHz mono speech from any voice stack and draws the frames, so the persona, the voice and the language model are yours to choose. Keep the speech-to-speech stack your assistant already uses, resample its reply audio to 16 kHz mono if it returns another rate, and feed it to the avatar. The avatar renders on the device, inside your app.
You need:
- an app with a voice assistant that produces reply audio (any speech recognition, language model and voice);
- Xcode 26 and an iPhone or iPad, or Android Studio and a physical arm64 Android phone;
- an API secret (Creator plan or higher).
Steps
6 steps
The code in these steps is from Companion app and Android: Run your first avatar.
Choose where the face renders
Where Use The conversation In your app, on the device the Swift package (iPhone, iPad, Mac) or the Android SDK (arm64) stays with your app and the services it uses In your own Python code, on a Mac or a Linux machine the Python SDK: rendertakes 16 kHz mono audiostays with your code and the services it uses In the bitHuman cloud or a browser tab the web embed or a LiveKit agent with a bitHuman cloud avatar runs on bitHuman’s servers with the web embed This page follows the first row. Android, and Essence 2 on iPhone and iPad, need a physical device, not an emulator or the Simulator.
Expected
One SDK: the Swift package or the Android SDK.
Pick the avatar
Use the
wise-pupsample (Expression 2, agent codeA23WJF0199) for a character, or thesofia-ramirezsample (Essence 2, agent codeA52DHS2219) for a photoreal person, while you build. Or create your own from one portrait.Expected
An agent code.
Add the SDK and set your API secret
- iPhone and iPad: the Swift package, product
Expression2orEssence2Kit(install). - Android: the Android SDK,
expression2-androidoressence2-android(install).
One call covers the avatar’s download and the session. Set it before anything downloads:
Swift
Expression2Credential.set(ProcessInfo.processInfo.environment["BITHUMAN_API_SECRET"] ?? "")Kotlin
Expression2Credential.set(BuildConfig.BITHUMAN_API_SECRET) // before fetch() and create()On Android,
BuildConfig.BITHUMAN_API_SECRETcomes from your app’sbuild.gradle.kts(setBITHUMAN_API_SECRETin the environment before you build, or readlocal.propertiesas on the Android page):android { buildFeatures { buildConfig = true } defaultConfig { buildConfigField("String", "BITHUMAN_API_SECRET", "\"${System.getenv("BITHUMAN_API_SECRET") ?: ""}\"") } }That field compiles the secret into the APK, so use it for local builds only. Keep the secret out of the app bundle you ship: fetch it from your backend at runtime (What a shipped app holds).
Expected
No refusal when the avatar opens. Without a secret, opening it fails and says why.
- iPhone and iPad: the Swift package, product
Resample speech to 16 kHz
The engines take 16 kHz mono speech. A voice service that returns another rate needs one conversion first; a service that returns 24 kHz 16-bit PCM, for example, needs a 24 kHz to 16 kHz converter. Keep one converter for the whole reply and pass each chunk through it before you feed the avatar:
Swift
// Add to your app: 24 kHz 16-bit mono PCM in, the 16 kHz [Float] that feed(_:) takes out. import AVFoundation final class To16k { private let from = AVAudioFormat(commonFormat: .pcmFormatInt16, sampleRate: 24_000, channels: 1, interleaved: false)! private let to = AVAudioFormat(commonFormat: .pcmFormatFloat32, sampleRate: 16_000, channels: 1, interleaved: false)! private lazy var converter = AVAudioConverter(from: from, to: to)! func convert(_ pcm24k: Data) -> [Float] { let n = AVAudioFrameCount(pcm24k.count / 2) guard n > 0, let input = AVAudioPCMBuffer(pcmFormat: from, frameCapacity: n), let output = AVAudioPCMBuffer(pcmFormat: to, frameCapacity: n) else { return [] } input.frameLength = n pcm24k.withUnsafeBytes { input.int16ChannelData![0].update(from: $0.bindMemory(to: Int16.self).baseAddress!, count: Int(n)) } var given = false _ = converter.convert(to: output, error: nil) { _, status in if given { status.pointee = .noDataNow; return nil } given = true status.pointee = .haveData return input } return Array(UnsafeBufferPointer(start: output.floatChannelData![0], count: Int(output.frameLength))) } }Kotlin
// Add to your app: 24 kHz 16-bit mono PCM in; floats() for Expression 2, pcm16() for Essence 2. import java.nio.ByteBuffer import java.nio.ByteOrder class To16k { private var rest = FloatArray(0) // input not used yet private var pos = 0.0 // read position in the input, in samples fun floats(pcm24k: ByteArray): FloatArray { val s = ByteBuffer.wrap(pcm24k).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer() val input = rest + FloatArray(s.remaining()) { s.get() / 32768f } val out = ArrayList<Float>() while (pos + 1 < input.size) { val i = pos.toInt() val f = (pos - i).toFloat() out.add(input[i] * (1 - f) + input[i + 1] * f) pos += 1.5 // 24 000 / 16 000 } rest = input.copyOfRange(pos.toInt(), input.size) pos -= pos.toInt() return out.toFloatArray() } fun pcm16(pcm24k: ByteArray): ByteArray { val f = floats(pcm24k) val b = ByteBuffer.allocate(f.size * 2).order(ByteOrder.LITTLE_ENDIAN) for (x in f) b.putShort((x.coerceIn(-1f, 1f) * 32767).toInt().toShort()) return b.array() } }Expected
Each 24 kHz chunk comes out as two thirds as many 16 kHz samples, and the lips keep pace with the voice. Speech fed at the wrong rate makes the mouth run slow and long.
Feed the reply and draw the frames
Your speech recognition hears the user, your language model writes the reply, and your voice turns it into audio. Feed that audio as 16 kHz mono, say where the reply ends, and draw the frames the avatar hands back. On Android, with Expression 2 and the
wise-pupsample (callrenderoff the main thread):import ai.bithuman.expression2.Expression2Avatar import ai.bithuman.expression2.Expression2Credential import ai.bithuman.expression2.Expression2ModelStore import ai.bithuman.expression2.Expression2Options import android.content.Context import android.graphics.Bitmap /** [pcm16k] is 16 kHz mono float32 in [-1, 1]. */ fun render(context: Context, pcm16k: FloatArray, show: (Bitmap) -> Unit) { Expression2Credential.set(BuildConfig.BITHUMAN_API_SECRET) // before fetch() and create() val model = Expression2ModelStore(context).fetch("A23WJF0199") // ~160 MB, first run only Expression2Avatar.create(context, model, Expression2Options()).use { avatar -> val frame = avatar.newFrameBitmap() // 416 x 720, allocate once avatar.feed(pcm16k) avatar.flushTail() // end of the utterance while (true) { if (avatar.pull(frame) != null) { show(frame); continue } if (!avatar.hasPendingTail && avatar.queuedFrames == 0) break Thread.sleep(10) // null means "not ready yet" } } }Pass
To16k().floats(chunk)aspcm16k. In Swift,feed(samples)andflushTail()do the same, and the reply’s audio starts on its first frame (audioTime == 0): the complete loop is in iOS & iPadOS: First frame, and an Essence 2 version is in Companion app: Speak a reply.Expected
The lips follow the reply's audio from its first word, and idle motion returns when the reply ends.
Let the user interrupt
When your speech recognition hears the user start talking over the avatar, stop your audio player and drop the rest of the reply:
interrupt()in Swift; in Kotlin,resetState(true)for Expression 2 orresetAudio()for Essence 2 (Android: Integrate into your app). Idle motion continues from the current frame.Expected
The mouth stops with the voice, and the next reply starts cleanly.
All steps done. Next: make it your own.
How it works
The avatar renders inside your app, from 16 kHz mono speech to picture frames; it does not listen, think or speak on its own. When the avatar renders in your app on the device and you use your own voice and language services, bitHuman receives usage metering only, never audio, video or conversation text. A session bills active session time, talking or idle, for as long as the avatar is open (pricing).
Make it your own
- A full companion screen: closing the avatar with the screen and an Essence 2 version of every step are on Companion app.
- A character from your own art: Turn a drawing, mascot or pet photo into a talking character.
- One codebase: the Flutter plugin wraps both engines.
- The conversation on bitHuman’s servers instead: the web embed runs a managed agent with your persona; with the web embed, the conversation runs on bitHuman’s servers, even when the avatar renders in the tab.
Troubleshooting
| Symptom | Fix |
|---|---|
| Opening the avatar fails with a metering refusal | Set the API secret before the download and the first create, on the Creator plan or higher. |
| It works on a phone but not in the Simulator or an emulator | Expected: Essence 2 and the Android SDK need a physical device; Expression 2 also runs in the iOS Simulator. |
| The mouth runs slow and the reply lasts too long | The speech is not 16 kHz: resample it before you feed it. |
| The lips run ahead of the voice | Start the reply’s audio with its first speech frame (audioTime == 0 in Swift), not when you feed it. |
| The reply is cut short | Mark the end of each reply once: flushTail(), or endOfAudio() for Essence 2 on Android. |