Performance

How fast Essence 2 and Expression 2 render on iPhone, Android, a WebGPU browser, a Mac, a Linux PC with no GPU and the bitHuman cloud, in times real time, with the raw data as performance.json.

× real time is seconds of avatar video rendered per second, rounded down to one decimal: at 1.0× or more, an avatar holds a live conversation.

Every configuration we publish renders faster than real time.

  • 1.0×: holds a live conversation
  • CPU only (no GPU)
  • held for 10 minutes
  • Bars run to 10×; a longer bar shows its number

Essence 2

On a phone
ConfigurationEssence 2, × real time
iPhone 15iPhone · Swift package2.1× real time
Details for iPhone 15, iPhone · Swift packageiPhone 15 · Swift package 2.17.3 · measured 2026-09-27 · 28.4 s speech clip · 54 frames rendered per second
iPhone 15iPhone · Swift packageheld 10 min1.3× real time
Details for iPhone 15, iPhone · Swift packageiPhone 15 · Swift package 2.15.0 · held 10 min · measured 2026-09-25 · 28.4 s speech clip · 33 frames rendered per second
Samsung Galaxy S25+Android2.0× real time
Details for Samsung Galaxy S25+, AndroidSamsung Galaxy S25+ · essence2-android 0.7.0 · measured 2026-09-25 · 16.2 s speech clip · 52 frames rendered per second
Samsung Galaxy S25+Androidheld 10 min1.4× real time
Details for Samsung Galaxy S25+, AndroidSamsung Galaxy S25+ · essence2-android 0.8.1 · held 10 min · measured 2026-09-27 · 28.4 s speech clip · 37 frames rendered per second
In the browser
ConfigurationEssence 2, × real time
Chrome on Apple M4Web browser (WebGPU)1.7× real time
Details for Chrome on Apple M4, Web browser (WebGPU)Chrome on Apple M4 · web viewer · measured 2026-09-27 · 28.4 s speech clip · 43 frames rendered per second
Chrome on Apple M4Web browser (WebGPU)held 10 min2.1× real time
Details for Chrome on Apple M4, Web browser (WebGPU)Chrome on Apple M4 · web viewer · held 10 min · measured 2026-09-27 · 1800.2 s speech clip · 54 frames rendered per second
On a computer
ConfigurationEssence 2, × real time
Apple M4macOS · Swift package4.8× real time
Details for Apple M4, macOS · Swift packageApple M4 · Swift package 2.15.0 · measured 2026-09-24 · 28.4 s speech clip · 120 frames rendered per second
Apple M4macOS · CLI4.2× real time
Details for Apple M4, macOS · CLIApple M4 · CLI 2.8.1 · measured 2026-09-27 · 28.4 s speech clip · 106 frames rendered per second
Apple M4macOS · Python6.9× real time
Details for Apple M4, macOS · PythonApple M4 · bithuman 2.11.12 · measured 2026-09-25 · 28.4 s speech clip · 174 frames rendered per second
Intel Core i7-13700F (x86_64)Linux · CLICPU only (no GPU)2.0× real time
Details for Intel Core i7-13700F (x86_64), Linux · CLIIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 · 28.4 s speech clip · 50 frames rendered per second
Intel Core i7-13700F (x86_64)Linux · PythonCPU only (no GPU)1.9× real time
Details for Intel Core i7-13700F (x86_64), Linux · PythonIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26 · 28.4 s speech clip · 49 frames rendered per second
bitHuman cloud
ConfigurationEssence 2, × real time
NVIDIA RTX 4090Cloud API · GPU4.1× real time
Details for NVIDIA RTX 4090, Cloud API · GPUNVIDIA RTX 4090 · cloud API · measured 2026-09-27 · 28.4 s speech clip · 104 frames rendered per second
Apple M4 MaxCloud API · Apple silicon2.8× real time
Details for Apple M4 Max, Cloud API · Apple siliconApple M4 Max · cloud API · measured 2026-09-26 · 28.4 s speech clip · 71 frames rendered per second
x86 server CPUCloud API · CPUCPU only (no GPU)1.1× real time
Details for x86 server CPU, Cloud API · CPUx86 server CPU, CPU only (no GPU) · cloud API · measured 2026-09-25 · 28.4 s speech clip · 28 frames rendered per second

Expression 2

On a phone
ConfigurationExpression 2, × real time
iPhone 15iPhone · Swift package5.5× real time
Details for iPhone 15, iPhone · Swift packageiPhone 15 · Swift package 2.18.0 · measured 2026-09-27 · 16.2 s speech clip · 111 frames rendered per second
iPhone 15iPhone · Swift packageheld 10 min5.1× real time
Details for iPhone 15, iPhone · Swift packageiPhone 15 · Swift package 2.15.0 · held 10 min · measured 2026-09-25 · 16.2 s speech clip · 103 frames rendered per second
Samsung Galaxy S25+Android2.4× real time
Details for Samsung Galaxy S25+, AndroidSamsung Galaxy S25+ · expression2-android 0.4.10 · measured 2026-09-23 · 21.5 s speech clip · 48 frames rendered per second
Samsung Galaxy S25+Androidheld 10 min2.2× real time
Details for Samsung Galaxy S25+, AndroidSamsung Galaxy S25+ · expression2-android 0.5.2 · held 10 min · measured 2026-09-27 · 21.5 s speech clip · 44 frames rendered per second
In the browser
ConfigurationExpression 2, × real time
Chrome on Apple M4Web browser (WebGPU)1.9× real time
Details for Chrome on Apple M4, Web browser (WebGPU)Chrome on Apple M4 · web viewer · measured 2026-09-27 · 28.4 s speech clip · 39 frames rendered per second
Chrome on Apple M4Web browser (WebGPU)held 10 min2.0× real time
Details for Chrome on Apple M4, Web browser (WebGPU)Chrome on Apple M4 · web viewer · held 10 min · measured 2026-09-25 · 1800.2 s speech clip · 41 frames rendered per second
On a computer
ConfigurationExpression 2, × real time
Apple M4macOS · Swift package8.8× real time
Details for Apple M4, macOS · Swift packageApple M4 · Swift package 2.15.0 · measured 2026-09-24 · 28.4 s speech clip · 177 frames rendered per second
Apple M4macOS · CLI8.4× real time
Details for Apple M4, macOS · CLIApple M4 · CLI 2.8.1 · measured 2026-09-27 · 28.4 s speech clip · 168 frames rendered per second
Apple M4macOS · Python8.4× real time
Details for Apple M4, macOS · PythonApple M4 · bithuman 2.11.12 · measured 2026-09-25 · 28.4 s speech clip · 169 frames rendered per second
Intel Core i7-13700F (x86_64)Linux · CLICPU only (no GPU)2.2× real time
Details for Intel Core i7-13700F (x86_64), Linux · CLIIntel Core i7-13700F (x86_64), CPU only (no GPU) · CLI 2.8.1 · measured 2026-09-27 · 28.4 s speech clip · 44 frames rendered per second
Intel Core i7-13700F (x86_64)Linux · PythonCPU only (no GPU)2.3× real time
Details for Intel Core i7-13700F (x86_64), Linux · PythonIntel Core i7-13700F (x86_64), CPU only (no GPU) · bithuman 2.11.13 · measured 2026-09-26 · 28.4 s speech clip · 47 frames rendered per second
bitHuman cloud
ConfigurationExpression 2, × real time
NVIDIA RTX 4090Cloud API · GPU17.0× real time
Details for NVIDIA RTX 4090, Cloud API · GPUNVIDIA RTX 4090 · cloud API · measured 2026-09-23 · 28.4 s speech clip · 340 frames rendered per second
Apple M4 MaxCloud API · Apple silicon5.5× real time
Details for Apple M4 Max, Cloud API · Apple siliconApple M4 Max · cloud API · measured 2026-09-23 · 28.4 s speech clip · 111 frames rendered per second
x86 server CPUCloud API · CPUCPU only (no GPU)1.3× real time
Details for x86 server CPU, Cloud API · CPUx86 server CPU, CPU only (no GPU) · cloud API · measured 2026-09-24 · 28.4 s speech clip · 27 frames rendered per second

14 published configurations, each measured on the named hardware and release. Select Details for the release, date and clip.

Runs onHardwareEssence 2Expression 2
Cloud API · GPUNVIDIA RTX 4090104 fps 4.1×340 fps 17.0×
Cloud API · Apple siliconApple M4 Max71 fps 2.8×111 fps 5.5×
Cloud API · CPUx86 server CPU28 fps 1.1×27 fps 1.3×
macOS · CLIApple M4106 fps 4.2×168 fps 8.4×
macOS · PythonApple M4174 fps 6.9×169 fps 8.4×
macOS · Swift packageApple M4120 fps 4.8×177 fps 8.8×
Linux · CLIIntel Core i7-13700F (x86_64)50 fps 2.0×44 fps 2.2×
Linux · PythonIntel Core i7-13700F (x86_64)49 fps 1.9×47 fps 2.3×
iPhone · Swift packageiPhone 1554 fps 2.1×111 fps 5.5×
AndroidSamsung Galaxy S25+52 fps 2.0×48 fps 2.4×
Web browser (WebGPU)Chrome on Apple M443 fps 1.7×39 fps 1.9×

Each figure is one avatar session rendering as fast as the hardware allows. On the cloud API the service picks the tier for each session; the three Cloud API rows show each tier.

The sections below go platform by platform, on-device first. How we measure has the method, the releases measured, memory, the time to a finished video and the raw data.

Mobile

First a short burst on each phone, then one session held for 10 minutes.

Runs onHardwareEssence 2Expression 2
iPhone · Swift packageiPhone 1554 fps 2.1×111 fps 5.5×
AndroidSamsung Galaxy S25+52 fps 2.0×48 fps 2.4×

Measured in September 2026 on Swift package 2.17.3, Swift package 2.18.0, essence2-android 0.7.0 and expression2-android 0.4.10.

  • The first table is one render of a speech clip on a cool phone, as fast as the phone allows.
  • Some phone figures use a shorter speech clip than the other platforms; each cell’s clip is in performance.json.
  • Only an iPhone 15 and a Samsung Galaxy S25+ are measured. Other phones render at other rates.

Setup for each SDK: Apple, Android.

Held for 10 minutes

Runs onHardwareEssence 2Expression 2
iPhone · Swift packageiPhone 1533 fps 1.3×103 fps 5.1×
AndroidSamsung Galaxy S25+37 fps 1.4×44 fps 2.2×
Web browser (WebGPU)Chrome on Apple M454 fps 2.1×41 fps 2.0×

One session held open for ten minutes from a cool start on Swift package 2.15.0, essence2-android 0.8.1, expression2-android 0.5.2 and web viewer, rendering as fast as the device allows. Each number is the median 30-second stretch of the slowest of that row’s sessions (four on iPhone · Swift package, three on Android, three on Web browser (WebGPU) Essence 2, four on Web browser (WebGPU) Expression 2); the slowest single stretch was lower (iPhone · Swift package Essence 2 31 fps; iPhone · Swift package Expression 2 100 fps; Android Essence 2 32 fps; Android Expression 2 37 fps; Web browser (WebGPU) Essence 2 50 fps; Web browser (WebGPU) Expression 2 41 fps). A phone warms up over a long conversation and slows its processor to stay cool, so a kiosk or any screen that renders all day should plan on this number rather than the short-burst rate.

Memory over the ten minutes (the probe’s own reading at the start and the end of the held window): Android Essence 2 memory (PSS) 0.9 GB to 0.9 GB (+0 MB); Android Expression 2 memory (PSS) 0.7 GB to 0.8 GB (+34 MB); Web browser (WebGPU) Essence 2 Chrome tab memory footprint 3.1 GB to 3.0 GB (-55 MB).

The Web browser row is the engine’s render throughput with WebGPU in Chrome on an Apple M4, measured in a visible (headed) browser window. It is not the frame rate a visitor sees on the page, where the voice and the display share the browser with the engine.

Web

These figures are for the avatar rendering in the visitor’s own tab (render=local). By default the avatar renders on the cloud API and streams to the page.

Runs onHardwareEssence 2Expression 2
Web browser (WebGPU)Chrome on Apple M443 fps 1.7×39 fps 1.9×

Measured in September 2026 on the hosted web viewer.

  • Each figure is the engine’s render throughput with WebGPU in Chrome on an Apple M4, measured in an automated browser.
  • It is not the frame rate a visitor sees on the page, where the voice and the display share the browser with the engine.
  • Only Chrome on an Apple M4 is measured. Other browsers, other GPUs and devices without WebGPU are not.

Setup and the WebGPU check: Web.

Desktop

Pick the row for the product you use: the CLI, Python and the Swift package render at different rates on the same machine.

Runs onHardwareEssence 2Expression 2
macOS · CLIApple M4106 fps 4.2×168 fps 8.4×
macOS · PythonApple M4174 fps 6.9×169 fps 8.4×
macOS · Swift packageApple M4120 fps 4.8×177 fps 8.8×
Linux · CLIIntel Core i7-13700F (x86_64)50 fps 2.0×44 fps 2.2×
Linux · PythonIntel Core i7-13700F (x86_64)49 fps 1.9×47 fps 2.3×

Measured in September 2026 on CLI 2.8.1, bithuman 2.11.12, bithuman 2.11.13 and Swift package 2.15.0.

  • Each figure is one render of a reference speech clip, as fast as the machine allows, with nothing else running.
  • The macOS CLI Essence 2 figure was measured with BITHUMAN_THREADS=8; by default the CLI uses one thread per CPU it may use, up to 16. The Linux CLI uses default settings.
  • Only these two machines are measured: an Apple M4 Mac and an Intel Core i7-13700F desktop. Other processors render at other rates.

Memory per render is on How we measure. Setup for each product: CLI, Python, Apple.

Cloud

By default the service picks the tier for each session; each row is one tier. To benchmark one tier you can pin it (below); in production, let the service choose.

Runs onHardwareEssence 2Expression 2
Cloud API · GPUNVIDIA RTX 4090104 fps 4.1×340 fps 17.0×
Cloud API · Apple siliconApple M4 Max71 fps 2.8×111 fps 5.5×
Cloud API · CPUx86 server CPU28 fps 1.1×27 fps 1.3×

Measured in September 2026 on the cloud API.

  • Each figure is how fast one finished video is delivered, including encoding the video file, on a server with no other sessions.
  • With other sessions on the same server, a session can render more slowly than shown.
  • A live conversation plays at the model’s own rate, 25 fps for Essence 2 and 20 fps for Expression 2. Speed above that makes a video file finish sooner; it does not put more frames on screen.

Pin a tier for a benchmark

To measure one tier, append ?model= with a tier slug to the viewer or embed URL:

https://www.bithuman.ai/embed/A23WJF0199?model=expression-2-apple
ModelTier slugs
essence-2essence-2-gpu · essence-2-apple · essence-2-cpu
expression-2expression-2-gpu · expression-2-apple · expression-2-cpu
  • A recognized slug pins the session to that tier: if the tier is unavailable, the session fails rather than playing elsewhere.
  • An unrecognized slug is ignored and the session plays as usual. If a pin seems to have no effect, check the spelling.
  • To be told about a typo, set the embed token’s model field instead: an unknown value is refused with a 400 listing the accepted names when you mint the token.

In production, omit ?model= and let the service choose.