Changelog
Release notes and version history for the bitHuman platform.
Note Product-level changes only. For per-version notes, see the Python SDK release history and the SDK releases.
September 2026
The avatar’s source video plays in place, and long audio stops being cut short — Swift SDK 2.13.6 / Essence2 engine 1.7.0 (2026-09-16)
Package tag 2.13.6 on the SwiftPM package;
it ships Essence 2 engine 1.7.0. No Swift surface change — from: resolves it and
nothing in your code moves. Tags 2.13.3 and earlier stay published and keep resolving
to the engine versions they always did.
- Long audio stops being cut short, and this is the one to read. Driving 75 seconds of audio through Essence 2 in one call previously returned 1209 of the 1875 frames the audio entitles you to — and returned success while doing it. If you were driving long audio in a single call, you were losing the end of it with nothing to tell you. 1.7.0 returns 1885: every frame, plus the short tail the engine has always added. Shorter clips are unaffected and return exactly the frame counts they did before.
- About 1.5 GB less memory on a 1080p identity. Earlier releases expanded the identity’s whole source video into memory before the avatar could speak, and drew the face onto that copy. 1.7.0 decodes one frame ahead of what it is drawing and draws onto the decoded frame directly, in idle and in speech. Measured against the copy previous releases drew onto, the result reads 45 dB PSNR — the difference is not visible, and it was reviewed on a side-by-side before it shipped.
- The idle animation plays whole. It runs from the first frame to the last and wraps only at the authored end, where the clip is designed to be seamless, instead of cutting early.
- A frame-buffer correctness fix. A buffer could hand a reader the slot still being written — a torn frame carrying a valid label, about one walk in ten thousand where two readers are active. Apple’s path has a single reader and could not reach it; fixed anyway.
The measured iPhone 15 rate is on the performance page; it did not regress with the memory saving.
Every frame of a reply comes out, and the driver video plays in place — essence2-android 0.5.8 (2026-09-16)
ai.bithuman:essence2-android:0.5.8 on Maven Central. No Kotlin surface change;
0.5.7 and earlier stay on Central and are superseded. Read this if your app
feeds audio faster than real time, or if it has ever run 0.5.7 on the same device.
- Every frame of a reply comes out. In 0.5.7 the motion frontier advanced only
behind
feed()/endOfAudio(), one block per call: an app that pushed more than 320 ms of audio per call fell behind its own audio by the difference, and afterendOfAudioexactly one more block ever came out. Measured on a Galaxy S25+ through an un-paced transport, 0.5.7 delivered 72–77 % of every reply’s frames — the rest of the audio played under a frozen last frame. 0.5.8 extends the frontier on a thread of the SDK’s own, woken by a feed and by a pull that finds the queue low; the same handset and script deliver 1828 of 1827 expected frames over 8 replies, hold 0, stale 0, speaker under-runs 0. Nothing to change in your code; feed as the audio arrives. - The driver video plays in place — resident memory 2969 → 1613 MB on the same
identity. 0.5.7 expanded
target_frames.mp4to 251 JPEG files at the first session and decoded the whole clip into memory at every open (1.56 GB for a 1080p identity). 0.5.8 plays it through a decode cursor — the phone’s hardware H.264 decoder (MediaCodec) one frame ahead of the paste — and writes nothing into the bundle directory. The principle every bitHuman surface now follows: play video in place rather than loading it into memory. - A device that ran 0.5.7 keeps its JPEG files, and 0.5.8 walks them; a fresh install decodes the mp4. Both are correct. On the mp4 path the first frame of a reply arrives a little later and less evenly than on the JPEG path (same script, same handset: time-to-first-audio +26 to +284 ms across takes, up to 12 stale frames and 2 speaker under-runs per take against 0 / 0) — the reply’s first block waits on a seek to the clip’s keyframe and the decoder’s warm-up. The next release pre-seeks that block when an utterance opens.
- The picture moves by one lossy generation. The paste now lands on the decoded frame itself rather than on 0.5.7’s JPEG re-encode of it: 45 dB PSNR against the old canvas on the region the paste leaves untouched; the mouth region is unchanged.
- A torn-frame race in the driver cursor is fixed before it shipped. Two consumers reading the cursor 13 frames apart could receive a frame the decoder was still writing, under the right label, about once in 10,000 reads. No published Android artifact carried it (0.5.7 has no cursor); 0.5.8 does not either.
The idle clip plays whole, decoded in place — expression2-android 0.4.7 (2026-09-16)
ai.bithuman:expression2-android:0.4.7 on Maven Central. Read this before
upgrading if your app reads Expression2Avatar.idleLoop.
- The idle clip plays from its first frame to its last and wraps there. 0.4.6
held the clip’s first 48 frames as a
List<Bitmap>and wrapped at 2.4 s — a cut the clip’s author never made. 0.4.7 decodes the clip in place withMediaCodec, one frame at a time, and wraps where the file ends; resident memory is independent of the clip’s length. idleLoopchanges type. It is no longer aList<Bitmap>; takeExpression2IdleLoop(next(bitmap)draws the next frame into your bitmap and reports the wrap). Code that indexed the old list does not compile against 0.4.7.0.4.6and earlier stay on Central and are superseded, not withdrawn.
pip install bithuman resolves 2.11.0 — the 3.x line is withdrawn from PyPI, and a bithuman<3 pin gets the same engine (2026-09-16)
bithuman 2.11.0 is what PyPI serves now, to every resolver: an unconstrained
pip install bithuman, a bithuman<3 pin, and pip install livekit-plugins-bithuman
(whose own pin is bithuman<3,>=0.5.25) all resolve it, on Python 3.11, 3.12 and 3.13
from a fresh environment, and pip check is clean on each. The 3.x releases
(3.0.0 through 3.1.10) were deleted from PyPI on 2026-09-16; a pin on any of them no
longer resolves.
Why the number goes backwards. 3.0.0 cut the exported surface from 32 names to 7 and
announced the break with a major bump — so a bithuman<3 pin never received it, and
resolved 2.3.4 instead: an old engine, but a working one. 2.11.0 is the 3.x engine
(every native half byte-identical to 3.1.10’s, on all three platforms) published where
that pin can reach it, with the whole 2.x import surface carried alongside: every name
the published 2.10.0 wheel exported — AsyncBithuman, Bithuman, AsyncAvatar,
AudioChunk, VideoControl, VideoFrame, Emotion, the 2.x exception kinds and the
rest — imports and works, at package level, next to the open() / render() surface
the Python reference documents. Nothing is a stub: AsyncBithuman is
the streaming class our own serving binds to.
One name changed meaning and is said out loud: Avatar is what open() returns
(.render()), as in 3.x; the 2.x synchronous class stays reachable as Bithuman, and
bithuman.Avatar.load(...) raises rather than pretending. The default picture is the 2.x
one again — AsyncBithuman yields 1280x720 unless told otherwise — and Essence 2 opens on
the streaming class too.
The Python reference is regenerated from the 2.11.0 wheel. This is the Python library; the CLI, the Apple and Android SDKs and the browser build ship their own engine and are not covered by this note.
The idle clip plays whole, decoded in place, on iPhone and Mac — Swift SDK 2.13.5 / Expression2 2.6.3 (2026-09-16)
Package tag 2.13.5 on the SwiftPM package;
from: resolves it. It ships the Expression2 engine at 2.6.3. Read this before
upgrading if your app reads engine.idleLoop.
- The idle clip plays from its first frame to its last and wraps there. The
identity’s
idle.mp4is a 10 s, 200-frame loop authored so that its last frame leads into its first. 2.6.2 held the first 48 frames and wrapped at 2.4 s — a cut the clip’s author never made, visible as a jump every few seconds. 2.6.3 plays the video in place: one hardware decoder runs a few frames ahead of the display and the wrap is the file’s own end. The principle is the one every bitHuman surface now follows: play video in place rather than loading it into memory — efficient compute and efficient memory management, which only requires the right implementation. Measured on the published bytes on an M4 Mac: the loop wraps at frame 199 → 0 every 200 frames with a seam smaller than the step between two ordinary frames, a 60 s clip costs the same resident memory as the 10 s one (within 3 MB), and an idle frame costs about 0.3 ms. idleLoopis gone from the public surface.idleLoop: [[UInt8]]— the list that invited the cap — does not exist in 2.6.3; code that reads it does not compile. TakeidleNextPixelBuffer()(the decoder’s ownCVPixelBuffer, no copy — hand it to a texture or a sample-buffer layer) oridle(into:)(the same frame as BGR bytes).idleFrameCount,idleIndexandidleWrapssay where the loop is;idleUnavailableReasonsays why no clip plays when it does not.pullPos()is unchanged from 2.6.2:(frame, speech, isSpeech, pos).- The
BithumanEngineProtocolproduct drops itsidleLooprequirement and gainsidleNextPixelBuffer()with a default ofnil.
An utterance is exactly as long as its audio, and every frame says where it belongs — Swift SDK 2.13.4 / Expression2 2.6.2 (2026-09-16)
Expression2 2.6.2 was the first release since 2.6.0 whose engine bytes moved
(2.6.1 re-hosted 2.6.0’s archive byte-for-byte). What changed for an app:
- Tail 0 and head 0. A fed utterance is delivered as exactly
round(seconds × 20)frames — no invented frames after the audio ends — and the first frame is the first audio frame.pullPos()reaches the shipped interface for the first time:(frame, speech, isSpeech, pos), whereposis the frame’s own audio position in 16 kHz samples, so a presenter pairs a frame with its sound by arithmetic instead of by counting. - Back-pressure instead of silent discard. When the app stops pulling, the engine parks its producer at 64 queued frames rather than dropping the oldest; the shipped 2.6.1 binary destroyed 455 of 565 frames on a paced consumer.
isSpeechis per-frame voice activity — this frame’s own 40 ms of fed audio has energy — with one definition on every platform;speechbeside it is the legacy flag and keeps its old meaning.
The idle clip is an SDK member, the tail is 0, and isSpeech means one thing — expression2-android 0.4.6 (2026-09-16)
ai.bithuman:expression2-android:0.4.6 on Maven Central (0.4.1 and earlier stay
and are superseded; 0.4.5 was never published).
Expression2Avatar.idleLoop— the identity’s own idle clip, from the same store and manifest as the weights (idle.mp4). Read from the published AAR, in 0.4.6 it is aList<Bitmap>of the clip’s first 48 frames (IDLE_LOOP_FRAMES) — a cap copied from the Apple SDK’s old premise, which wraps a 10 s clip at 2.4 s. The next release replaces it with a cursor that plays the whole clip in place, decoded byMediaCodecone frame at a time; that changes the member’s type, and the note for it will say so.- Tail 0: a segment of n fed samples is delivered as
round(n / 800)frames and not one more; before, the last chunk always yielded 21 frames. Expression2Frame.isSpeechis per-frame voice activity with the same definition as the Apple SDK’s;audioSampleis the frame’s own 16 kHz position.
Renders longer than 48 seconds — bithuman 3.1.10 (2026-09-15)
bithuman 3.1.10 on PyPI. Essence 2 had a per-render maximum of 48.0 s / 1200
frames: the positional table has a fixed row count, an utterance has as many
frames as its audio, and the first bounded the second. Longer audio was
refused, with the remedy in the message — split it into parts of 48 s or less.
That bound is gone. Past the table the second half of the rows repeats, and the
first 1200 frames read the table exactly as before, so a render that fit under
the old ceiling is unchanged. Read from the published wheels: the refusal text
appears nowhere in 3.1.10, and _offline.py is byte-identical across the macOS
arm64 and Linux x86_64 builds.
The public API did not change. Every name, signature and exception on the Python reference is the same as the release before it; only the behaviour on long audio moved. That page is regenerated from these bytes.
This is the Python library. The CLI, the Apple and Android SDKs and the browser build ship their own engine and are not covered by this note.
BITHUMAN_UNMETERED is gone from the Android SDK, and a frame-source constructor loses an argument — essence2-android 0.5.7 (2026-09-15)
ai.bithuman:essence2-android:0.5.7 on Maven Central. Read this before
upgrading if you construct the frame source yourself.
- The unmetered development variable is gone. Read from the published AAR,
BITHUMAN_UNMETERED,UNMETERED_ENVand the unmetered banner appear zero times in 0.5.7; the same search finds the variable in 0.5.6, and findsSelfHostMeterandMeteringRefusedin both. With the CLI ignoring it from 2.6.20 and the public Python wheels refusing with or without it, no shipping surface now has an environment variable that renders free — see pricing. ElevateFramestakes one argument fewer. The class sits on theai.bithuman.elevatepackage — a legacy name kept for compatibility, which a developer still types. Its constructor was(String, String, int, String, boolean)in 0.5.6 and is(String, String, int, boolean)in 0.5.7 — the execution-provider string is no longer accepted. Code that passed it will not compile against 0.5.7; delete the argument. Nothing on this site taught that parameter.- The
0.2.0through0.5.6artifacts stay on Central and are superseded, not withdrawn.
Every render path needs a credential, on both platforms — cli-v2.6.20 (2026-09-14)
CLI cli-v2.6.20, macOS arm64 and Linux x86_64 built from one commit.
Read this before upgrading if anything you run renders without signing in.
bithuman runnow refuses without a credential on Linux too. It stops before serving a frame — exit 77,METERING_REFUSED, in about two seconds — where 2.6.19 on Linux rendered indefinitely. macOS behaves as it did in 2.6.19. The platforms word it differently and ask for the same two things: Linux says “the render host refused this session: its credential was rejected. Runbithuman login, or set BITHUMAN_API_SECRET to the account this session should be billed to.”; macOS says “refusing to serve: no api-secret is available, so this session cannot be attributed to an account … runbithuman loginor set BITHUMAN_API_SECRET to the API secret of the account this session should be billed to.”- No environment variable renders for free any more. The CLI ignores
BITHUMAN_UNMETERED=1completely: with it set,runstill exits 77 on both platforms, and no session prints an unmetered or not-being-billed banner. In 2.6.19 it still bought an unmetered Linuxrun. The Python SDK, the Docker container and the Swift SDK are unchanged — see pricing. bithuman renderis unchanged — exit 77,NOT_SIGNED_IN, no output file written, on both platforms, as in 2.6.19.--host 0.0.0.0is refused unless you say you meant it. It exits 2 withPUBLIC_BIND_REFUSEDand leaves nothing listening — “—host 0.0.0.0 binds every interface, which would expose this session to your whole network.” On 2.6.19 the same command bound the wildcard and served.--allow-public-bindstill opts in, and it still binds: with a credential and the flag,runlistens on0.0.0.0— checked against the kernel’s socket table, with the same credential and no flag refusing and listening on nothing. The refusal is a gate on one flag, not a blanket ban on exposing a session. An unparseable--hostnow exits 2 withBAD_HOST, where 2.6.19 exited 1.bithuman mcp toolsstill lists 28 tools, and the engine inside is still 3.1.8 (ABI 7).
bithuman render needs a credential, and on macOS so does bithuman run — cli-v2.6.19 (2026-09-14)
CLI cli-v2.6.19, macOS arm64 and Linux x86_64 built from one commit.
Read this before upgrading if anything you run renders without signing in.
- Every render needs a credential. A render stops before the first frame
with exit code 77, reported as
NOT_SIGNED_IN, and says “not signed in, or the credential is not valid — runbithuman login, or set BITHUMAN_API_SECRET”. That is what you get for all three cases: no credential, one this machine cannot use, and one the service rejects. Until now a rejected credential kept rendering for 300 seconds behind a countdown, and no credential at all rendered while saying it was not charging you. Getting a key is free and takes a moment:bithuman loginopens your browser, orbithuman login --deviceprints a code for an SSH session. bithuman runrefuses on macOS, and is not yet covered on Linux. Which half of the product meters the session differs by platform: the CLI does it on macOS arm64, the engine does it on Linux, and this release fixed the CLI half only.- macOS. No credential, or one the service rejects, ends the session in
a few seconds with exit 77 and
METERING_REFUSED— “refusing to serve: no api-secret is available, so this session cannot be attributed to an account”, or “refusing to serve: the API secret was rejected — revoked, or from another environment. (401)”. Nothing is served and no output is written. - Linux. It renders, and with no credential at all it keeps
rendering — ”★ UNMETERED RENDER: no BITHUMAN_API_SECRET is set, so this
render cannot be attributed to an account. Proceeding anyway — metering is
FAIL-OPEN” — with no countdown and no refusal. The 300-second grace
applies only to a credential the service actively rejects: that logs
”★ UNMETERED RENDER — CREDENTIAL REJECTED (401) … Rendering continues for
another 300 s of grace”, renders through four once-a-minute beats, and on
the fifth logs “REFUSED: the credential has been rejected (401) for
300 s” — the session then ends and the process exits 70, not 77.
Treat a Linux
runas unenforced until a release says otherwise.
- macOS. No credential, or one the service rejects, ends the session in
a few seconds with exit 77 and
BITHUMAN_UNMETERED=1is gone wherever the CLI is the meter — that isrenderon both platforms andrunon macOS, where setting it changes nothing. The engine meter that servesrunon Linux still honours it: that session prints ”★ BITHUMAN_UNMETERED is set — THIS RENDER IS NOT BEING BILLED. No usage will reach the ledger. This must never be set in production.” and renders.- An unreachable meter still renders, and is never refused. If our service
cannot be reached, the render continues and says so — being unable to ask is
not the same as being told no. An operator who wants a validated credential
before any frame sets
BITHUMAN_METER_ENFORCE=1, which refuses that case too. - Downloading is unchanged:
bithuman pull <slug>of a showcase avatar still needs no account. bithuman auth …is gone.auth login,auth logoutandauth statusduplicated the top-levellogin,logoutandwhoami— use those — andauth tokenis nowbithuman token. If a script callsbithuman auth, change it.- Failures that returned 64 start returning documented exit codes.
initoutside a terminal returns 2 (INTERACTIVE_ONLY); a script matching on 64 needs updating. Ctrl-C is now published as 130, which it always returned. Correction: this entry also named a bad host and a refused public bind here. Checked against the published binary, neither changed in 2.6.19 — a bad--hoststill exited 1 and--host 0.0.0.0still bound the wildcard. Both landed in 2.6.20, above. --jsonerrors carry ahintbeside the cause when there is a next step.bithuman mcp toolslists 28 tools, adding localpullandrender.- The engine inside moves to 3.1.8 (ABI 7).
bithuman 3.1.8: on a Mac, Expression 2 stops waiting at the end of each utterance (2026-09-14)
bithuman 3.1.8 on PyPI, for Apple Silicon macOS (14 or newer), Linux x86_64
and Linux aarch64 (CPython 3.10–3.14).
- The last frames of an utterance arrive as soon as they are ready. The Mac
wheels carry the render host
cli-v2.6.18ships, byte for byte, and that host no longer waits out a fixed pause before handing over the end of an utterance. - The output does not change. The frames are byte for byte what 3.1.7’s host delivered, checked frame by frame with a determinism control.
- Linux is unchanged.
Expression 2 renders much faster on a Mac, and every frame is what 2.6.17 produced — cli-v2.6.18 (2026-09-14)
CLI cli-v2.6.18, macOS arm64 and Linux x86_64 built from one commit. If you
installed 2.6.17, this is a drop-in upgrade with nothing to change on your
side.
- The end of a render no longer waits on a timer. On a Mac the renderer paused on fixed timers before handing over the last frames of a render, and under load those pauses grew past a second. It now hands them over as soon as they are ready. Measured frame rates are on the performance page.
- The output does not change. With the same avatar and the same audio, 2.6.17 and 2.6.18 deliver the same frames byte for byte — checked frame by frame over a full 28-second clip, with repeated runs of 2.6.17 agreeing with each other as the control.
- Linux renders are unchanged. The engine inside moves to 3.1.7 (ABI 7);
bithuman --versionprints it beside the CLI’s own version.
Earlier versions stay resolvable; a BITHUMAN_VERSION=cli-v2.6.17 pin keeps
working.
bithuman 3.1.7: Expression 2 on a Mac renders on the same CoreML engine as the macOS CLI (2026-09-14)
bithuman 3.1.7 on PyPI, for Apple Silicon macOS (14 or newer), Linux x86_64
and Linux aarch64 (CPython 3.10–3.14).
- On a Mac, Expression 2 renders on the same CoreML host and engine as the
macOS CLI. The Mac wheels carry the host
cli-v2.6.17ships, byte for byte, and a clean install renders the quickstart avatar on it: with logging atINFOthe library says “expression-2 on the CoreML host … the same host and engine as the macOS CLI”. Nothing you call changes. Linux is unchanged. - An avatar that cannot use CoreML now says why. If an avatar was packed without the Apple Silicon decoder, the library says so first and renders on the CPU instead, where the warning used to end in a cut-off host log.
- To see which version you have, run
python -c "from importlib.metadata import version; print(version('bithuman'))".
Expression 2 renders faster on a Mac, and every frame is what 2.6.16 produced — cli-v2.6.17 (2026-09-14)
CLI cli-v2.6.17, macOS arm64 and Linux x86_64 built from one commit. If you
installed 2.6.16, this is a drop-in upgrade with nothing to change on your
side.
- On Apple Silicon, an Expression 2 render now overlaps its steps. For each short stretch of video it used to produce the frames, finish them into pictures and hand them to the video encoder one after another. The next stretch is now produced while the previous one is finished and encoded.
- The output does not change. With the same avatar and the same audio, 2.6.16 and 2.6.17 deliver the same frames byte for byte — checked frame by frame over a full 28-second clip, with repeated runs of 2.6.16 agreeing with each other as the control.
- Linux renders are unchanged. The engine inside moves to 3.1.6 (ABI 7);
bithuman --versionprints it beside the CLI’s own version.
Earlier versions stay resolvable; a BITHUMAN_VERSION=cli-v2.6.16 pin keeps
working.
bithuman 3.1.6: Essence 2 uses more of your machine’s cores (2026-09-14)
bithuman 3.1.6 on PyPI, for Apple Silicon macOS (14 or newer), Linux x86_64
and Linux aarch64 (CPython 3.10–3.14). The Linux wheels published first and the
Mac wheels about half an hour later; if pip install bithuman on a Mac gave you
3.1.5 in that window, run pip install --upgrade bithuman.
- Essence 2 uses more of your machine’s cores. The thread count reaches the part of the renderer that does most of the work, which until now stayed at four whatever you asked for. The default becomes the smaller of your core count and 16.
- The thread count does not change a single frame. The release was tested at four, eight and sixteen threads, and every frame matched.
Essence 2 renders faster again on Linux, and every frame is what 2.6.15 produced — cli-v2.6.16 (2026-09-14)
CLI cli-v2.6.16, macOS arm64 and Linux x86_64 built from one commit. If you
installed 2.6.15 in the last hour, this is a drop-in upgrade with nothing to
change on your side.
- Building each finished frame now spreads across the threads the renderer already had. It was doing most of that work on one thread while the rest of the machine waited. On a 24-thread Intel desktop, building a frame drops from about 36 ms to about 26 ms at the old thread count, and from about 29 ms to about 17 ms at the thread count 2.6.15 introduced.
- The output does not change at all. Every frame is byte-for-byte what 2.6.15 produced — checked at four thread counts on three avatars, and over a 200-frame render where ten runs produced one identical result.
- macOS is unchanged, and so is the engine inside:
bithuman --versionstill reports 3.1.5 (ABI 7) beside the CLI’s own version.
Earlier versions stay resolvable; a BITHUMAN_VERSION=cli-v2.6.15 pin keeps
working.
Essence 2 on Android delivers frames sooner, and a model file that changed is picked up — essence2-android 0.5.6 (2026-09-14)
ai.bithuman:essence2-android:0.5.6 on Maven Central. Code written against
0.5.5 compiles unchanged — nothing was removed or re-typed, and the two
additions below are new types you can ignore until you want them.
- Frames arrive sooner, and the frames themselves are unchanged. The render now uses the phone’s GPU alongside its CPU, and a step that ran as several operations runs as one. The measured rate on a Galaxy S25+ is on the performance page.
- Interrupting the avatar stops it immediately. Until now, a barge-in delivered one more frame of the sentence it was already speaking.
- A model file that changes is picked up on the next session. When
Essence2ModelStore.fetch()finds an identity already on the device, it now asks the download service whether any member of it changed: if one has, the device installs the new file; if the service cannot be reached, it opens the copy it had already verified, so an app with no network keeps working. That costs one small request on a cache hit. Two new types report the outcome:Essence2ModelStore.Revalidated(UNCHANGED,UPDATED,KEPT_UNREACHABLE,KEPT_REFUSED,KEPT_FAILED) andMemberChanged. 0.2.0through0.5.5stay on Central;0.5.1and0.5.2cannot install a model on a handset. The FFmpeg relink materials for this version are on FFmpeg / LGPL, re-measured on the published0.5.6artifacts.
Essence 2 renders faster on Linux, and Linux video encoding costs far less processor time — cli-v2.6.15 (2026-09-14)
CLI cli-v2.6.15, macOS arm64 and Linux x86_64 built from one commit; it is
what Homebrew and the universal installer give you. Both changes are on Linux,
and the picture is the same.
- An Essence 2 render on Linux is about 1.5x faster. The frame-assembly work ran on four cores no matter what your machine had; it now uses more of them. Measured on a 24-thread Intel desktop at 1080p, under load. Frame rates for every platform are on the performance page.
- Video encoding on Linux takes about 60% less processor time.
bithuman rendernow compresses with the same encoder settings the bitHuman cloud uses. Both settings were measured on the same frames against an uncompressed reference: the picture is equal or slightly better, and files are about a fifth larger. - macOS is unchanged — it keeps the hardware video encoder 2.6.14 introduced.
- The engine inside moves to 3.1.5 (ABI 7).
bithuman --versionprints it beside the CLI’s own version, and the CLI page shows what that output looks like.
A second Linux speedup is already built and will arrive as its own release, with its own entry here. Nothing you install today needs changing for it.
Earlier versions stay resolvable; a BITHUMAN_VERSION=cli-v2.6.14 pin keeps
working.
Essence 2’s audio step does less work on iPhone and Mac — Swift SDK 2.13.3 (2026-09-14)
Swift package tag 2.13.3 ships Essence 2 engine 1.6.3. If you depend on
the package with from:, you already resolve it, and nothing in your code
changes.
- Essence 2’s audio step computes only the part of each audio window the renderer reads — the same change the CLI took in 2.6.13 and the Python library in 3.1.5. An Essence 2 render on an iPhone is substantially faster for it. The measured rate for each platform is on the performance page.
- Nothing else in the package changes. The
bitHumanKitandExpression2products are the same binaries 2.13.2 shipped; 2.13.2 shipped Essence 2 engine 1.6.2. - Essence 2 in your own iOS or macOS app still works from 2.13.2 — the install section has the pin.
Essence 2 on Android is fast at its default settings — essence2-android 0.5.5 (2026-09-13)
ai.bithuman:essence2-android:0.5.5 on Maven Central. If you are on 0.5.3,
change the version and nothing else — every class and method your code calls
is the same in 0.5.5.
- The default settings are now the fast ones. 0.5.3 left speed on the table unless you overrode its settings: its default thread count was too low for a current handset, and it assembled every 1080p output frame on the CPU. 0.5.5 corrects the thread default and assembles each output frame on the GPU of a Snapdragon (Adreno) handset by default. The frame rate a Galaxy S25+ reaches at those defaults, measured on the published library, is on the performance page.
- There is no 0.5.4 — it was never published, so 0.5.3 is followed by 0.5.5.
0.2.0through0.5.3stay on Central;0.5.1and0.5.2cannot install a model on a handset. The coordinate and the troubleshooting table are on the Android SDK page, and the FFmpeg relink kit for this version is on FFmpeg / LGPL.
bithuman 3.1.5: Essence 2’s audio step does less work, and a render that stops early raises (2026-09-13)
bithuman 3.1.5 on PyPI: pip install --upgrade bithuman. The same wheels as
3.1.4 — CPython 3.10–3.14 on macOS arm64 (14 or newer), Linux x86_64 and Linux
aarch64.
- Essence 2’s audio step computes only what the renderer reads. Each audio window is now processed only as far as the part the renderer actually uses, so the audio side of a render does less work than on 3.1.4. Delivered frames are unchanged — compared frame by frame against 3.1.4’s audio step on Linux before publishing. The library fetches the new audio files once, on first use, and checks them by digest.
- A render that stops early raises instead of returning a short video.
avatar.render(...)now checks that it delivered the frames your audio calls for, and raisesFailednaming both counts — “that render stopped early — N frames came out of the M this audio should produce” — so you can render again rather than ship a clip that ends before its audio does. On 3.1.4 such a render returned normally. A live stream has no fixed length and is never flagged. - If you pinned 3.1.4, move the pin. 3.1.4 keeps working and keeps fetching the audio files it was built for, but it runs the larger audio step.
bithuman render on a Mac compresses video on the Mac’s own hardware encoder — cli-v2.6.14 (2026-09-13)
CLI cli-v2.6.14, macOS arm64 and Linux x86_64 built from one commit; it is
what Homebrew and the universal installer give you. One change, and it is on
macOS: bithuman render compresses the output video on your Mac’s hardware
video encoder instead of on the CPU.
- Renders are faster, and your Mac stays usable while they run. The CPU was spending more time compressing the video than rendering it. On an Apple M4 the same render runs about 1.7x faster, and compressing it takes about half of one core instead of about five. Measured frame rates are on the performance page.
- The picture quality is the same. The encoder setting was chosen to match what 2.6.13 produced, measured frame by frame against an uncompressed reference — not to make the file smaller.
- Speed is steadier from run to run. Two identical renders used to differ by up to 1.67x in speed depending on what else the Mac was doing; now they differ by about 1%.
- Nothing fails without the hardware encoder. If your
ffmpegdoes not have Apple’s hardware encoder — a custom build, say — you get the CPU encoder as before. Every render prints which encoder it used, andBITHUMAN_FORCE_X264=1pins the CPU encoder. - Linux is unchanged: same encoder, same settings, same bytes. The engine inside is the same one 2.6.13 carried; only the CLI’s own version moves.
Known issue, and it is not new: on macOS, rendering the same input twice can
produce slightly different output — every frame stays in its place, but some
pixel values can differ slightly between runs. 2.6.11 through 2.6.13 behave
the same way. A fix is in progress.
Earlier versions stay resolvable; a BITHUMAN_VERSION=cli-v2.6.13 pin keeps
working.
Essence 2’s audio step does less than half the work — upgrade to cli-v2.6.13 (2026-09-13)
CLI cli-v2.6.13, macOS arm64 and Linux x86_64 built from one commit. If you
are on any earlier 2.6.x, this is the one to install: the speed-up below
reached almost nobody until it landed.
Rendering an avatar turns your audio into motion in short steps. Those steps now run over only the part of each audio window the renderer actually reads, so the audio side of a render costs less than half what it did. Every delivered frame is identical — we compared full renders frame by frame on macOS and on Linux before publishing.
Getting it onto developers’ machines took three tries, and the first two are why
2.6.13 exists:
cli-v2.6.10started fetching the faster audio files alongside the shared audio model, in the same download, verified by checksum. On Linux the files as first published could not be read by the runtime the CLI ships, so Linux kept the slower path; they were re-published in a form every runtime reads.cli-v2.6.11looked for those files beside whichever audio model the CLI actually resolves, so machines that already had the model — most machines — stopped being skipped.cli-v2.6.12made the step itself do less than half the work, but only replaced the files on a machine with an empty cache. If you had rendered before, you kept the older files and the older speed.cli-v2.6.13checks the files already on your machine against the ones it expects and replaces any that differ, on the first render. Nothing to configure.
Known issue, and it is not new: on macOS, rendering the same input twice can
produce slightly different output — every frame stays in its place, but some
pixel values can differ slightly between runs. 2.6.11 and 2.6.12 behave the
same way. A fix is in progress.
Earlier versions stay resolvable; a BITHUMAN_VERSION=cli-v2.6.12 pin keeps
working. Measured frame rates are on the performance page.
bithuman render --json reports its own steady-state rate (2026-09-13)
CLI cli-v2.6.9. render --json now includes render_fps beside
render_seconds: frames per second measured from the first audio pushed to the
last frame delivered, with model load and start-up excluded. It is the number
the performance page publishes, so you can reproduce that
figure on your own machine instead of timing the whole process yourself.
fps in the same object is unchanged and still means the output video’s
frame rate. Both fields are null rather than 0 when they cannot be measured.
The Apple tier’s force slugs are essence-2-apple and expression-2-apple (2026-09-13)
The ?model= slug that pins a session to the cloud’s Apple tier is now spelled
after the tier — essence-2-apple and expression-2-apple — on
embed, viewer and share URLs, and in the
tier tables. Nothing you
already have breaks: the older essence-2-ane and expression-2-ane
spellings stay accepted forever, so saved links, embeds and share tokens keep
routing to the same tier.
bithuman render is about 1.5x faster on Apple Silicon and 1.3x on Linux (2026-09-12)
CLI cli-v2.6.8, macOS arm64 and Linux x86_64 built from one commit, published
2026-09-12. The version installed before it is 2.6.6 — cli-v2.6.7 was built
and withdrawn before it was ever published, so the entry below first reaches
developers here. From the release notes:
renderwas opening your avatar on the wrong render model.bithuman renderandbithuman runare the same engine on the same avatar, and they asked it for different render models:runthe one-frame model,renderthe multi-frame one driven a frame at a time, which is the shape it is worst at. The model is now chosen once, in one place, for both commands, and a check outside both refuses any future build that reintroduces the split. Against 2.6.6, end to end,renderis about 1.5x faster on Apple Silicon and about 1.3x faster on Linux; the delivered MP4 is bit-identical.- Your very first render is slower, and only your first — it fetches the shared audio model and builds the accelerator’s compiled-model cache. On a short clip most of the Mac wall is one-time work, so longer clips converge on the faster render-loop rate.
- The last CPU step of the renderer moved to the GPU on Macs. Writing each face region back into the delivered frame now runs as a Metal compute pass, byte-exact against the CPU path, with the on/off switch deleted rather than defaulted.
- What did not change: the bundled engine core stays at 3.1.3 (ABI 7), so
there is no new Python wheel;
bithuman renderstill refusesessence-1avatars on both platforms, naming the Python package to use instead. The macOS tarball is Developer ID signed and notarized.
cli-v2.6.6 stays resolvable; a BITHUMAN_VERSION=cli-v2.6.6 pin keeps
working. Measured frame rates are on the performance page.
The CLI’s Essence 2 engine is BUILT again, and its ffmpeg libraries travel with it (2026-09-11)
cli-v2.6.7was built and withdrawn before it was published; everything in this entry ships incli-v2.6.8(above).
CLI cli-v2.6.7, macOS arm64 and Linux x86_64 from one commit
(66613942f5f0). The essence engine core moves to 3.1.3 (ABI 7).
- An offline
essence-2render is 3.2x faster. The engine core inside the 2.6.6 tarball (lib/lible_core.*) was a hand-pinned build that carried none of the engine’s vector routines — measured on the published 2.6.6 asset itself: 885,304 bytes, zero vector targets. 2.6.7 builds that core from the same commit the release pins: 1,137,920 bytes, eight vector targets. Same machine, same avatar, same held-out human voice, 100 frames of 1920x1080, unpaced: 1.08 fps → 3.48–3.77 fps on Linux x86_64, and 8.59–8.71 fps on an Apple M4.essence-2through the CLI is still slower than realtime; this release is a correctness fix, not the end of that work. - The Linux tarball now carries the ffmpeg libraries its own engine links.
It shipped
libavutil.so.59alone; the engine also needslibavcodec.so.61andlibavformat.so.61, so on a machine without a system ffmpeg 7 anessence-2render exited 69 rather than rendering. All three now travel in the tarball. The release script’s dependency walk had been running with the vendor directory off the loader path — it saw “not found”, copied nothing, and reported success; it now refuses to build a tarball with any unresolved library. expression-2on macOS, re-measured on the published arm64 tarball itself: 300 frames of 1280x720 at 25 fps, 30.5 fps, mouth-to-audio lag 0 frames, and the receipt’sframesequals the frame count in the file.
cli-v2.6.6 is marked superseded; its assets stay downloadable, and a
BITHUMAN_VERSION=cli-v2.6.6 pin keeps working.
A free avatar by CODE, a clip that is in sync, and an Android default that is fast (2026-09-11)
CLI cli-v2.6.6, ai.bithuman:expression2-android:0.4.1 and
ai.bithuman:essence2-android:0.5.2 on Maven Central, and bithuman 3.1.2 on
PyPI.
bithuman pull <CODE>gets you any avatar in the gallery, with no account, no key and no credits. Every rowbithuman listprints carries the CODE to pull —pulltakes it directly, writes the.imxand tells you whether your machine can render it ("runnable_locally": true). The CLI page is two commands from nothing to a talking face.- A rendered clip’s mouth is in sync with its audio. Before this release
bithuman renderput ten warm-up frames — 400 ms of a closed, still face — in front of the first spoken frame, one constant offset for the whole file, on every platform, and the receipt reported every frame as good. If you trimmed 400 ms of audio or nudged a track in an editor to compensate, undo it. The finished file also holds every frame the receipt counts, andrender --jsonnames what it left out:lead_in_frames_dropped. - An avatar file you downloaded yourself renders on Linux.
renderandrunon a file fetched with a browser orcurlused to exit 69 naming a shared engine file as missing — a file the installer had already put beside the binary. They take the copy that ships in the tarball now, and it never reaches the network to do it. Andbithuman run <CODE>on Linux x86_64 no longer opens a cloud session for an avatar that same machine renders in seconds:runasks the questionrenderasks, so the two cannot disagree. - On Android, a bare
Expression2Options()is the fast one. Depend onexpression2-android:0.4.1and the defaults ask for the accelerator — the measured frame rate on a current handset is on the Android SDK page. On0.3.1the same bare options stayed on the CPU and you had to name a routing to get off it.essence2-androidis0.5.2. pip install bithumanis 3.1.2. The CLI is not on PyPI: install it with Homebrew or the universal installer (the CLI page).
Android — the Kotlin example is a whole project, and essence-1 2.3.6 cannot authenticate on a phone (2026-09-09)
Two Android findings from a walk of the published pages on a Galaxy S25+
(SM-S936U1, Snapdragon 8 Elite, Android 16).
- Kotlin / Android — Hello, avatar is now a
complete project rather than fragments — every file in full, in the order you
create them, plus the audio reader and the playback clock. Parsed back out of
the served page into an empty directory and run:
app-debug.apk3,474,583 B, 5.72 s of speech → 117 frames of 416×720,acc=CPU. Two toolchain steps the site never wrote down (ANDROID_HOME/local.properties, and a JDK 17 launcher — AGP 8.7.3 rejects a newer one with an error whose whole body is the string26.0.2.1) are now in Install. ai.bithuman:sdk:2.3.6(essence-1) resolves, compiles and installs, and then cannot authenticate on a device.Avatar.loadthrowsbe_auth_authenticate: status=11 … SSL peer certificate … was not OK: the published native library carries no CA trust store. It is not your network and not your key — the same handset reached that exact endpoint over a public Google Trust Services chain in the same minute — and there is no app-side workaround on this version. Use expression-2 on Android, which needs no key. The measurement.
CLI 2.6.4 — a rejected key gets 300 seconds, then the session stops (2026-09-07)
cli-v2.6.4 (published 2026-09-07 23:30Z on the Homebrew tap; the formula
pins it) — bithuman-x86_64-unknown-linux-gnu.tar.gz (sha256
42094b2c912b3b3b4be364aed18893892d247d1e1070c9bb82225fa5e7f26f1a) and
bithuman-aarch64-apple-darwin.tar.gz (sha256
ed827aaa0b3918100e6c6776ca0527d7b7cabb8e4618f3ce91ef437f205f1bbc, Developer
ID signed and notarized, verified quarantined), both from one commit
(01325a3). Engine core unchanged.
- A key the service rejects gets a grace of 300 seconds, then the session
stops. The CLI checks your key with the service when a self-hosted
session starts and once a minute while it runs. If the service cannot be
reached (no network, a timeout, a 5xx on our side) the session renders,
prints a loud
★ UNMETERED RENDERline saying why, and keeps trying — it never stops for this, however long it lasts. If the service rejects the key (HTTP 401, 402 or 403 — revoked, from another environment, or out of credits) the session keeps rendering for 300 seconds from the first rejection, prints a line once a minute naming the seconds of grace left and the fix, and re-checks the key every minute; a key accepted again clears the clock, and a key still rejected at 300 seconds stops the session —runcloses the preview andrenderexitsMETERING_REFUSED(77) with no output. Before 2.6.4 a rejected key rendered on indefinitely behind the loud line. Measured on the published tarballs from a fresh home directory, on Linux x86_64 and on an Apple Silicon Mac, three sessions each: with an invented key the session printed five countdown lines and ended by itself 306 s (Linux) / 305 s (macOS) after it came up; with a good key revoked once the session was up, the first rejected check started the clock and the session stopped 300 s after it, exit 77, on both; with the metering service unreachable the session was still rendering 345 s later with no rejection and no refusal. The published 2.6.3 tarball, same invented-key arm, was still rendering after 420 s — the control. - Billing is unchanged. A live session bills wall-clock, a render bills the clip it writes, a download is free — as in 2.6.3. On the published 2.6.4 Linux tarball a 92 s Essence 2 session recorded 91.8 s in two acknowledged beats.
- The same rule, with the same number, applies to the Python package (3.0.4) and to the Apple engine below; for the Android SDK it is landed in the source and ships in the next coordinate (details).
Apple engine essence2-v1.4.0 / Swift package 2.10.0 — the same 300-second rule (2026-09-07)
essence2-v1.4.0 (published 2026-09-07 23:37Z; Package.swift at tag
v2.10.0 pins the engine archive at checksum
75b1919b848a0a8e13bdfe51999739813b610a42dad25d9fc5a3a4e408e29808; ONNX
Runtime and the resources archive carried forward byte-identical from
essence2-v1.3.0). A key the service rejects renders for a grace of 300
seconds from the first rejection behind a line once a minute, re-checked every
minute; still rejected at 300 seconds the engine stops —
be_essence2_pull_frame and be_essence2_idle_frame return -3 from then on.
A meter that cannot be reached still never stops a render. Billing is
unchanged from essence2-v1.3.0. Details on the Swift page.
CLI 2.6.3 — a live self-hosted session is billed on wall-clock (2026-09-07)
cli-v2.6.3 (published 2026-09-07 12:59Z on the Homebrew tap; superseded by
2.6.4) — bithuman-x86_64-unknown-linux-gnu.tar.gz (sha256
bf2c7b6414ed9d2fe8e00db929471ce82f405159c58c051733f3de6fdb94ecd6) and
bithuman-aarch64-apple-darwin.tar.gz (sha256
14ee0490a6bec87f26357bcdeb77160834ffdad6434200d77fdc3c806d043506, Developer
ID signed and notarized), both from one commit (b7a1005). Engine core
unchanged.
- A live self-hosted session bills wall-clock, which is what the pricing
page defines.
bithuman run <code>.imxon an Essence 2 or Expression 2 avatar bills the seconds the session was live, idle animation included, at 2 credits per minute (pricing); an offlinebithuman renderstill bills the duration of the clip it writes;bithuman pullis still free. 2.6.2 counted frames delivered ÷ fps instead, so a preview on a machine whose engine paints below nominal fps under-claimed — measured on an Apple Silicon Mac, the published 2.6.2 binary held a 92 s Essence 2 session and recorded 8.0 s of it, for 0 credits. On the published 2.6.3 tarball, same machine and clip, the same session records 92.1 s. Verified on the published bytes from a fresh home directory on Linux x86_64 and on an Apple Silicon Mac, both families, with the 2.6.2 macOS binary as the control — the self-host guide. - The live preview holds its nominal frame rate. On some Macs 2.6.2’s
preview settled at about a third of nominal with no viewers and an idle
engine, because it trusted
sleepto return on time and never made up a late wake. The preview now paces on an absolute clock: measured on the Mac that had it, Expression 2 settles at 20.0 fps against a 20 fps target where it previously ran at 5.8. Where the engine itself is the limit the preview still runs below nominal — the pacer cannot invent frames — but it no longer adds delay of its own. - The Linux tarball’s
PROVENANCE.jsonsays whether its tree was clean (dirty:false) and names the engine SDK revision, as the macOS half already did, plus the source revision of the Expression 2 render host it carries. - The release lane is tracked in the CLI repository, so the release and the proof that drives it can be re-run by someone other than the person who cut it.
CLI 2.6.2 — self-hosted sessions on macOS are metered, and the help tells the truth (2026-09-07)
cli-v2.6.2 (published 2026-09-07 08:49Z on the Homebrew tap; superseded by
2.6.3) — bithuman-x86_64-unknown-linux-gnu.tar.gz (sha256
1248f34feea05c8f3adab0312376e3704643296ce8012718c7a75aba032db03f) and
bithuman-aarch64-apple-darwin.tar.gz (sha256
0dab98763ecf7b25414abfbfecfc9071edcbda859eb88616e6dc1942b6373025, Developer
ID signed and notarized), both from one commit (679b9a6). Engine core
unchanged.
- A self-hosted essence-2 or expression-2 session is billed at the
published self-hosted rate on macOS and Linux alike — 2 credits per minute
(pricing). Before 2.6.2 only expression-2 on Linux was
metered:
bithuman run <code>.imxandbithuman renderon an essence-2 model were not metered on any platform, and on a Mac no session was. A credit minute is the pricing page’s — “wall-clock time a session is live and the engine is rendering”, idle animation included — and an offlinebithuman renderbills the duration of the clip it writes (the definition); one usage row per session; downloading a model is free. Known gap, fixed in 2.6.3: cli-v2.6.2 counts frames delivered ÷ fps as the served time, which under-counts a preview that paints below nominal fps. Metering never stops a render: with no sign-in, or a rejected or depleted key, the session renders behind a loud★ UNMETERED RENDERline, andBITHUMAN_METER_ENFORCE=1turns those three cases into a refusal. Proven on the published tarballs from a fresh home on Linux and on an Apple Silicon Mac, with the 2.6.1 macOS binary as the silent control — the self-host guide. bithuman run --helpsays where a model renders. It no longer claims essence-2 / expression-2 have “no local runtime yet” (false since 2.6.1): a local.imxrenders on this machine for essence-1, essence-2 and expression-2 (bithuman pull <CODE>fetches one); an agent code for essence-2 / expression-2 opens a live cloud session, as before;--cloudforces one. The routing did not change; the words did.bithuman doctoron macOS no longer tells you topip install bithuman-cliwhen the conversation worker is not set up yet — the firstbithuman runsets it up.- The Linux tarball’s build time is no longer in the future. 2.6.1’s
PROVENANCE.jsonsaid 2026-09-08;built_atis now the commit’s time on both platforms, and a tarball stamped ahead of the clock is refused before it can be published. - A live preview ends its session cleanly on the first Ctrl-C.
CLI 2.6.1 — essence-2 renders locally, on Linux and on macOS (2026-09-07)
cli-v2.6.1 (published 2026-09-07 05:04Z on the Homebrew tap; superseded by
2.6.2) ships the essence-2 runtime inside the CLI tarball on both
platforms — bithuman-x86_64-unknown-linux-gnu.tar.gz (sha256
5aef085a0686fc4f05b83ee50a26262a522f0f232c8f8717d157d29a5b18c1d6) and
bithuman-aarch64-apple-darwin.tar.gz (sha256
0fb359a8b709e2606af1f7e26b1df1ac705c8dc1da6f954131641b012d4f953c, Developer
ID signed and notarized). A downloaded essence-2 avatar now renders on your
own machine, offline, exactly the way expression-2 already did:
bithuman pull <AGENT_CODE> --model essence-2 # → <AGENT_CODE>.imx
bithuman render <AGENT_CODE>.imx -a speech.wav -o out.mp4 # exit 0; 5 s of audio → 125 frames at 25 fps
bithuman run <AGENT_CODE>.imx # local server
- The shared audio encoder is fetched on the first essence-2 render —
about 377 MB, once per machine, from the public release coordinate, checked
by content digest, into
~/.bithuman/engines/essence-2/— and reused after that. Nothing to stage by hand, no environment variable, no extra install step. The first play performs a licence check with the cloud, so it needs the sign-inpull <AGENT_CODE>already needs. - Fail-closed. An essence-2 model file that is incomplete — a required model member missing — is refused with exit 69 and no output file. The CLI never substitutes a generated mouth for the one the avatar recorded.
- expression-2 is unchanged. The 2.6.0 line below — “essence-2 does not render locally from these tarballs” — is closed, and so is every earlier note about a runtime to stage beside the binary.
- Proven from the published tarball alone — fresh home directory, empty environment — on Linux x86_64 and on an Apple Silicon Mac: Verified transcript.
essence-2 reaches Apple and Android as public coordinates, and the CLI moves to 2.6.0 (2026-09-07)
One line per release, each dated from the release itself and each checked anonymously on 2026-09-07 before it was written here:
- 2026-09-06 16:12Z — essence-2 Apple engine
essence2-v1.1.0. The first public release of the on-device essence-2 engine for iOS and macOS: an engine archive and an ONNX Runtime archive, each with a.sha256sidecar, fetchable with no credential. It refuses rather than drawing a mouth the avatar never recorded. - 2026-09-06 16:42Z — Swift SDK
v2.7.0. Adds theEssence2product (manifest only), pointing atessence2-v1.1.0. Its only module wasCLibEssence2. - 2026-09-06 18:50Z — the essence-2 shared audio encoder is published on a public release coordinate (377,625,424 B, SHA-256
95c35c86…), so the CLI and the Python package fetch it themselves instead of asking you to find it. - 2026-09-06 21:32Z — CLI
cli-v2.6.0, macOS arm64 and Linux x86_64 from one commit, with aPROVENANCE.jsonin each tarball. Fixed:bithuman runandbithuman pullagree on a container’s name;bithuman renderon macOS no longer truncates a clip; a refused render leaves no file behind. Added: the shared audio encoder is fetched once per machine and digest-checked on every use. Known and stated in the release: essence-2 does not render locally from these tarballs (renderexits 69 for it) — closed the same day bycli-v2.6.1; expression-2 renders locally on both platforms. The Python extra for offline rendering is now spelledbithuman[offline]. - 2026-09-06 — Android
ai.bithuman:essence2-android:0.3.0. The first essence-2 AAR whose engine refuses, with a thrown exception, rather than drawing a mouth of its own — but it judged a bundle by a descriptive list in its manifest and refused complete bundles. Superseded the same night; do not build against it. - 2026-09-08 02:17Z — Android
ai.bithuman:essence2-android:0.5.1. The version to use. A self-hosted session is metered at the published rate —0.5.0(2026-09-07 15:21Z) was the first Android version to meter at all; every version through0.4.0rendered free — and0.5.1adds the one rule every bitHuman runtime follows when the key check does not come back clean: a rejected key renders for a five-minute grace behind a countdown line, then every render call throwsMeteringRefused; a metering service that cannot be reached never stops a render. Driven on a Galaxy S25+ through the exact bytes uploaded, before the press: five arms green, including the invented key refused at 300 s and the unreachable service still rendering at 345 s; the same arm on0.5.0went red.0.2.0through0.5.0still resolve — use none of them. Android SDK. - 2026-09-07 01:20Z — essence-2 Apple engine
essence2-v1.2.0. One rule for when a model renders — all four recorded-mouth files present, or a refusal naming the missing one — on every platform;import Essence2compiles; the resources archive rides on the same release, so the coordinate is complete on its own. - 2026-09-07 01:30Z — Swift SDK
v2.8.0.Essence2points atessence2-v1.2.0. Pinfrom: "2.8.0". Resolves and builds for iOS device, iOS simulator and macOS from a consumer outside any bitHuman repository. Still missing: an in-app model download route that accepts a runtime token — Essence 2 on-device. - 2026-09-07 01:45Z — Android
ai.bithuman:essence2-android:0.4.0. The version to use. The same one rule as the Apple engine; an in-SDK model store (Essence2ModelStore— no default host yet, you pass the mirror); the product-named Kotlin packageai.bithuman.essence2beside the legacyai.bithuman.elevate, kept for compatibility;INTERNETmerged into your app. Driven on a Galaxy S25+ through the published bytes, 11 of 11 tests green.0.2.0can show a mouth the avatar never recorded without telling you and0.3.0refuses complete bundles — use neither. Android SDK. - 2026-09-07 — Python
bithuman3.0.0. A clean break: thirty-two public names become eight (bithuman.open,Avatar,Avatar.render,AvatarError,InvalidAvatar,NotSupported,NotAuthorised,Failed), frames are RGB, the key comes fromBITHUMAN_API_SECRETonly, and essence-2 and expression-2 open through the same call on macOS and Linux (bithuman[expression-2]for the latter). An essence-2 avatar missing its recorded-mouth data is refused atopen. The offline route isbithuman.offline/bithuman[offline](the 2.x spellings warn until 4.0.0), and the shared audio encoder is fetched and digest-checked for you.pip install "bithuman<3"stays on 2.9.0. Python SDK. ai.bithuman:expression2-android:0.3.1(2026-09-04) is unchanged and current — see its entry. The Kotlin hello page now carries an expression-2 and an essence-2 example, both compiled against the published AARs: Kotlin / Android — Hello, avatar (rewritten on 2026-09-09 as a complete project).
Also corrected on 2026-09-07: the scope matrix on where each model runs no longer carries a “Not ruled” column — the cloud CPU serving tier is in scope for essence-2 and expression-2 by the 2026-09-04 dispatch ruling, and essence-1 is served from the cloud’s Apple tier only (2026-09-05).
Swift SDK 2.6.0 — Expression 2 can be handed a model (2026-09-06)
Tag v2.6.0 on the SwiftPM package. Pin from: "2.6.0". Everything in it
is additive: if you are on from: "2.5.0" or from: "2.5.1" you pick it up
automatically and nothing you have written stops compiling.
★ The engine can now be given a model. Through 2.5.1 the only initializer
was Expression2Engine(), which searched an environment variable or the app
bundle and left isReady == false when it found nothing — so an app that had
downloaded its own avatar had no way to point the engine at it. 2.6.0 adds:
Expression2Engine.create(modelPath:sharedEngineDir:warmSpeech:)and the instanceload(modelPath:…);Expression2Engine.create(avatarContainer:sharedEngineContainer:sharedEngineDir:stagingDir:warmSpeech:), which opens the<code>.avatarthatGET /v1/agent/{code}/model/downloadreturns;Expression2Container—isContainer,members(of:),read(_:from:),readManifest(_:),unpack(_:to:)— plusExpression2ContainerError(the file is wrong) andExpression2LoadError(the contents are wrong, includingnotAnAvatarDirectory(path:)).
This withdraws a sentence this site published. The Swift SDK page said there was “no supported way to hand them to this product … no unpacking route is published or supported”. True of 2.5.x; false as of 2.6.0, and the page has been rewritten rather than softened.
★ Three binary targets now, not two. UnifiedModelHeader.xcframework ships
with this release because the engine’s own module interface imports it. You
never write that import — attach the Expression2 product and all three
targets come with it. A hand-rolled dependency on only the two 2.5.0 targets
fails at import with no such module 'UnifiedModelHeader'.
bitHumanKit is untouched: still the v2.4.0 asset, same URL, same
checksum. Expression2 still ships no model weights, so resolving it does
not by itself get you a rendering avatar — what changed is that you can now hand
it one.
Measured on the published zips, downloaded anonymously and re-hashed against the
checksums the manifest pins (all match) — ios-arm64 slice, aggregated over the
nine emitted .swiftinterface files, with two unchanged symbols and a nonsense
token as controls:
token v2.5.0 v2.6.0
create(modelPath 0 9
Expression2Container 0 45
notAnAvatarDirectory 0 9
public init() (control) 9 9
a token in neither (control) 0 0
expression-2 Android is 0.3.1, and google() is no longer required (2026-09-04)
ai.bithuman:expression2-android:0.3.1 reached Maven Central at
2026-09-04T11:47:17Z (maven-metadata.xml latest/release = 0.3.1),
2,742,085 B. It supersedes 0.3.0 as the version this site documents.
What changed for a consumer. 0.3.1’s POM declares only
org.jetbrains.kotlin:kotlin-stdlib:2.0.21. 0.3.0’s also declared
com.google.ai.edge.litert:litert:2.2.0, which is 404 on Maven Central, so
a 0.3.0 build needed google() in its repositories and failed at
checkReleaseAarMetadata without it. mavenCentral() alone now resolves it.
A build pinned to 0.3.0 still needs google() — Central never replaces a
published POM.
What did not change. The engine is the same binary: libexpr2jni.so is
446,200 B in both and differs in exactly 20 bytes at offsets 736–755 (the
GNU build-id), and the bundled libLiteRt.so (5,508,376 B) is byte-identical.
classes.jar goes from 32 to 41 entries, adding nine Bhci* classes and
removing none — so every measurement this site published against 0.3.0 still
describes 0.3.1.
★ Still arm64-v8a only, as are all three ai.bithuman Android artifacts.
An x86_64 emulator resolves and installs and then throws
UnsatisfiedLinkError at the first System.loadLibrary; use a physical arm64
device or an arm64-v8a system image.
Essence 2’s head upsample reaches the Android path, and the phone figure is measured (2026-09-03)
★ Correction to the 2026-09-02 entry below. That entry announced a rebuilt head-upsampling step and published a 3.00× CPU speedup. Both statements stand, but the entry let a reader infer something that was not true: that an Android device got the win. It did not.
Why not. An identity’s renderer ships as two graphs — a batched one and a single-frame one — and they are rewritten independently. The 2026-09-02 rollout rewrote only the batched graph. Android runs the single-frame graph (its batched path is unavailable while the sharp mouth-interior pass is attached, and that pass is always attached), so the change reached nothing a phone executes. The single-frame graph was rewritten and deployed on 2026-09-03, on the same single identity. The rewrite is now on 1 of 52 published identities, in both graphs; the other 51 have it in neither.
And now it is measured on a phone. Both graphs, benchmarked head to head on
a Galaxy S25+ (SM-S936U1, Snapdragon 8 Elite / SM8750), ONNX Runtime
1.26.0 CPU execution provider, batch 1 (the shape Android runs), 4
intra-op threads pinned to the four big cores, screen held awake, 8 interleaved
and rotated repeats, medians, 80 of 80 samples passing the clock and contention
guards. This is an offline benchmark of the renderer graph — not a live
session, and not a run through the published Android SDK’s own API.
| Renderer graph | Cooled ms/frame | fps | RTF | Sustained ms/frame | fps | RTF |
|---|---|---|---|---|---|---|
| Previous step (cubic resize) | 109.23 | 9.15 | 2.73 | 160.04 | 6.25 | 4.00 |
| Rebuilt step | 49.00 | 20.41 | 1.23 | 83.38 | 11.99 | 2.08 |
2.23× cooled, 1.92× sustained. The deployed identity, benchmarked as itself, agrees to 0.27% / 0.30%. Both controls fired: a byte-identical duplicate measured 1.003× / 1.016×, inside the floor; a deliberately 33.8%-heavier arm measured slower, 0.968× / 0.938×. The sustained figure is a lower bound — the burn preceding it is fixed work, not fixed time, so the faster arm enters its window hotter (50.7 °C vs 47.8 °C) and against a lower clock ceiling (1.958 vs 2.438 GHz).
★ Essence 2 still does not render in real time on a flagship phone. Even cooled and rebuilt, 20.41 fps is below the 25 fps a session consumes (RTF 1.22); sustained it is 11.99 fps. Against the internal real-time bar (RTF ≤ 0.50, ≥ 40 fps) the rebuilt graph is 2.45× short cooled and 4.17× short sustained.
★ The 7.54 → 23.87 fps figure is not an Android figure. It is a developer workstation — Threadripper PRO 5955WX, x86-64, batch 24 — and it is neither a phone nor the deployed CPU worker. The handset rows above supersede it for every on-device claim. The share of the forward pass taken by the replaced step is 68.6% on that x86 part but 57.9% on the Snapdragon, which is why the same rewrite is worth 3.00× there and 2.23× here; shares and speedups do not carry across silicon.
Unchanged, and stated again so it is not read as a general speedup: on the GPU tier the rewrite is 0.971× — about 3% slower, and no GPU speedup should be expected or quoted. On the Apple tier there is nothing to gain: that build’s converter has expressed this step the rewritten way since 2026-08-30, and the step is 6.06% of that tier’s forward pass, which caps any work on it at 1.06×.
No other Android device has been measured and no figure is projected for one. See Performance and the Android SDK page.
essence-2 lands on Maven Central — both families now have a public Android SDK (2026-09-03)
ai.bithuman:essence2-android:0.2.0 is published to Maven Central and resolves
anonymously, with no credential. With
ai.bithuman:expression2-android:0.3.0 (2026-09-02) and
ai.bithuman:sdk:2.3.6 (essence-1), the ai.bithuman group now lists three
artifacts and every model the scope ruling puts on the Android lane has a
coordinate that resolves.
implementation("ai.bithuman:essence2-android:0.2.0") // essence-2, minSdk 29, arm64-v8a
Ship-state, stated plainly. This artifact ships knowingly under the 2026-08-30 “base offering first” ruling, and two things are below bar:
- it fails the
PARITY_U8gate at 2 levels; - sustained throughput is 1.63x short of the accepted bar — a 1,000-second Hexagon run reads RTF 0.9959 / 20.08 fps against an accepted RTF 0.61 / 32.7 fps.
No Gradle project outside bitHuman has been compiled against it yet, and no
render through this artifact’s own API has been taken — what is established is
the coordinate, the bytes, the checksum, the declared minSdk and the native
payload. Updated the same day: the renderer graph it carries has since been
benchmarked on a Snapdragon 8 Elite handset — see the entry above. The Android SDK page carries the
measurements and both negative controls.
FFmpeg / LGPL. The AAR links FFmpeg 7.1 statically, so LGPL-2.1 §6(a)
applies, and the relink materials are published beside the AAR at a
repo1.maven.org URL baked into the shipped META-INF/NOTICE.txt. The kit has
15 entries — the object archive, the real link command, and FFmpeg’s complete
corresponding source. Every claim about it is checkable from the published
bytes: FFmpeg / LGPL — the Android relink offer.
expression2-android carries no FFmpeg and needs no such offer.
CLI 2.5.1 — macOS and Linux back on one version (2026-09-03)
cli-v2.5.1 publishes both aarch64-apple-darwin and
x86_64-unknown-linux-gnu. The 2.5.0 split — where macOS moved ahead and
Linux was stuck three releases back on cli-v2.4.2 — is closed, and the
BITHUMAN_VERSION=cli-v2.4.2 pin this site used to recommend on Linux should
be dropped. The unpinned universal installer is now correct on both.
pull --model <family> is in the Linux build too; the note saying it was
macOS-only is withdrawn.
Still not published, and never has been: x86_64-apple-darwin (Intel Mac)
and, since cli-v2.3.27, aarch64-unknown-linux-gnu (Linux ARM). On those two
targets install.sh resolves a download that 404s and exits 1. See
Downloads for the four-target probe.
bithuman render is unchanged and still limited: rc=0 for expression-2,
rc=69 for essence-2 — the shipped lib/libonnxruntime.so.1 is built at
VERS_1.20.1 while every lible_core.so requires VERS_1.26.0, so copying
a file in does not fix it — and rc=70 for essence-1. Details and the
controls: what the CLI actually does.
CLI 2.5.0 — bithuman pull --model, and the first signed macOS tarball (2026-09-02)
bithuman pull <CODE> --model <FAMILY> gives the CLI a door to something the
download endpoint has always had. Before it, bithuman pull <CODE> could only
hand you the agent’s birth model, so an agent created as one family and
later given another returned the first one silently, with nothing saying another
family existed.
--model <FAMILY>is forwarded verbatim; the endpoint owns the vocabulary and answers an unknown name with a400naming the accepted set.- The no-flag path now names the families it did not hand you, read off the
download response’s own headers — no second request.
--jsongainsother_modelsandmodel_source; a missing header leavesother_modelsabsent, never[]. --modelon a showcase slug is refused (exit 66) rather than silently ignored.bithuman list --mineandbithuman auth statusverify against the platform API, andservetells the brain when audio finishes playing, not only when it is cut off.
cli-v2.5.0 was also the first Developer ID signed and notarized macOS tarball.
It shipped macOS only; Linux caught up a day later in
cli-v2.5.1.
Expression 2 self-hosting on Linux is fail-open, by owner ruling (2026-09-02)
The rebuilt Linux engine that ships inside CLI 2.5.1
(engines/linux-x64-1.0.0.engine) has metering enforcement off. A render
with no credential proceeds, behind a ★ UNMETERED RENDER banner on
stderr; the meter is still running and still beats wherever a credential exists.
This was deliberate. The engine’s own source carries the ruling: shipping the
LGPL remediation fail-closed “would have switched billing on for every
existing self-hoster at the moment they upgraded, with no notice — a pricing
change riding in on a licence fix.” Both switches (ENFORCE_DEFAULT and its
twin METER_ENFORCE_DEFAULT) read False in the shipped bytes.
Scope: expression-2 only. The engine declares PRODUCT = "expression-2",
and essence-2’s bithuman.tessera_offline stays fail-closed — no
credential there still raises MeteringNotArmedError and produces no frames.
Enforcement is expected to return; treat unmetered rendering as a grace period,
not a price.
Essence 2’s head upsample is rebuilt — a CPU-tier speedup, same picture (2026-09-02)
Corrected 2026-09-02. The first version of this entry said the Apple tier serves the previous head upsample. It does not: that tier’s CoreML build already expressed the step the rewritten way. The Apple row below is the measured replacement.
Essence 2’s head-upsampling step has been rebuilt. It is an internal graph
change: the API, the session contract, the ?model= tier slugs and the price
are unchanged, there is nothing to opt into, and the picture is the same — the
new and previous builds agree to 167.85 dB PSNR on the same identity and the
same frames, below one step of an 8-bit pixel, and the change was reviewed side
by side on video before it was accepted.
It rolls out per identity, the way the 2026-07-27 renderer change did: an identity picks it up when its bundle is rebuilt, and serves the previous build until then. The first identity was served on 2026-09-02. One identity is not a fleet: most identities are still on the previous build.
The speedup is a CPU-tier speedup, and only a CPU-tier speedup. The step it replaces is 68.6% of the forward pass under the ONNX Runtime CPU execution provider and 0.48% of it on CUDA, so the gain does not transfer:
- CPU (Threadripper PRO 5955WX, ORT CPU provider, batch 24, 4 threads, sustained, no throttling): 3.00× on the model step; the full delivered path goes 7.54 → 23.87 fps. Still short of 25 fps — that tier remains an offline and last-resort tier.
- GPU (RTX 4090, ORT CUDA provider, batch 24, as the deployed worker is configured): 0.971×, about 3% slower, against a 0.083% noise floor. No gain, and none should be expected. The absolute cost there — 1075 fps before, 1044 fps after on that step — is far enough above a session’s 25 fps that the difference is not observable.
- Apple (M4 Max on the serving host, CoreML, GPU compute, batch 1, fp32, cooled and unthrottled): no gain — the step was already built this way there. The rebuilt graph is genuinely not shipped to that tier, but its CoreML converter has expressed this step as the same padded 5×5 depthwise convolution plus pixel shuffle since before the rollout began, so the change is not a change there. That step costs 6.06% of the forward pass on that tier (renderer model 1.834 ms/frame, 545 fps, 0.36% noise floor), which caps any further work on it at 1.06×. Batch 1, so not comparable with the batch-24 rows above.
- Browser-local: not measured.
See The renderer.
The “Apple Neural Engine” tier is renamed Apple, and a false performance claim is withdrawn (2026-09-02)
Essence 2’s cloud Apple tier was documented as the Apple Neural Engine tier, with a per-frame throughput figure attached and attributed to every operation in the graph running on the Neural Engine. That attribution was wrong, so the number went with it.
What is true. The tier runs on Apple Silicon Macs through CoreML, and every serving worker on those hosts binds the GPU compute unit — verified on the production hosts on 2026-09-02. The Neural Engine is not off-limits: on a minority of identities the renderer resolves to a half-precision graph the Neural Engine accepts and runs there. Measured head to head, it was about 2.2× slower than the Metal GPU and slightly further from the reference picture, so the GPU is not a fallback — it is the fastest and most faithful unit on that machine. Both “it runs on the Neural Engine” and “it can never touch the Neural Engine” are false; the page now says the measured thing.
No replacement number is published. The withdrawn figure was a model-in-isolation reading that a live session never sees, and the per-model, per-compute-unit protocol used for Essence 2’s CPU table has not been run for the Apple or GPU tiers. Picking one of the figures in circulation is what produced the error, so the page says so instead. See Serving tiers.
Expression 2 is the opposite case, and is now documented separately. Its
Apple members are exported at half precision and CoreML’s own per-operation
compute plan places 84–100% of their operations on the Neural Engine — none
on the GPU. The two models share the word “Apple” and the historical -ane
slug and nothing else. See
Serving tiers.
Nothing you can write changed. essence-2-ane, expression-2-ane, the
?model= force slugs, saved links and every API field keep working exactly as
before. Only the prose name of the tier changed, from “Apple Neural Engine” to
“Apple”.
Linux wheels restored for bithuman 2.10.0 (2026-09-02)
2.10.0 was published for macOS first and carried no Linux files for about a
day, so between 2026-09-01 and 2026-09-02 pip install bithuman on Linux
silently resolved to the previous release, 2.9.0. All ten Linux wheels
(cp310–cp314 × manylinux_2_28 x86_64 and aarch64) are on PyPI now, published
from the same measured build as the macOS wheels. If you installed in that
window, run pip install -U bithuman and check bithuman.__version__. See
Downloads.
August 2026
409 MODEL_NOT_GENERATED now tells you how to fix it (2026-08-17)
The model gate used to state the problem and stop — "agent <code>'s expression-1 model hasn't been generated yet" — which read as “this agent
can’t do that model”. It can. Every 409 now names the remedy: the exact
model-add call and its cost when
the agent qualifies, or the missing asset when it doesn’t. expression-1 is the
clearest case and got its own wording (“isn’t enabled on this agent yet”):
nothing is ever trained for it, so any agent with an image and a voice —
including an Essence 1 agent — enables it with one free, instant call and can
then render Expression 1 talking videos immediately. See
Using Expression 1 on an existing agent.
Same gate, same 409, same “before any charge” guarantee — only the message
changed.
Essence 2 self-hosted — offline CPU rendering ships in Python SDK 2.9.0 (2026-08-02)
The essence-2 model now self-hosts on your own CPU servers. Python SDK
2.9.0 (Linux x86_64 and aarch64, Python 3.10–3.14) adds
bithuman.tessera_offline — install the bithuman[tessera] extra and
render the downloaded <code>.lebundle.imx to frames or an mp4 entirely on
your hardware, no GPU required, teeth-refinement stage included. Measured
end-to-end: ~22–31 FPS on a 16-core desktop (the higher band when the
bundle carries the CPU acceleration member). The runtime ships together with
its metering: a valid BITHUMAN_API_SECRET is required, sessions bill at
the self-hosted rate (2 credits/min), and without a key the renderer is
fail-closed — zero frames. Live streaming from your own server still runs
through the cloud. Quickstart:
Self-hosted → Essence 2.
July 2026
Essence 2 — sharper mouth and teeth, and a much smaller model file (2026-07-27)
Essence 2 now renders each identity through a new unified renderer. The mouth interior — the teeth especially — is rendered sharply rather than being averaged out of the source frames, and it shows most on wide-open speech: measured against each identity’s own previous build, mouth-region fidelity improved roughly 2× to 4.7× across the launch gallery. That ratio is LPIPS — a learned perceptual image-distance metric — computed only inside the mouth-interior mask of the reference render, on each identity’s held-out frames, so it is a per-identity improvement factor and not a cross-identity score (the renderer). It was also checked frame by frame by eye, not only by metrics. Mouth motion is also re-centred and wider, so speech reads as more dynamic.
Two practical consequences:
- The downloadable model got about 5× smaller. A
<code>.lebundle.imxfromGET /v1/agent/{code}/model/downloadis now roughly 85–105 MB instead of several hundred. ReadContent-Lengthrather than hard-coding a size. - No serving-cost or pricing change. Measured warm and end to end, the new renderer costs nothing extra to serve. Rates are unchanged: 4 credits/min cloud, 2 self-hosted, 500 credits to create.
Nothing in the API, the session contract, or the tier slugs changed. New creations get the new renderer automatically; existing agents move over as they are retrained, so an older agent keeps serving its current build (and its larger model file) until then. Runs on all three runtimes — cloud GPU, Apple Silicon, and CPU.
Expression 2 — faster creation, same quality bar (2026-07-22)
Expression 2 agent creation now runs the adaptive-ladder recipe by default: it starts from a short, efficient training schedule and climbs to more training only when an identity needs it to pass the same quality checks. In head-to-head testing this reaches essentially the same quality as the previous full-length recipe at roughly 40% of the training time and cost. Creation now takes about 1 to 1.5 hours (a recent cold-start run measured ~1h40m; runs trend faster as the shared training pool stays warm). No API, pricing, or serving changes — new creations get the faster path automatically.
Run Expression 2 locally from the CLI (2026-07-16)
The bitHuman CLI now renders expression-2 avatars on your
own hardware. bithuman run with no arguments is a zero-config quickstart: it
fetches the free Wise Pup avatar and renders it live — on macOS (Apple
Silicon) via CoreML / Apple Neural Engine, and on Linux x86_64 via LiteRT;
Windows is coming. Each avatar is one self-contained
.imx file and the render engine ships inside the CLI,
so a fresh install runs its first avatar with no extra setup — the CLI downloads
only your platform’s slice (about 26 MB on macOS, 63 MB on Linux). See
Local rendering by platform.
Expression 2 — smaller, sharper serving model (2026-07-16)
Expression 2 now serves each identity through a more compact per-identity model — roughly 6× smaller and faster to run than the previous build — with sharper rendering of the mouth and teeth. The result is a crisper avatar at a lighter serving cost. The change is live across the gallery identities and is applied to new creations automatically. Serving surfaces, the platform contract (push audio in, drain video out), the APIs, and pricing are unchanged — existing agents get the improvement with no action needed.
Expression 2 — adaptive per-identity training (2026-07-15)
Expression 2 agent creation now runs an adaptive training recipe: every agent must pass the same quality checks as before, and an identity that needs more work automatically gets more training rather than a lower bar. In practice creation completes in about 1 to 1.5 hours — see Expression 2 for the updated expectations. The Expression 2 identities in the gallery have been refreshed with models trained under the new recipe, with the same quality checks enforced. No action is needed: existing agents, integrations, APIs, and pricing are unchanged.
Agent creation is image-only (2026-07-10)
The video creation input is removed for all models (essence-1,
expression-1, essence-2, expression-2):
- Provide a portrait
image(or let the prompt generate one) — bitHuman generates the identity video internally, always 10 seconds, authored so idle loops seam perfectly (first frame == last frame). User footage can’t guarantee that loop contract, which is why it’s no longer accepted. - Never send
videotoPOST /v1/agent/generate: as enforcement rolls out platform-wide, requests carrying it are rejected with400 VIDEO_INPUT_NOT_SUPPORTEDbefore anything is billed — never silently ignored. video_aspect_ratiois removed with the video input;durationis deprecated (accepted but ignored — the internally generated identity video is always 10 seconds).- Existing agents are unaffected, and
POST /v1/files/uploadstill accepts video files as assets — video just isn’t a creation input.
Essence 2 naming settled (2026-07-10)
essence-2is the standard tier name — the light-name retirement completed (the formeressence-2-lightwas consolidated intoessence-2on 2026-07-05): the standard photoreal model, optimized to run everywhere (GPU / Apple Silicon / CPU / WebGPU-WASM), and the default. See Essence 2. The premium tier of the family (previouslyessence-2-quality) became an internal model and is no longer offered publicly; see Naming & migration.- Rates unchanged.
essence-2stays 4 credits/min cloud, 0.5× when self-hosted; creation stays 500 credits.GET /v1/pricingadvertises the canonical names only inagent_generation.by_modelandtalking_video.rates; deprecated aliases are not advertised. - Docs moved. The model guide now lives at
/concepts/essence-2; the old URLs
(
/concepts/essence-2-light,/concepts/essence-2-quality) redirect.
Expression 2 creation price: 2000 credits (2026-07-10)
Creation pricing is now per engine:
expression-2creation (and model-add) costs 2000 credits — up from 500. Expression 2 is the fully generative engine; each per-identity train runs substantially more GPU time than an Essence 2 train, and the price now reflects that cost.- Essence 2 stays at 500 credits; v1 stays at 250.
autobills the routed model’s rate — 500 when your subject routes toessence-2(photorealistic person), 2000 when it routes toexpression-2(cartoon / animal / stylized character). The dashboard shows the range before you generate;GET /v1/pricingadvertisesautoat the 2000 ceiling so callers never see a number lower than the possible charge.
Essence 2 & Expression 2 — launch rollout begins; model pages refreshed (2026-07-10)
The second-generation models reach their announced launch date and the
rollout is underway. Creation access opens progressively (a v2 creation
ahead of your account’s access returns
503 MODEL_NOT_YET_AVAILABLE and bills nothing;
the dashboard’s v2 creation entries ship separately from the API). Alongside the
rollout, the model documentation gained the shipping characteristics:
essence-2— photorealistic people; animates real identity footage at its native resolution (full-HD 1080p identity video by default) at ~25 fps; serves GPU → Apple Silicon → CPU (corrected 2026-09-02: this entry originally said “fully on-device on Apple Silicon”; no on-device Essence 2 build has been published — the Apple tier is bitHuman’s own hardware), and a browser-local tier is rolling out (?render=local, WebGPU with WASM fallback) as per-identity web bundles publish.expression-2— stylized and universal characters; fully generative across the whole 416×720 scene at 20 fps from a single photo (no face detection or cropping anywhere in the pipeline), which is why any character morphology animates naturally; serves GPU → Apple Silicon → CPU; its on-device Apple engine shipped later, in Swift SDK 2.5.0 — engine only, with no model bundle published.- The family overview’s device matrix and creation guide were refreshed to match.
Plan concurrency, offline licensing preview, and one pricing page (2026-07-10)
Rounding out the launch — plan allowances and a documentation overhaul:
- Concurrent avatar sessions are now a plan allowance — Creator 3,
Pro 10, Business 50, Enterprise 200, Custom unlimited. Enforcement is
rolling out: once active, a session start beyond the allowance returns
403 CONCURRENCY_LIMIT_REACHED, and live sessions are never cut off mid-stream by the limit. See Session concurrency. - Offline licensing is coming soon — run avatars fully self-hosted with per-device, per-model signed credit bundles minted through your online account: Business $999/year prepacks 120,000 credits (Essence 2 + Expression 2); Enterprise $1,999/year prepacks 240,000 credits. Self-hosted minutes meter at half the cloud rate. Preview at Pricing → Offline licensing.
- Pricing is now the single home of every number — per-model serving rates (cloud and self-hosted), creation credits, talking-video rates, and the plan table live there; other pages link to it instead of repeating figures.
- Naming and migration history has one home — every alias, retired name, and response-name lag is consolidated at Models → Naming & migration.
- Every API operation ships a runnable example — all 33 operations in the interactive API reference now carry copy-paste curl samples with realistic bodies and next-step hints.
- Android documentation restored — the Kotlin / Android SDK page and the Android hello example are reachable again, and the voice reference URLs consolidated at Text to speech.
Multi-agent avatar rooms — audio binds to the launching agent (2026-07-09)
The cloud avatar now pins its audio to the agent that starts the
AvatarSession (via the LiveKit lk.publish_on_behalf attribute), fixing
wrong-agent audio binding in rooms with more than one agent participant. The
avatar previously bound to the first agent it saw, so with a facilitator +
persona in the same room it could latch onto the wrong agent — staying silent
for the persona and never returning playback_started/playback_finished.
Server-side fix; no SDK or plugin upgrade required. See
LiveKit → Multiple agents.
essence-2-light consolidated into essence-2; force-tier slugs (2026-07-05)
The Essence 2 request surface is now just essence-2 (plus the explicit
essence-2-quality reference tier):
- The
essence-2-lightname is retired. Create and render withmodel: "essence-2"— the light tier is what it serves. Requests namingessence-2-light(or the oldessence-2-light-aneslug) get a targeted400pointing atessence-2. Existing agents and saved links keep working (retired values route to theessence-2chain), andessence-2-lightremains the internal family name you’ll still see insupported_models,409messages, and model downloads. - Serving chains + force tiers. By default
essence-2andexpression-2sessions route down a serving chain (GPU → Apple Silicon → CPU) with automatic overflow. New force-tier slugs —essence-2-gpu/essence-2-ane/essence-2-cpuandexpression-2-gpu/expression-2-cpu/expression-2-ane— pin one tier for benchmarking/placement testing and never overflow. See tier pinning. - Talking videos:
POST /v1/video/generateacceptsessence-2(4 credits/min) in place of the retired name;essence-2-quality(8) andexpression-2(4) unchanged. - Where each model runs: the family overview gains a device/runtime matrix (cloud tiers, self-hosted, on-device Apple Silicon, browser-local status).
Android / Kotlin SDK docs restored (2026-07-04)
The Android SDK page and the Kotlin hello-avatar example are back. The on-device Essence runtime for Android — ai.bithuman:sdk:2.3.6, a self-contained arm64-v8a AAR on Maven Central — is unchanged and installable; only its documentation had been removed. It’s pinned at 2.3.6 (Essence, Engine ABI v7, Beta) and renders Essence .imx models fully on-device.
Pick-for-me creation, model adds & downloads (2026-07-02)
Named as of today: these two tiers were called Essence 2 Light and
Essence 2 Quality when this shipped; Light is essence-2 now and
Quality became an internal model that is no longer offered publicly — both
retired names and the migration are documented under
Naming & migration.
The model-release UX wave — one creation surface across all five model families, plus post-creation adds and artifact downloads:
model: "auto"— let the platform pick.POST /v1/agent/generatenow acceptsauto: an LLM classifies your input (the image if provided, else the prompt) and routes it — a photorealistic person →essence-2, a cartoon / animal / exotic creature →expression-2. It’s the default selection in the dashboard’s create flow; API callers send it explicitly (an omittedmodelkeeps the historicalessence-1default). Charges the routed model’s 500-credit rate.- The Essence 2 subject gate. Explicit
essence-2*creations require a photorealistic human subject — anything else is rejected with a clean422 MODEL_SUBJECT_MISMATCHbefore billing and before any agent row is created (autoroutes instead of rejecting). See the subject gate. - Per-model creation pricing. Creation is billed per model — 500 credits for the second generation (
essence-2,essence-2-quality,essence-2-light,expression-2,auto), 250 for v1 (essence-1,expression-1).GET /v1/pricingnow returns the per-model map (agent_generation.by_model) — the old flat field is gone. POST /v1/agent/{code}/models— add a model to an existing agent. No re-creation: addessence-1(250),essence-2(500),expression-2(500), orexpression-1(free, instant — the shared v1 engine drives the agent’s existing image + voice, nothing trained). Async adds poll viasupported_models; failures auto-refund; re-POSTing never double-charges.GET /v1/agent/{code}/model/download— download your generated model. A 302 to the artifact (?redirect=falsefor JSON):essence-1→.imx,essence-2-light→.lebundle.imx(licensed weights),essence-2-quality→.pkl,expression-2→.avatar(the Mac-runnable CoreML build). Per-family error matrix including the poll-able404 MODEL_ARTIFACT_NOT_READY.- The CLI recognizes every model family.
bithuman run/info/pullnow sniff any bitHuman artifact and answer honestly:essence-1.imxruns locally as always;.lebundle.imx/.pkl/.avatarare recognized with a clear handoff to where they run (launch matrix). New:bithuman pull <AGENT_CODE>downloads your own agent’s model through the endpoint above.
Official model guides + natural idle for the second generation (2026-07-02)
Named as of today: these two tiers were called Essence 2 Light and
Essence 2 Quality when this shipped; Light is essence-2 now and
Quality became an internal model that is no longer offered publicly — both
retired names and the migration are documented under
Naming & migration.
- Per-model official documentation. Each second-generation model now has a full product guide — what it is, how creation works (inputs, pipeline steps, realistic durations), serving tiers and
?model=pinning, idle behavior, pricing, and limits: Expression 2, Essence 2 — plus a new session behavior & troubleshooting guide covering connect latency (warm first line vs scale-from-zero overflow), idle vs speaking behavior, and the common errors. - Expression 2: real-footage idle on every creation. During silences the avatar now plays a looping clip derived from the identity itself — cropped from your source footage when available, or captured from the trained model’s rest pose for photo-only creations — instead of generated idle frames. Baked in automatically at creation; existing agents’ idle clips were regenerated.
- Forward-only looping. Idle and base-video loops now always play forward, wrapping from the last frame back to the first — footage never plays in reverse. Applies to
expression-2(all tiers, including on-device) andessence-2-light(idle and speech, all tiers). supported_models+ early model gate. Agent responses (status, get, list, and the embed-token response) now includesupported_models— the canonical model families the agent can be launched as right now.POST /v1/embed-tokens/requestaccepts an optionalmodelfield, validated up front; requestingexpression-2/essence-2-lightbefore the agent’s trained model exists returns a clean409 MODEL_NOT_GENERATED(“agent<code>’s<model>model hasn’t been generated yet”) — on talking video, before any charge. (Update, later on 2026-07-02:essence-2-quality— originally never gated here — is now gated on the agent’s source video, the footage its identity prepares from; see the model-release entry above.) A live?model=override to an ungenerated model now ends the session cleanly withavatar_error: "model_not_generated"instead of hanging.
Announced — Essence 2 & Expression 2 (launching July 10, 2026)
bitHuman’s two second-generation avatar models — essence-2 and expression-2 — are announced and launch July 10, 2026 on every surface (the REST API, the embed widget, the dashboard, and the SDKs). Until then, essence-1 and expression-1 are available today. See Essence 2 & Expression 2 for the full guide.
expression-2— the second-generation expression engine. Audio-driven, real-time avatar video from a single photo: agent creation trains a small per-identity model, then the engine synthesizes fully generated motion live. (Update 2026-07-02: per-model creation-time expectations are now documented — roughly 45 minutes forexpression-2; see the per-model guides.) Serves on three tiers — gpu, cpu, and Apple (the Apple tier — the slug is historical; serving tiers). 4 credits/min cloud · 2 credits/min self-hosted.essence-2-quality— the highest-fidelity tier of the Essence family: a heavy GPU renderer for close-up, hero-quality output on cloud GPUs. 8 credits/min cloud · 4 credits/min self-hosted.essence-2-light— the cost-effective tier: an efficient renderer that runs across gpu, cpu, and Apple — including fully on-device, where audio and video never leave your hardware. 4 credits/min cloud · 2 credits/min self-hosted.
All three are train-on-create via POST /v1/agent/generate (500 credits, one-time) and serve through the existing session flows unchanged. The v1 models (essence-1, expression-1) remain fully supported at 250 credits creation.
June 2026
Talking video generation — new API (2026-06-29)
- New endpoints:
POST /v1/video/generate+GET /v1/video/{job_id}. Render a finished talking-video mp4 of one of your agents from text or audio. With text input, the agent’s own voice speaks your script; with audio input, your hostedaudio_urldrives the render directly. The API is asynchronous — submit a job, then poll for the public mp4 URL, output duration, and credits charged. Launch engines:expression-2(4 credits/min) andessence-2-quality(8 credits/min), billed per minute of output rounded up; a failed render is automatically refunded. Limits: 120 seconds of output, 5000 characters of text. See Talking video generation and the Video API reference.
Agent generation — v2 model names accepted (2026-06-29)
POST /v1/agent/generatenow accepts the v2 model names. Themodelparameter takesessence-2-quality,expression-2, andessence-2-lightas supported generation targets (alongsideessence-1/expression-1). The v2 models launch July 10, 2026 (upcoming). Update 2026-06-30: the legacy aliases (elevate,embody,embody-gpu,essence-2-mobile) were retired ahead of GA — requests using them now return a400 VALIDATION_ERRORnaming the current model list. Share links are unaffected.
Model naming — versioned public taxonomy (2026-06-26)
- The avatar model families now have versioned public names. The
modelparameter on agent generation (and the viewer’s?model=selector) accepts the consolidated namesessence-1,essence-2-quality,essence-2-light,expression-1, andexpression-2. Essence 2 ships in two tiers — Quality (essence-2-quality, the high-fidelity cloud GPU renderer) and Light (essence-2-light, the efficient run-everywhere renderer). The older valuesessenceandexpressionmap toessence-1/expression-1; the pre-release codename values (elevate,embody,essence-2-mobile) were transitional aliases and have since been retired (see the 2026-06-30 note above — they now return a validation error naming the current model list). Share links are unaffected. Documentation, dashboards, and app labels now use the new family names.
Python SDK bithuman 2.3.10 (2026-06-23) — self-hosted streaming lag fix
- Streaming no longer degrades over a long turn. Self-hosted streaming now holds a steady frame rate for the full length of a turn (long utterances used to slow down as they grew), with byte-identical output. The audio stream also resets at the start of each turn so idle frames can’t shift lip-sync.
Python SDK bithuman 2.3.9 (2026-06-23) — barge-in / interrupt fix
- Interrupt (barge-in) no longer wedges the runtime. Interrupting the avatar mid-utterance previously froze it after the first barge-in (the interrupt path shared the terminal stop signal). 2.3.9 routes interrupts through a separate event, drains in-flight frames, and resumes on a fresh runtime — so a user can talk over the avatar repeatedly without it getting stuck.
- Recommended LiveKit stack:
bithuman2.3.9+ withlivekit-plugins-bithuman1.6.3 andlivekit-agents1.6.x (pluspillow).
Python SDK bithuman 2.3.8 (2026-06-16)
- Maintenance release on the 2.3 line (2.3.5–2.3.7 were not published).
Python SDK bithuman 2.3.4 (2026-06-12) — Linux CA auto-discovery
- Linux CA auto-discovery. The SDK now finds your distro’s CA bundle automatically on Linux — self-hosted auth (
AsyncBithuman.create()) works zero-config on Debian, Ubuntu, SUSE, and Alpine-glibc layouts. The/etc/pki/tls/certs/ca-bundle.crtsymlink workaround needed on ≤ 2.3.3 is obsolete. Thanks to the customer report that pinned down the Debian/UbuntuProblem with the SSL CA certfailure. - Env-var override preserved.
CURL_CA_BUNDLE/SSL_CERT_FILEtake precedence over auto-discovery when set — a stale or wrong value will still break auth, so unset them unless they point at a valid bundle. - macOS wheel tags. The 2.3.4 macOS wheels are tagged for macOS 26+ (arm64). On older macOS, pip reports
No matching distribution found— see the Python SDK page for options.
May 2026
2.3.0 (2026-05-28) — layered architecture + PyPI wheel split
- PyPI wheel split.
pip install bithumanis now the Python SDK library only (~5 MB) —from bithuman import AsyncBithumanstill works. The bitHuman CLI moved to a siblingbithuman-cliwheel, published beside the Homebrew formula and the universal installer — all three delivered the same Rust binary, which printedlibessence 1.19.1 ABI 7 / bithuman 2.3.0onbithuman --version. That wheel is no longer published; the CLI comes from Homebrew or the universal installer today (the CLI page). - CLI surface trimmed. The binary now exposes exactly six runtime subcommands:
run,render,info,pull,list,doctor(plusinitfor scaffolding a new project — seven in total). Legacy 1.x verbs (voice,text,avatar,stream,speak,action,generate,asr,tts,models pull|list,cleanup) were removed during the 2.x line and stay removed. - Wheel matrix. The Python library
bithumanships on PyPI for macOS arm64 and Linux x86_64 + aarch64 (manylinux). The CLI wheel was macOS Apple Silicon only and is no longer published at all — install the CLI via Homebrew or the universalinstall.sh/ tarball, on macOS and Linux alike. Python 3.10–3.14. - Repo layout. Public source lives in two repos:
bithuman-sdk-public(since archived; examples now live inhomebrew-bithuman/Examples) — docs source, runnable examples, and landing pages — andhomebrew-bithuman— the Homebrew tap, universalinstall.sh, and tarball release mirror. The engine and language SDKs ship as prebuilt, statically linked artifacts on PyPI and SwiftPM. BITHUMAN_BRAIN_*→BITHUMAN_AGENT_*env-var rename (carried through from Wave 5 of the 2.x line):BITHUMAN_AGENT_PORT,BITHUMAN_AGENT_PYTHON,BITHUMAN_AGENT_SCRIPT. The oldBITHUMAN_BRAIN_*names are still read with a deprecation warning.- No external API breaks. Python (
from bithuman import AsyncBithuman) and Swift (import Bithuman) public APIs are unchanged from 2.2.x. Migration for existingpip install bithuman && bithuman runusers was install-time only: install the CLI separately to keep thebithumancommand (the CLI page is the one writer for how). - Engine ABI bumps to
v7(libessence 1.19.1) — addsbe_runtime_tick_compose_from_mel(compose a tick directly from a mel feed). Additive on top of v6; old SDK builds keep working. (be_set_default_audio_encoderis an additive, ABI-unchanged entry point and did not bump the ABI.) - LiveKit integration. The upstream pin-relaxation PR (livekit/agents#5882) has since merged —
livekit-plugins-bithuman(1.6.3) now pinsbithuman<3,>=0.5.25, sopip install bithuman livekit-plugins-bithumanresolves cleanly. - Removed surfaces. The
bithuman.utilsandbithuman.audioPython modules are gone from the slim 2.3.0 wheel (helpers are inlined into the examples). Elevate was removed from the cloud model family but is retained as the on-device engine (vendoredlibelevate, used by AvatarUIKit and theexpression/iphonesample app) — it was not deleted from the platform.
Python SDK bithuman 2.2.2 (2026-05-25) — Linux CLI tarballs restored
- CI-only cleanup release; no API / runtime changes. Same Python wheel content as 2.2.1.
- Linux CLI tarballs (
bithuman-x86_64-unknown-linux-gnu.tar.gzandbithuman-aarch64-unknown-linux-gnu.tar.gz) ship on the GitHub Release again — they had been missing since 2.0.1 because of two container-build blockers, both now fixed inmain. - Pin
bithuman==2.2.2if you wantpip installAND the standalone Linux CLI binary from the same tag;==2.2.1is fine for wheel-only consumers.
Python SDK bithuman 2.2.1 (2026-05-25) — the on-device brain
Note 2.2.0 was skipped; 2.2.1 is the first published build of this release, with identical source content. Install 2.2.1.
- A
[local]extra on the CLI wheel added a fully on-device conversation brain tobithuman run, flipped on withBITHUMAN_LOCAL=1; no API key required, no outbound network. The wheel is gone and the brain is not: local mode names the packages to install today. - Stack:
whisper.cpp(STT) +llama.cpp(LLM, default Qwen 2.5 0.5B-Instruct Q4_K_M) + Supertonic 3 (TTS, 31 languages, voice M1 default) + Silero VAD. All in-process — no Ollama or other server. - All three backends have first-party iOS C++ cores, so the same
.gguf/.bin/.onnxmodel files are reusable when porting to mobile. - New plugins live in
livekit.plugins.bithuman.{WhisperSTT, LlamaCppLLM, SupertonicTTS}alongsideAvatarSession. The avatar-only install path is unchanged (heavy deps are lazy-imported). - Tuning via env vars:
BITHUMAN_LOCAL_WHISPER,BITHUMAN_LOCAL_LLM,BITHUMAN_LOCAL_LLM_FILE,BITHUMAN_LOCAL_VOICE,BITHUMAN_LOCAL_LANG,BITHUMAN_INSTRUCTIONS. See Python SDK. - Footprint: ~860 MB on disk (auto-downloaded from HuggingFace on first run), ~1.5 GB RAM, ~717 ms warm load, ~1.4 s warm end-to-end on Apple Silicon.
- Cloud path (
BITHUMAN_LOCALunset,OPENAI_API_KEYset) is byte-for-byte unchanged.
Python SDK bithuman 2.1.0 (2026-05-24) — figure → avatar
- Retired legacy “figure” terminology. CLI flag
--figures-rootis now--avatars-root(old name kept as a deprecated alias). Default cache moved from~/.cache/bithuman/figuresto~/.cache/bithuman/avatars. - No runtime behavior change; alignment with the public-facing “avatar” product term.
Python SDK bithuman 2.0.2 (2026-05-24) — graceful drain
bithuman runnow cancels active sessions and waits up to 2 s for libessence/HDF5 teardown before the process unwinds. Eliminates theH5F.c: decrementing file ID failed+ exit 134 SIGABRT on Ctrl-C / LaunchDaemon stop. Required for production-style supervisors.
Python SDK bithuman 2.0.1 (2026-05-24)
AsyncBithuman.cleanupis nowasync—await b.cleanup()works (was raisingTypeErrorand segfaulting at interpreter shutdown).- CLI error message polish:
bithuman pull <bad-slug>andbithuman renderno longer reference renamed subcommands. essence-render --helpshows the correct prog name (wasbithuman).
Python SDK bithuman 2.0.0 (2026-05-22) — bundled-CLI release
pip install bithumannow ships abithumanconsole-script that runs the full talk-to-your-avatar stack (Rust CLI + embedded livekit-server + the agent brain (STT/LLM/TTS) + browser UI). One install, one command, one URL — same Rust binary as the Homebrew CLI.- The runtime library API (
import bithuman,AsyncBithuman,from bithuman import Avatar) is unchanged — existing library consumers keep working. - The legacy 1.x Python CLI is preserved as the
essence-renderconsole-script. - Wheels: macOS arm64, Linux x86_64, Linux aarch64. Python 3.10+.
- Quickstart:
pip install bithuman && bithuman run— see the quickstart for the full flow.
v1.18.5 (2026-05-18)
- Unified
bithuman: onepip install bithuman= full prior1.11.3API + native engine, 100% backward-compatible (==1.11.3code runs unchanged). - Native engine: far faster cold load + lower memory than pure-Python, exact output parity. Loads fresh console
.imxTAR exports natively. - Python 3.9–3.14 (Linux x86_64/ARM64, macOS Apple Silicon). Pin
>=1.18.5(1.18.0–1.18.4 predate the unification; Windows / macOS-Intel stay on==1.11.3).
v1.17.x (2026-05-14)
bithuman avatar --openai— workstation Realtime, browser-rendered avatar.voice/textauto-pick cloud vs--local; explicit flags override.- Interactive TUI for
voice(mic/bot meters + transcript);BITHUMAN_NO_TUI=1opts out. - Flutter plugin renamed
bithuman_avatar→bithuman(one Dart codebase, mac/iOS). - Canonical OpenAI Realtime path is now the Rust CLI’s
--openaimode.
v1.16.0 (2026-05-14)
- Streaming API on Swift (
pushAudio/frames()/resetStream()). Flat per-tick cost on long sessions. - Default Realtime model:
gpt-realtime-mini.
v1.12.0 (2026-05-12)
- First unified release: Python, Swift, CLI from one source, identical output.
- Linux + Windows Python wheels (no WSL).
April 2026
- Chat Widget v5 — text/voice/video in one floating widget; themes, FAB styles, JS API (
open/close/setTheme/destroy). - FAQ KB — search always runs; removed dedup that dropped valid results.
- Voice — Siri-style animation; multilingual TTS (+11 languages including Thai, Chinese, and Arabic).
- Streaming — instant text to UI without waiting for audio sync.
March 2026
- Platform UI — sidebar (Explore / Library / Billing / Developer); Explore replaces Community, Library replaces My Agents; credit balance in top nav.
- Docs — screenshots + navigation refreshed for the new UI.
February 2026
- Expression Avatar v2 — 24% faster pipeline; no concurrent-session artifacts.
- Self-hosted GPU container — up to 8 sessions/GPU; ~50 s cold / 4–6 s warm; ~5 GB weights auto-cached.
- Examples overhaul — fixed Compose
env_file; standardized.env.example; addedAGENTS.md,llms.txt, OpenAPI spec. - REST API —
/v1/agent/{code}/speak,/v1/agent/{code}/add-context; consistent error codes. - SDK —
livekit-plugins-bithumanExpression support;bithuman.AvatarSessionunified cloud/CPU/GPU; animal mode for Essence.
January 2026
- Essence Avatar — CPU-only
.imxrendering, 25 FPS, Linux / macOS / Windows. - Platform API — agent generation, CRUD, file upload, dynamics/gestures.
- Integrations — LiveKit cloud plugin, iframe embed (JWT), webhooks, Flutter example.
Note Feature requests and bugs: GitHub and Discord. See the full community guide.