Meta unveils Muse Realtime Avatar: expressive video avatars synced to Muse voice in under a second
On 23 September 2026 Meta introduced Muse Realtime Avatar, turning Muse Realtime Voice into interactive avatars conditioned on a photo or illustration. Streams run at 448×768 and about 870 ms latency; Muse is for ages 18+.
On 23 September 2026 at Meta Connect, Meta introduced Muse Realtime Avatar — a system that turns Muse Realtime Voice into expressive, interactive avatars conditioned on reference media such as a photo, an illustration, or even animals and objects. The company also announced Ray-Ban Meta Audio glasses and said Muse is coming to glasses; Muse remains available for ages 18 and older.
How it works
The avatar is driven by an audio-driven Diffusion Transformer that shares a speech-token stream with the voice model so lip motion and expression stay in sync. Meta says it streams a 448×768 portrait at 25 frames per second, with roughly 870 milliseconds of latency from the end of a user turn to the first byte of synced voice and video.
Distillation cuts compute sharply: a teacher model that used 120 evaluations per chunk is replaced by a student that needs only two unguided evaluations — a 60× reduction. On a single GB200 GPU, Meta reports 12 concurrent realtime video sessions, an 8× capacity gain versus a BF16 baseline.
Comparisons and safety
In live-call comparisons, human raters preferred Muse Realtime Avatar overall against Runway Characters and HeyGen LiveAvatar, according to Meta. For safety, Meta applies Meta Video Seal, an invisible watermark, with no added latency.
Why it matters
Realtime, reference-conditioned avatars tighten the link between conversational AI and on-screen presence. Meta’s published latency, capacity and watermark details give a concrete baseline for how such systems may ship — while the 18+ age gate and watermarking signal the company’s stated safety framing for Muse.
Also available in: Саха тылаРусскийEspañol中文Portuguêsالعربية