Cover image

One Model, Two Routes: Why Audio-Omni Unifies Audio by Splitting Its Controls

TL;DR for operators A product that must understand an instruction, preserve a voice or reference sound, and stay synchronized with video is handling several different kinds of control. Audio-Omni’s strongest architectural evidence says those controls should not be forced through the same interface. The system combines a frozen multimodal language model with a trainable audio generator. High-level meaning and transcript information are supplied as flexible context; synchronization and acoustic-reference information are attached directly to the evolving audio representation. In the paper’s conditioning ablation, that allocation performs best across text-to-audio, video-to-audio, text-to-speech, and audio editing. ...

September 14, 2026 · 9 min · Zelina