Cover image

OmniAvatar’s Metrics & Training: Under the Hood of Next-Gen Avatars

TL;DR for operators OmniAvatar is best read as a shift from “make the mouth move” to “make the person perform.” The paper introduces an audio-driven avatar video generation system that takes a reference image, an audio clip, and a text prompt, then generates facial and semi-body video with synchronised speech, adaptive body motion, and prompt-controlled scene elements.1 ...

June 24, 2025 · 16 min · Zelina
Cover image

Blind Trust, Fragile Brains: Why LoRA and Prompts Need a Confidence-Aware Backbone

TL;DR for operators LoRA and prompts are attractive because they make model adaptation feel almost too easy: add a few examples, attach a small adapter, nudge the model into a domain, and call it customised. The uncomfortable part is that adaptation changes not only what a model says, but how confidently it says it. A compliance assistant that becomes slightly more domain-specific but far more overconfident has not been improved. It has been promoted beyond its competence, a classic corporate move. ...

March 25, 2025 · 14 min · Zelina