Cover image

When Open Artifacts Still Hide the Workflow

A Central Kurdish TTS audit shows why downloadable weights, data, and evaluation code are not enough to establish reproducibility, measurement validity, linguistic scope, or reuse rights.

September 30, 2026 · 7 min · Zelina
Cover image

Before the Model Speaks, Check Whether the Audio Does

SURE-Voice shows why voice-agent reliability may depend as much on deciding whether audio should reach the model as on what the model generates afterward.

September 29, 2026 · 6 min · Zelina
Cover image

Don’t Make the LLM Serialize the Whole Workflow

A workflow-generation study suggests that when LLMs understand the requested actions but struggle to serialize dense graph structure, deterministic compilation can improve reliability without immediately requiring a stronger model.

September 29, 2026 · 8 min · Zelina
Cover image

Restore the Image, Preserve the Poisoning Signal

Imperfect Restoration Poisoning shows that data protection can recover image quality without fully restoring learnability, changing how teams should think about the trade-off between perceptual usability and resistance to model training.

September 29, 2026 · 7 min · Zelina
Cover image

Reward the Right Thing: GUI Agents Need Better Success Criteria, Not Just Better Judges

AdaptRubric shows that GUI reward quality depends not only on what a verifier can see, but on whether the system has defined the right task-specific success criteria before asking for a judgment.

September 29, 2026 · 7 min · Zelina
Cover image

Teach the Primitive, Not the Picture

SpatialBlock suggests that deliberately simplified synthetic tasks can transfer better than more realistic spatial labels when they isolate reusable spatial operations.

September 29, 2026 · 7 min · Zelina
Cover image

The Grader Is Part of the System: What RLVR Verifier Audits Reveal

A category-level audit shows why mathematical reward verifiers should be governed as configurable evaluation infrastructure rather than summarized by one accuracy number.

September 29, 2026 · 7 min · Zelina
Cover image

The Graph Isn’t the Verifier: What LCoT-GV Actually Learns From Long Reasoning Chains

LCoT-GV shows that graph-based reasoning verification depends far more on semantic step content than on structural metadata alone—and that verifier performance remains sharply domain-dependent.

September 29, 2026 · 7 min · Zelina
Cover image

Before the Word Arrives: How LLMs Use Sound to Choose a or an

A mechanistic study shows that LLMs can use a causal phonological feature to choose allomorphs before the conditioning word is generated—and that direct questioning may miss the mechanism.

September 28, 2026 · 7 min · Zelina
Cover image

One Forecast, Many Explanations: Why Time-Series Attribution Needs a Horizon Axis

A new time-series explanation framework shows that different forecast steps often rely on different parts of history, making one importance map an incomplete account of a multi-step forecast.

September 28, 2026 · 7 min · Zelina