Cover image

The 99% Problem: When a Stroke Benchmark Looks Ready Before It Is

Near-perfect stroke-prediction accuracy can justify deeper validation, but not deployment, when the score depends on a heavily rebalanced single-dataset benchmark.

September 27, 2026 · 7 min · Zelina
Cover image

Yesterday’s Gold, Today’s Bias: Reusing Human Feedback After a Model Upgrade

StalePO shows how production translation teams may recover useful corrections from legacy post-edits without training an upgraded model to imitate an older system.

September 27, 2026 · 7 min · Zelina
Cover image

Audit the Crowd Before the Attack: Forecasting Multi-Agent Capture from Benign Logs

A multi-agent safety study shows that normal interaction logs can forecast later collective behavior more accurately than isolated-agent testing within a controlled LLM-agent protocol.

September 26, 2026 · 7 min · Zelina
Cover image

Before the Solver: Clarification Needs Its Own Readiness Gate

A new benchmark shows that optimization agents can recover more missing requirements by questioning users more systematically, but still struggle to know when the specification is actually ready to model.

September 26, 2026 · 7 min · Zelina
Cover image

Control the Caption by Training What to Omit

FoCUS shows that controllable generation improves when training rewards both requested content and the suppression of off-scope content.

September 26, 2026 · 7 min · Zelina
Cover image

Forgotten Until Asked Differently: Unlearning Needs an Adversarial Sign-Off

A unified LLM-unlearning benchmark shows why passing ordinary forgetting tests is not enough evidence that targeted information is practically inaccessible.

September 26, 2026 · 7 min · Zelina
Cover image

Grow While You Roll: Training the Draft Head Inside the RL Run

GrowMTP shows that speculative decoding can be learned and amortized inside an LLM reinforcement-learning run, but its experiments also show why acceptance rate is the wrong metric to govern the accelerator alone.

September 26, 2026 · 7 min · Zelina
Cover image

More Shots, More Coverage? Measure the Set, Not the Sample

Validated Task Coverage shows why the best single LLM response, or the most visibly diverse generator, may not produce the best finite set of useful candidates.

September 26, 2026 · 7 min · Zelina
Cover image

The Verifier Already Knows: Turn Pass/Fail Checks Into Training Credit

VICT shows how explicit terminal checks can become selective, auditable training credit for long-horizon LLM agents instead of remaining a single final reward.

September 26, 2026 · 7 min · Zelina
Cover image

Human Demonstrations Need a Relevance Filter Before VLA Post-Training

ReWeight shows why robotics teams should select and weight human demonstrations by behavioral compatibility rather than treating them as uniformly useful training data.

September 25, 2026 · 7 min · Zelina