The 99% Problem: When a Stroke Benchmark Looks Ready Before It Is
Near-perfect stroke-prediction accuracy can justify deeper validation, but not deployment, when the score depends on a heavily rebalanced single-dataset benchmark.
Near-perfect stroke-prediction accuracy can justify deeper validation, but not deployment, when the score depends on a heavily rebalanced single-dataset benchmark.
StalePO shows how production translation teams may recover useful corrections from legacy post-edits without training an upgraded model to imitate an older system.
A multi-agent safety study shows that normal interaction logs can forecast later collective behavior more accurately than isolated-agent testing within a controlled LLM-agent protocol.
A new benchmark shows that optimization agents can recover more missing requirements by questioning users more systematically, but still struggle to know when the specification is actually ready to model.
FoCUS shows that controllable generation improves when training rewards both requested content and the suppression of off-scope content.
A unified LLM-unlearning benchmark shows why passing ordinary forgetting tests is not enough evidence that targeted information is practically inaccessible.
GrowMTP shows that speculative decoding can be learned and amortized inside an LLM reinforcement-learning run, but its experiments also show why acceptance rate is the wrong metric to govern the accelerator alone.
Validated Task Coverage shows why the best single LLM response, or the most visibly diverse generator, may not produce the best finite set of useful candidates.
VICT shows how explicit terminal checks can become selective, auditable training credit for long-horizon LLM agents instead of remaining a single final reward.
ReWeight shows why robotics teams should select and weight human demonstrations by behavioral compatibility rather than treating them as uniformly useful training data.