A Good Score Is Not Permission to Rewrite the Playbook
R² Flow shows why self-improving agents need separate mechanisms for learning what works, judging what helps, and approving permanent workflow changes.
R² Flow shows why self-improving agents need separate mechanisms for learning what works, judging what helps, and approving permanent workflow changes.
A reasoning-error taxonomy becomes more operationally valuable when recurring failures are converted into specific inference-time checks rather than generic instructions to be careful.
A developmental study of language models shows why highly accurate internal probes should not be treated as evidence that the same representation can reliably steer behavior.
A new agent-study framework suggests that reusable preparation can improve frozen agents and reduce repeated test-time search—but only when the preparation strategy fits the environment.
USplit-VQA shows that moving multimodal computation off constrained clients can sharply reduce memory and communication, but the partition itself becomes a model-quality and security decision.
A symbolic legal decision-support framework shows how an unresolved compliance check can identify the exact information needed to continue reasoning.
Magenta shows why reliable agent verification needs a semantic check before formal execution verification—and targeted repair after failures.
ViS-CoT shows that better product-video extraction starts with better evidence, while staged reasoning becomes valuable only after that evidence is available.
MM-IFEval-Pro shows how deterministic multimodal instruction testing can improve compliance and expose image-embedded instruction conflicts without confusing those gains with broader model capability.
A new credit-assignment study shows why precise turn-level rewards can underperform when the verifier sees only a small fraction of the workflow that produced success.