When Models Guess the Verb by Looking at the Drawer
A case-first reading of RCORE shows why video models can still confuse actions when object priors overpower temporal evidence.
A case-first reading of RCORE shows why video models can still confuse actions when object priors overpower temporal evidence.
A mechanism-first reading of how explicit state dynamics can make LLM agents more temporally coherent, and why too much stability becomes its own failure mode.
A mechanism-first reading of Cosmos Policy, showing how latent frame injection turns a video diffusion model into a robot policy, world model, and planner.
DeepBound shows how a neural node selector can help branch-and-bound solvers find strong feasible solutions earlier without replacing exact MILP machinery.
A tournament-style prompt evaluation study shows why educational AI teams need evidence, not just elegant prompt wording.
A practical reading of why multimodal climate fact-checking needs evidence orchestration, not just a larger vision-language model with a browser attached.
A diagnostic study of RL-trained Lean provers shows that more inference samples can repeat the same failed strategy, while tactic-level structural hints recover proofs that random sampling misses.
A mechanism-first reading of why LLM unlearning can look successful at the output layer while membership traces remain detectable inside model representations.
A mechanism-first reading of how an agentic DISARM pipeline turns disinformation investigation from expert taxonomy work into auditable, semi-automated evidence production.
A mechanism-first reading of why multi-agent LLM systems can improve results without adding information: they factor constraints across agents and stabilize solutions that single update dynamics may not reach.