Cover image

Gated Sparse Attention: Speed Without the Sink

A mechanism-first reading of Gated Sparse Attention, showing how sparsity, gating, and adaptive token selection jointly target long-context cost, attention sinks, and training instability.

January 24, 2026 · 17 min · Zelina
Cover image

Learning to Discover at Test Time: When Search Learns Back

A mechanism-first reading of TTT-Discover, where test-time search becomes test-time learning for verifiable discovery problems.

January 24, 2026 · 18 min · Zelina
Cover image

PyraTok: When Video Tokens Finally Learn to Speak Human

A mechanism-first reading of PyraTok, showing why language-aligned multi-scale video tokenization matters for generation, understanding, and enterprise video AI.

January 24, 2026 · 15 min · Zelina
Cover image

Training Models to Explain Themselves: Counterfactuals as a First-Class Objective

A mechanism-first reading of counterfactual training: why better recourse may require changing the model, not just improving the explanation generator.

January 24, 2026 · 16 min · Zelina
Cover image

Triage by Token: When Context Clues Quietly Override Clinical Judgment

How proxy-variable testing exposes a quiet failure mode in LLM-based emergency triage: models can change acuity judgments when non-clinical context enters the prompt.

January 24, 2026 · 13 min · Zelina
Cover image

When LLMs Get a Laptop: Why Sandboxes Might Be the Real AGI Benchmark

A mechanism-first reading of LLM-in-Sandbox, showing why giving models a minimal computer environment may matter more than adding another clever prompt.

January 24, 2026 · 16 min · Zelina
Cover image

When Models Guess the Verb by Looking at the Drawer

A case-first reading of RCORE shows why video models can still confuse actions when object priors overpower temporal evidence.

January 24, 2026 · 17 min · Zelina
Cover image

Affective Inertia: Teaching LLM Agents to Remember Who They Are

A mechanism-first reading of how explicit state dynamics can make LLM agents more temporally coherent, and why too much stability becomes its own failure mode.

January 23, 2026 · 15 min · Zelina
Cover image

Cosmos Policy: When Video Models Stop Watching and Start Acting

A mechanism-first reading of Cosmos Policy, showing how latent frame injection turns a video diffusion model into a robot policy, world model, and planner.

January 23, 2026 · 16 min · Zelina
Cover image

Learning the Fast Lane: When MILP Solvers Start Remembering Where the Answer Is

DeepBound shows how a neural node selector can help branch-and-bound solvers find strong feasible solutions earlier without replacing exact MILP machinery.

January 23, 2026 · 17 min · Zelina