Cover image

DRIFT-BENCH: When Agents Stop Asking and Start Breaking

A business-focused reading of DRIFT-BENCH, showing why agent reliability depends less on asking more questions and more on knowing when clarification helps, when it harms, and when execution must stop.

February 3, 2026 · 17 min · Zelina
Cover image

Identity Crisis: How a Trivial Trick Teaches LLMs to Think Backwards

A mechanism-first reading of why identity-bridge data can weaken the reversal curse in autoregressive LLMs—and why the useful trick is more delicate than it first looks.

February 3, 2026 · 18 min · Zelina
Cover image

No More Bit-Length Anxiety: Policy Iteration Goes Strongly Polynomial

A mechanism-first reading of why robust policy iteration for $L_\infty$ robust MDPs is not merely convergent, but strongly polynomial under fixed discount.

February 3, 2026 · 16 min · Zelina
Cover image

RAudit: When Models Think Too Much and Still Get It Wrong

RAudit shows why longer reasoning, stronger judges, and harsher critique can reveal LLM failures—but can also amplify them.

February 3, 2026 · 17 min · Zelina
Cover image

Seeing Is Not Reasoning: Why Mental Imagery Still Breaks Multimodal AI

A mechanism-first reading of MentisOculi, and why explicit visual thoughts still fail to become reliable reasoning evidence for multimodal AI.

February 3, 2026 · 18 min · Zelina
Cover image

Small Models, Big Mouths: Why Game AI Doesn’t Need Giant Brains

A mechanism-first reading of DefameLM: why narrowly scoped small language models may be more practical than giant cloud LLMs for real-time game AI and some business automation loops.

February 3, 2026 · 17 min · Zelina
Cover image

Thinking in Panels: Why Comics Might Beat Video for Multimodal Reasoning

A business-focused reading of Thinking with Comics, a paper arguing that comic panels may offer a cheaper and more structured middle path between static images and video for multimodal reasoning.

February 3, 2026 · 17 min · Zelina
Cover image

ThinkSafe: Teaching Models to Refuse Without Forgetting How to Think

A mechanism-first reading of ThinkSafe, a self-generated safety-alignment method that restores refusal behavior in reasoning models without paying the usual teacher-distillation tax.

February 3, 2026 · 15 min · Zelina
Cover image

When Language Learns to Doubt Itself: Self-Contradiction as an Upgrade Path for Multimodal AI

Self-contradiction in multimodal models is not just a failure signal; it may be a cheap diagnostic for aligning generation with understanding.

February 3, 2026 · 17 min · Zelina
Cover image

When LLMs Meet Time: Why Time-Series Reasoning Is Still Hard

A close reading of TSAQA shows why turning time series into question-answering tasks helps evaluate LLMs—but does not magically give them temporal reasoning.

February 3, 2026 · 16 min · Zelina