Cover image

Stuck on Repeat: Why LLMs Reinforce Their Own Bad Ideas

A mechanism-first reading of Martingale Score, a new unsupervised way to detect when LLM reasoning becomes prior-protecting rather than truth-seeking.

December 3, 2025 · 16 min · Zelina
Cover image

Blunders, Patterns, and Predictability: What n‑Gram Models Teach Us About Human Chess

A mechanism-first look at how skill-specific n-gram models turn chess move prediction from optimal play into human behavior modeling.

December 2, 2025 · 16 min · Zelina
Cover image

Checkmating the Hype: What LLM CHESS Reveals About 'Reasoning Models'

A mechanism-first reading of LLM Chess, showing why interactive benchmarks expose failures that static reasoning tests often miss.

December 2, 2025 · 17 min · Zelina
Cover image

From Building Blocks to Breakthroughs: Why RL Finally Teaches Models to Think

A mechanism-first reading of why reinforcement learning helps models compose memory and context only after supervised training has built the right atomic skills.

December 2, 2025 · 18 min · Zelina
Cover image

Ground and Pound: How Iterative Reasoning Quietly Redefines GUI Grounding

Chain-of-Ground shows that GUI grounding can improve not only by training larger models, but by forcing multimodal models to revisit their own visual hypotheses.

December 2, 2025 · 17 min · Zelina
Cover image

Roots of Understanding: When Transformers Try to Learn the Language of Numbers

A mechanism-first analysis of how a GPT-2-style transformer partially learns arithmetic structure from rooted-tree Dyck words—and why that is a benchmark lesson, not a factoring breakthrough.

December 2, 2025 · 15 min · Zelina
Cover image

Rules of Attraction: How LLMs Learn to Judge Better Than We Do

A mechanism-first reading of learned-rule-augmented LLM evaluators, and why the next AI judge may need better rubrics before bigger brains.

December 2, 2025 · 15 min · Zelina
Cover image

Short Paths, Sharp Minds: Why Knowledge Graph Distance Feels Like Cognitive Gravity

A mechanism-first reading of how graph distance can act as a surprise signal for knowledge-graph reasoning, and why the idea is useful before it is proven.

December 2, 2025 · 13 min · Zelina
Cover image

Eight Arms, One Mind: How OctoMed Turns Data Recipes into Medical Reasoning Power

OctoMed shows that medical reasoning gains may come less from bigger architectures and more from carefully mixed, trace-rich supervised fine-tuning data.

December 1, 2025 · 18 min · Zelina
Cover image

Forecasting the Forecasters: How Hierarchical LLM Meteorologists Rewrite Weather Reasoning

A mechanism-first reading of Hierarchical AI-Meteorologist, an LLM-agent system that turns forecast tables into multi-scale, explainable weather reports.

December 1, 2025 · 16 min · Zelina