Model Interpretability

Opening — Why this matters now Reasoning models are having a moment. Latent-space architectures promise to outgrow chain-of-thought without leaking tokens or ballooning costs. Benchmarks seem to agree. Some of these systems crack puzzles that leave large language models flat at zero. And yet, something feels off. This paper dissects a flagship example—the Hierarchical Reasoning Model (HRM)—and finds that its strongest results rest on a fragile foundation. The model often succeeds not by steadily reasoning, but by stumbling into the right answer and staying there. When it stumbles into the wrong one, it can stay there too. ...

Model Interpretability

Reasoning or Guessing? When Recursive Models Hit the Wrong Fixed Point

Thoughts, Exposed: Why Chain-of-Thought Monitoring Might Be AI Safety’s Best Fragile Hope