Cover image

A Good Score Is Not Permission to Rewrite the Playbook

TL;DR for operators An agent that repeatedly uses a procedure and receives high task reward still has not shown that the procedure deserves permanent promotion into its reusable playbook. R² Flow separates three questions that agent systems often blur together: how much a skill participates in successful execution, whether choosing that skill actually improves outcomes relative to alternatives, and whether there is enough independent evidence to authorize a persistent library change. Across eight recursive phases, the paper reports mean held-out OOD score rising from 71.2 to 81.02, with 33 of 37 committed edits improving held-out verified score. By contrast, reward-derived edit labels can raise training reward while lowering that verified score. ...

October 2, 2026 · 9 min · Zelina
Cover image

Self‑Improvement Without Self‑Destruction: Keeping Recursive AI Aligned

AI agents do not need to wake up one morning and declare independence to become difficult to govern. A more boring path is enough: generate an answer, critique it, revise it, score the revision, repeat. Add a little memory, a little tool use, a little automated evaluation, and suddenly “self-improvement” is no longer science-fiction wallpaper. It is an engineering loop. ...

March 9, 2026 · 13 min · Zelina
Cover image

Bottleneck or Breakout? Modeling the Compute Barrier to AI's Intelligence Explosion

TL;DR for operators The practical question is not whether AI will “think itself into godhood by Tuesday”. Charming as that spreadsheet would be, this paper is doing something narrower and more useful. Whitfill and Wu ask whether a software-only intelligence explosion can survive a compute bottleneck: if AI systems become good enough to replace human AI researchers, can that extra cognitive labour keep improving AI without a matching increase in research compute?1 ...

August 3, 2025 · 16 min · Zelina