A Good Score Is Not Permission to Rewrite the Playbook
TL;DR for operators An agent that repeatedly uses a procedure and receives high task reward still has not shown that the procedure deserves permanent promotion into its reusable playbook. R² Flow separates three questions that agent systems often blur together: how much a skill participates in successful execution, whether choosing that skill actually improves outcomes relative to alternatives, and whether there is enough independent evidence to authorize a persistent library change. Across eight recursive phases, the paper reports mean held-out OOD score rising from 71.2 to 81.02, with 33 of 37 committed edits improving held-out verified score. By contrast, reward-derived edit labels can raise training reward while lowering that verified score. ...