TL;DR for operators

An agent that repeatedly uses a procedure and receives high task reward still has not shown that the procedure deserves permanent promotion into its reusable playbook.

R² Flow separates three questions that agent systems often blur together: how much a skill participates in successful execution, whether choosing that skill actually improves outcomes relative to alternatives, and whether there is enough independent evidence to authorize a persistent library change. Across eight recursive phases, the paper reports mean held-out OOD score rising from 71.2 to 81.02, with 33 of 37 committed edits improving held-out verified score. By contrast, reward-derived edit labels can raise training reward while lowering that verified score.

For operators building agents that can modify reusable procedures, the relevant design pattern is therefore not simply “learn from successful trajectories.” It is a governed update loop: pool equivalent workflow evidence, distinguish importance from benefit, collect independent verification, validate proposed edits against quality and operating-cost constraints, version the accepted changes, and retain rollback.

The evidence is substantial within the paper’s benchmark protocol, but it is not evidence of indefinite or open-world self-improvement. OOD means six held-out benchmarks, most evaluations use fixed 128-record draws, reliable execution-grounded verifiers are assumed to exist, and recursive behavior is observed for eight phases.

A good task score is not approval to change the system

Consider an enterprise agent that repeatedly calls the same procedure, completes its tasks successfully, and then proposes adding that procedure permanently to its skill library. Two signals appear to support the change: the procedure was used frequently, and the resulting tasks scored well.

Neither signal answers the decision that matters.

Frequent use may indicate that the procedure sits on many successful execution paths, but not that it caused those outcomes. High reward may show that the complete workflow worked, but not that a permanent edit based on one component will improve later tasks. Once an agent can modify reusable workflows, optimization and change authorization become different control problems.

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution1 is built around that separation. It combines a shared-state orchestration graph, flow-based policy learning, independent verification, and versioned skill-library updates into a recursive system in which accepted edits alter the environment learned in the next phase.

Shared states stop equivalent workflows from competing for evidence

Suppose one successful workflow performs independent steps A and B in that order, while another performs B and A. A history-tree representation treats those traces as different prefixes even when the resulting operational state is equivalent.

R² Flow instead canonicalizes histories when their difference is only the ordering of independent events and the required dependency, artifact, observable, action, successor, and reward conditions remain satisfied. The result is a shared-state directed acyclic graph rather than a pure history tree.

That representation matters because evidence from equivalent procedures can accumulate at the same state rather than being fragmented across permutations.

The ablation results support treating this as more than implementation housekeeping. Every one of the 108 reported component-ablation cells falls below the full system, and removing shared states produces the largest average degradation among the component removals. The paper therefore links state representation directly to the quality of recursive learning.

For enterprise workflows, the narrower inference is that agent telemetry should not automatically identify a strategy with its exact event sequence. If two traces differ only in inconsequential ordering, learning systems may benefit from representing the operational state they reach rather than counting them as unrelated procedures.

Usage and usefulness are different quantities

The second separation is more important for skill governance.

R² Flow assigns a flow share to a skill: roughly, the fraction of reward-matched skill-call flow passing through it. This is an importance measure. It tells the system how centrally a skill appears in successful execution.

But flow share is nonnegative. A frequently invoked skill can therefore receive substantial share even when invoking it is worse than the available alternatives.

The paper adds a separate signed utility estimate. It asks whether choosing a particular skill event improves expected tempered reward relative to other legal actions from the same state, while discounting estimates associated with poor local flow balance. Positive and negative values can therefore distinguish helpful from harmful use.

That distinction becomes concrete in the harmful-skill experiment. Signed utility allows the full system to prune 11 of 12 injected harmful skills, whereas share-only ranking retains substantially more of them.

The operational lesson is not that flow share is defective. It answers a different question. A production system may need both:

Signal Question it answers Governance use
Flow share How much successful execution passes through this skill? Identify operationally central procedures
Signed utility Does choosing this skill improve outcomes relative to alternatives? Identify candidates for retention, refinement, or removal

A high-frequency procedure should receive attention. It should not automatically receive approval.

Self-modification needs evidence outside the reward being optimized

Even correctly identifying useful skills does not solve the approval problem. If the same task reward both trains behavior and decides which permanent edits are accepted, errors in reward attribution can propagate into the next generation of the system.

R² Flow therefore gathers independent verifier outcomes for selected skill invocations. These can come from execution-grounded checks such as unit tests, schema validation, assertions, or execution status. Context-specific Beta posteriors accumulate this evidence, while credible bounds distinguish a poorly evidenced skill from one with evidence of genuine weakness.

This is the paper’s most relevant governance mechanism.

Across eight recursive phases, 33 of 37 committed edits improve the held-out verified score, corresponding to reported edit precision of 0.89. Mean OOD score rises from 71.2 to 81.02 over the same sequence. The comparison with reward-derived edit labels is especially informative: those labels can increase training reward while reducing held-out verified performance.

The paper is therefore not merely adding another evaluator. It is separating the signal used to optimize behavior from the signal used to authorize persistent modification.

For an enterprise agent, Cognaptus would extend that design principle beyond model scores. A workflow change that looks beneficial in the agent’s optimization objective could still require an independently specified acceptance test tied to the actual operational risk: correctness, compliance, cost, latency, or another measurable constraint.

The recursive loop behaves more like controlled change management

R² Flow does not edit the skill library continuously after every improvement signal.

Within a phase, the orchestration policy is trained using its Recursive-Residual Trajectory Balance objective. The system monitors variance in complete-trajectory balance residuals and proposes a phase transition when learning appears to plateau, with additional checks including slope, relative decrease, and skill entropy.

Only then does the outer loop consider library changes.

Verifier evidence and credible bounds determine which skills qualify for actions such as retain, refine, split, consolidate, generate, or prune. Proposed changes then face paired non-inferiority validation covering success, tempered reward, token cost, and latency. Accepted edits are versioned; failed changes can be rolled back. Flow and verifier state are retained across phases rather than discarded.

This makes the recursion easier to interpret operationally:

learn → diagnose → propose → verify → validate → commit → continue learning

The library produced by one phase becomes part of the system trained in the next. Recursive self-improvement is therefore not one optimizer repeatedly rewriting itself. It is a sequence of controlled state transitions in which permanent changes face a different evidentiary standard from ordinary policy updates.

What the experiments establish

The broader benchmark results are consistent with the mechanism tests. Under the paper’s shared executor, tool, retrieval, and inference-budget protocol, R² Flow leads every reported IID and OOD benchmark row. Relative to the strongest reported baselines, the authors highlight gains of 2.03 IID exact-match points and 5.57 IID accuracy points on the corresponding averages, plus 3.83 OOD exact-match points and 2.85 OOD accuracy points.

The training-objective comparison also matters, but it tests a narrower question. R²TB outperforms GRPO, detailed balance, trajectory balance, tempered trajectory balance, and SubTB in the reported comparisons, with OOD improvements of 1.69 to 6.17 points over the alternatives. That comes at 7.5% more GPU time per training step than GRPO.

Executor-transfer experiments provide a different kind of evidence. Keeping the trained Supervisor fixed, the system improves every tested frozen executor across IID and OOD evaluation, with reported average gains between 15.70 and 24.64 points. This supports the possibility that orchestration and verified reusable skills can add value around different executors without fine-tuning each underlying model.

It does not establish executor-independent portability in general. The transfer tests still operate inside the paper’s defined architecture and benchmark protocol.

The production boundary is verifier quality

The paper provides unusually broad internal evidence for a self-improving agent system: controlled ablations, objective comparisons, executor transfer, harmful-skill injection, recursive phase analysis, and released implementation artifacts.

The largest gap between that evidence and production use is the verifier.

Unit tests, schema checks, assertions, and execution status can provide relatively crisp independent evidence. Many organizational procedures have no equivalent oracle. A research agent may produce a plausible synthesis whose quality cannot be reduced to a deterministic check; a customer-service procedure may optimize short-term resolution while degrading longer-term trust; a compliance workflow may depend on incomplete or changing interpretations.

In those environments, verifier-gated evolution remains an attractive architecture, but the verifier itself becomes part of the governance problem.

The other boundaries are temporal and distributional. Most benchmark evaluations use fixed 128-record draws, OOD refers to six prespecified held-out datasets rather than unrestricted deployment shift, and recursive improvement is demonstrated for eight phases. The results do not establish that improvement will continue indefinitely or remain stable as the environment changes.

Treat self-improvement as controlled authorization

R² Flow’s most transferable contribution is not the idea that an agent can accumulate skills. Systems already do that in various forms. The more consequential idea is that reusable self-modification requires several signals that should not be collapsed into one.

Equivalent workflow histories can share evidence. Frequently used skills can be separated from beneficial skills. Task reward can train behavior without automatically authorizing permanent edits. Independent checks can accumulate evidence before changes are committed. Validation, versioning, and rollback can sit between a proposed improvement and the next operational phase.

For organizations considering agents that can rewrite their own procedures, that architecture reframes the central question. The issue is not simply whether the agent can improve itself. It is what evidence the system should require before an observed improvement is allowed to become part of its future behavior.

Cognaptus: Automate the Present, Incubate the Future.


  1. Mingda Zhang and Qiang Huang and Yanjin Li and Zijia Wang and Qika Lin and Xiaoying Tang and Tiesunlong Shen (2026). R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution. arXiv:2609.33867. https://arxiv.org/abs/2609.33867 ↩︎