TL;DR for operators
A company that uses an expensive agent to diagnose robot failures during development has an attractive deployment option: capture what that agent learned, hand the procedures to a cheaper agent, and remove the expensive model from the routine control loop. The evidence here shows why that handoff needs its own validation cycle.
In the reported SimplerEnv Bridge experiments, the lighter Luna agent achieved 43.8% success without external guidance. Giving it the teacher’s initial playbook reduced success to 31.3%. After the playbook was repeatedly revised using Luna’s own execution feedback, success reached 66.7%. Recursive Harness Distillation (RHD) therefore treats transferred expertise as an external control artifact that must be adapted to the recipient, not as instructions that can simply be copied from a stronger model.1
The operating implication is specific: use expensive agent capability to discover and revise recovery procedures during development, but validate those procedures on trajectories produced by the lower-cost deployment agent. In the paper’s Bridge evaluation, Luna with the refined playbook had an estimated mean agent inference cost of $0.63 per episode, versus $3.43 for Astra on the same held-out test set. That cost comparison is attractive only because recipient-specific refinement preserved useful performance.
A stronger teacher can hand over worse behavior
The intuitive deployment design is straightforward. A strong agent watches a robot policy fail, identifies what went wrong, and records what should be done next time. A cheaper agent later follows those instructions.
The first handoff in this study did not work that way. Luna’s 43.8% Bridge success without a playbook fell to 31.3% when it received the initial teacher-authored version.
That result matters more than a simple failure of prompt wording. Once an agent intervenes in robot execution, it changes what happens next. A different instruction can change the robot policy’s proposed action; an action correction changes the next camera observation; that observation changes the recipient’s next decision. The recipient therefore encounters trajectories partly created by its own previous interventions.
Guidance learned from the teacher’s execution history can consequently be sensible in the teacher’s context and poorly matched to the recipient’s.
RHD is designed around that mismatch.
The playbook changes execution without changing model weights
The underlying robot controller is a vision-language-action policy: a model that converts visual observations and language instructions into robot actions. RHD leaves that policy frozen. It also leaves the teacher and recipient agent parameters unchanged.
Instead, learned execution knowledge lives in an external playbook. The playbook tells the recipient when and how to intervene around the frozen policy through three channels.
The first edits the language instruction presented to the robot policy. The second modifies feature attention during the policy’s inference process. The third edits the action sequence after the policy generates it but before execution.
This distinction matters operationally. The system is adapting behavior without another round of parameter training. The resulting knowledge is also represented as explicit intervention guidance rather than being absorbed invisibly into updated model weights.
But an external playbook still becomes part of the effective execution policy. Changing its rules changes the trajectories the robot generates. That is why evaluating the document independently of the recipient that uses it misses the central system interaction.
Recipient rollouts are the revision signal
RHD starts with intervention experience collected by the strong teacher, which produces an initial playbook. The lighter recipient then operates the frozen robot policy using that guidance.
Its execution trace records observations, intervention choices, robot-policy proposals, executed actions, and outcomes. The teacher reviews these recipient-generated traces and proposes revisions. Candidate playbooks are tested on shared development instances, and revisions are accepted when the recipient reaches the teacher’s empirical success target.
The methodological point is more important than the document-generation loop itself: the target is the recipient’s task success under the playbook, not similarity to the teacher’s decisions.
The initial-playbook failure provides direct evidence for this design choice. After recursive revision against Luna’s behavior, Bridge success moved from 31.3% with the initial playbook to 66.7% with the refined one. Relative to unguided Luna, that is a 22.9-percentage-point increase on the 48 held-out Bridge instances.
The experiments separate refinement from teacher capability
Several comparisons help identify what is doing the work.
| Test | Result | Interpretation |
|---|---|---|
| GR00T alone on Bridge | 41.7% | Frozen robot-policy baseline |
| Luna + GR00T, no playbook | 43.8% | Adding the recipient alone produces little aggregate gain |
| Luna + initial playbook | 31.3% | Teacher-authored guidance can initially hurt the recipient |
| Luna + refined playbook | 66.7% | Recipient-feedback refinement is the main Bridge result |
| Astra + same refined playbook | 79.2% | The stronger agent still executes the guidance more successfully |
| Luna-derived refined playbook used by Luna | 22.9% | A weaker teacher does not reproduce the main transfer result |
| Astra-derived refined playbook used by Luna | 66.7% | Teacher capability materially affects playbook quality |
The teacher comparison is an ablation, not evidence that any sufficiently long refinement process will eventually work. When Luna acts as both teacher and recipient, the refined playbook reaches only 22.9% held-out Bridge success. The main result therefore depends on expertise contributed by the stronger Astra teacher as well as adaptation to the recipient.
The intervention-site ablations provide another constraint. Removing instruction editing causes the largest measured performance drop among the three intervention channels. Attention and action interventions remain part of the tested harness, but instruction-level control contributed most in this configuration.
Physical trials show the pattern is not simulation-only
The physical evaluation uses a Franka Panda robot on Cube-to-Tray, Cube Stacking, and Button Pressing, with 25 trials per task.
The frozen task-fine-tuned $\pi_{0.5}$ policy succeeds in 37.3% of the 75 trials. Adding Luna without a playbook produces 0% overall success. With the refined adapted playbook, Luna plus the same frozen policy reaches 64.0%.
That 0% condition is especially informative. Merely inserting an agent with access to intervention tools does not establish useful orchestration. The external guidance is doing substantive work in the reported system.
At the same time, these results should not be read as evidence that RHD creates a general robot-control capability. The playbook is guidance for a particular policy, interface, recipient, and task setting.
The operating model moves expensive reasoning into development
What the paper shows: a strong teacher can accumulate intervention experience, encode it outside model weights, revise it using recipient execution, and hand the refined guidance to a cheaper agent. Under the study’s recorded-token pricing calculation, Luna averages $0.63 in agent inference cost per Bridge episode with the refined playbook, compared with $3.43 for Astra.
Cognaptus inference: for repeated robotic workflows, this supports an architecture in which high-capability inference is concentrated in development, exception analysis, and playbook revision rather than paid for on every routine episode. The artifact being operationalized is not just a prompt. It is a versioned control procedure tied to observed failure conditions, intervention choices, and outcome checks.
That also creates a governance opportunity. Because the adaptation resides outside model weights, teams can inspect which intervention rule fired, compare revisions, and trace a recovery procedure back to the execution evidence that motivated it.
The required control, however, is recipient-specific validation. A procedure should not move into production merely because the expert agent generated it successfully.
Where the evidence stops
The study covers four SimplerEnv Bridge tasks and three physical Franka Panda tasks, with specific frozen VLA checkpoints and one primary strong-teacher/light-recipient pairing. It reports controlled held-out comparisons and ablations, but no confidence intervals or statistical-significance tests for the success-rate differences.
The recursive acceptance process also has no guarantee of termination if the teacher cannot propose a playbook that reaches its empirical target. And the reported cost advantage depends on the model-specific API prices used in the study.
So the evidence supports a system-design proposition, not a universal replacement rule for strong agents: intervention expertise can sometimes be moved into inspectable external guidance and reused by a cheaper recipient, but successful transfer depends on both teacher capability and recipient-centered refinement.
The most consequential result is therefore the initial failure. The teacher already had useful intervention knowledge, yet its first distilled playbook made Luna worse. Only after the guidance was revised against Luna’s own behavior did the handoff become effective.
For operators building agent-controlled physical systems, that changes where validation belongs. Do not stop after testing whether the expert knows what to do. Test whether the deployment agent can use that knowledge successfully on the trajectories it creates.
Cognaptus: Automate the Present, Incubate the Future.
-
Seungyeon Kim and Junhoo Lee and Minkyu Kim and Baekseung Kim and Nojun Kwak (2026). Recursive Harness Distillation across Agents for Robot Manipulation. arXiv:2609.33378. https://arxiv.org/abs/2609.33378 ↩︎