TL;DR for operators
A capable agent can identify the right problem and still fail to finish the job. In the clinical simulation studied here, the stronger baseline framework averaged 4.39/5 on diagnosis, but only 2.94/5 on critical actions and 3.34/5 on timeliness.
Asclepius attacks that gap without changing the underlying model weights. It combines three forms of scaffolding: an operating manual refined from prior simulated shifts, an external library of clinical skills, and specialized subagents for triage, diagnosis, and treatment completion.
Across all ten patient batches, the full system improved critical actions by 0.73 points and timeliness by 0.45 points relative to the baseline, while diagnosis and disposition were statistically unchanged. More importantly for generalization, four batches were never shown during harness evolution. On those held-out cases, critical actions improved by 0.625 points, or about 22%, with $p=0.024$. The held-out gains in overall score and timeliness were positive but not statistically significant.
The result should not be read as evidence that self-evolving instructions alone solved the problem. Once skills and acting subagents were held fixed, replacing the original operating manual with the evolved one produced only a +0.12 overall gain across all batches ($p=0.364$) and +0.02 on held-out batches ($p=0.883$). The stronger claim is about the combined scaffold.
For high-stakes agent systems, this changes the evaluation question. Teams need to measure whether correct reasoning becomes complete, timely, correctly prioritized action under sustained workload—not merely whether the model can state the right answer.
The agent knew more than it executed
The paper Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents1 starts from a failure mode that is easy to miss in short agent benchmarks.
Its simulated emergency-department shifts last six hours. Twelve patients arrive over time. Tests return with delays, physiology evolves, resources are constrained, and the agent must keep several cases moving simultaneously.
Under those conditions, a strong Claude Code-based baseline does well on diagnosis: 4.39/5. Yet its critical-actions score is 2.94, and timeliness is 3.34.
That separation is the paper’s central evidence. The model can often determine what is clinically wrong while failing to reliably convert that judgment into all required actions at the right time.
Calling this an execution gap becomes meaningful only after seeing that mismatch. It is not simply another accuracy problem. It concerns what happens after a plausible answer has already been formed: remembering unfinished work, applying procedural knowledge, prioritizing deteriorating cases, and issuing the necessary tool actions.
For an operator choosing between a stronger model and a stronger surrounding system, those are different investment decisions.
Asclepius changes the operating scaffold, not the model weights
Asclepius keeps the underlying Claude Opus 4.6 backbone fixed and modifies the system around it.
The first component is a natural-language operating manual. Between simulated shifts, a proposer reviews previous manuals, patient-level scores, action records, and audit logs, then proposes revisions. Ten iterations are evaluated on six search batches, with version 7 selected by overall search-batch performance.
The second component is an external clinical skills library. Rather than expecting the model to reconstruct every high-stakes regimen from parametric memory during a crowded shift, condition-specific procedural knowledge can be retrieved when needed.
The third component is a set of isolated subagents for triage, diagnosis, and treatment completion. Their purpose is not to add a new medical knowledge source. They narrow context and divide recurrent decisions into smaller scopes. In the default acting configuration, they can also execute permitted tools directly.
These mechanisms address different points in the workflow: cross-shift behavioral correction, access to procedural knowledge, and per-turn specialization.
The component ablations support treating them as complementary. Skills alone and subagents alone produce relatively modest changes. The harness is the strongest individual component on pooled data, but the full system produces the largest overall improvement and the strongest simultaneous reduction in the paper’s three execution-failure counters.
The strongest generalization result is about critical actions
Across all ten batches, Asclepius raises the overall score from 3.80 to 4.10. Critical actions rise from 2.94 to 3.67, and timeliness from 3.34 to 3.79. Diagnosis changes from 4.39 to 4.38, while disposition moves from 4.52 to 4.54.
| Dimension | Baseline | Asclepius | Difference |
|---|---|---|---|
| Overall | 3.80 | 4.10 | +0.30 |
| Diagnosis | 4.39 | 4.38 | -0.02, not significant |
| Critical actions | 2.94 | 3.67 | +0.73 |
| Timeliness | 3.34 | 3.79 | +0.45 |
| Disposition | 4.52 | 4.54 | +0.03, not significant |
The pooled results are informative, but they combine cases used during harness search with cases that were not.
The four held-out batches therefore matter more for judging whether improvements survive beyond the trajectories that influenced manual evolution. Among those 48 patients, the critical-actions gain is +0.625 points, about 22%, with $p=0.024$.
The held-out timeliness gain is +0.312 and the overall gain +0.193, but neither is statistically significant. Diagnosis and disposition remain unchanged.
That narrower result is more defensible than saying the complete Asclepius performance improvement generalizes. The held-out evidence most clearly supports better completion of critical actions.
The evolved manual is not the whole explanation
The paper’s design invites a tempting interpretation: perhaps repeatedly rewriting the operating instructions is the main innovation.
Its own ablation evidence does not establish that.
When the skills library and acting subagents are held fixed, changing only from the original manual to the selected evolved manual produces +0.12 overall across all batches with $p=0.364$. On the held-out batches, the difference falls to +0.02 with $p=0.883$.
The harness-search trajectory is also non-monotonic. Version 7 is selected because it has the highest overall score on the search batches, but later revisions can lose ground on individual dimensions.
The paper therefore provides stronger evidence for scaffold composition than for autonomous prompt evolution as an independently effective intervention.
That distinction matters outside healthcare. A team observing improvement after adding memory rules, retrieved procedures, specialized agents, and new instructions simultaneously cannot safely attribute the gain to whichever component is easiest to describe.
Failure counters show where execution degrades
The paper supplements outcome scores with three trace-level counters.
For the baseline, the rate of “naked disposition”—disposing a patient without the expected supporting execution—rises from 16% early in the shift to 57% late. Under full Asclepius, it moves only from 10% to 12%.
Among correctly diagnosed patients, the proportion receiving zero treatment falls from 19% to 9%.
The timeliness difference between non-severe and severe patients falls from 1.13 points to 0.23.
These measures are not comprehensive definitions of instruction drift, treatment completeness, or equity. Their value is diagnostic: they show that deterioration under workload can occur along several execution pathways at once.
The full system’s strongest result is therefore not merely a higher aggregate score. It is that several distinct signs of long-horizon degradation shrink together.
Scoped tool execution may matter as much as specialization
One additional comparison tests how the subagents participate.
With advisory subagents, the specialist returns recommendations to the main agent, which must interpret and execute them. With acting subagents, the specialist has scoped access to the relevant tools.
The acting version scores 4.10 overall versus 4.00 for the advisory design, including a 0.26-point advantage in critical actions.
The paper attributes the difference to avoiding another translation step between recommendation and execution, including paraphrase and vocabulary-resolution failures.
This is a variant test rather than the paper’s central evidence, but it raises a practical systems question: if a specialist can make a well-bounded decision, routing every action back through a coordinator may reintroduce the failure that specialization was intended to remove.
What an operator can carry into production design
The paper directly shows comparative performance inside CES. The broader business implications require one additional inferential step.
For teams deploying agents into long-running workflows, Cognaptus would separate at least three evaluation layers:
- Reasoning quality: did the agent identify the correct state or decision?
- Execution quality: were required actions complete, correctly ordered, and timely?
- Workload resilience: does performance remain stable as tasks accumulate, compete, and change priority?
Asclepius also gives a concrete architecture for improving the second and third layers without fine-tuning model weights: externalize procedural knowledge, isolate recurring decision functions, and adapt operating instructions using trace evidence.
The paper does not establish that every workflow needs all three. It does show why evaluating them only through final-answer accuracy can miss the dominant failure mode.
There is also a cost. Full Asclepius uses about 1.21× the LLM calls and 1.52× the wall-clock time of the baseline per batch, before counting the separate development cost of outer-loop harness evolution. Better execution therefore arrives with additional inference and engineering overhead that needs to be justified by the cost of omissions and delays in the target process.
The evidence stops at the simulator boundary
The study provides reasonably strong comparative evidence within its experimental environment, but several boundaries constrain transfer.
CES uses structured patient cases derived from de-identified emergency-department records with simulated physiology. Outcomes are graded by LLM judges. Four additional judges from three model families reproduce the same direction of critical-action and timeliness improvements, which is a useful robustness check, but no expert-agreement study was conducted specifically on the Asclepius traces.
Each configuration is also run only once per batch. The statistical model accounts for patient and batch heterogeneity, not run-to-run stochastic variation.
Finally, the evaluation uses one backbone model, one default emergency-department configuration, and one language and guideline regime. The overall improvement also compresses from +0.36 on the six search batches to +0.19 on the four held-out batches, consistent with some adaptation to the search cases.
Production teams should therefore treat the architecture as a testable design hypothesis, not as validated evidence of clinical benefit.
Correct reasoning is only one checkpoint
Asclepius does not meaningfully improve diagnosis. That is precisely why the paper is useful.
Its strongest evidence concerns what happens after the model already appears capable: whether the surrounding system keeps the right work visible, supplies procedural knowledge when needed, delegates decisions cleanly, and converts recommendations into actions before timing matters.
For long-horizon agents, evaluating the answer is not enough. The unit of reliability is increasingly the workflow that follows it.
Cognaptus: Automate the Present, Incubate the Future.
-
Grace Chang Yuan and Xiaoman Zhang and Sung Eun Kim and Luyang Luo and Pranav Rajpurkar (2026). Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents. arXiv:2609.13543. https://arxiv.org/abs/2609.13543 ↩︎