TL;DR for operators
Kang and Zhang’s KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI1 starts from a practical question: after a clinical AI system generates work, how much of that work does the responsible clinician actually keep?
For documentation, the reported answer is striking. Across more than one million signed encounters, more than six contiguous months, and thirteen specialties, 97.99% of generated content was retained on a content-weighted basis. But that number is not 97.99% factual accuracy, and it does not establish that 97.99% of the clinician’s documentation effort disappeared.
The larger contribution is the measurement framework around the number. KnowBench requires deployment claims to disclose presentation and finalization rates, completeness, formatting treatment, alignment method, and encounter-level distributions. The current release withholds several of those companion statistics. For procurement and governance teams, that makes the reporting protocol at least as consequential as the headline percentage.
A signed note records what survived review, not how easy review was
A common clinical-AI workflow already produces an evaluation trace without creating a separate benchmark dataset. The system drafts a note, the clinician reviews and edits it, and the signed note records what remained after that review.
If most generated material survives unchanged, the system appears to have completed work that otherwise would have needed to be produced manually. If substantial portions are rewritten or deleted, more work has remained with the clinician.
KnowBench formalizes that intuition as Effort Reduction, or ER:
Here, $g$ is the generated artifact, $f$ is the clinician-finalized artifact, and $R(g,f)$ is the generated content retained after review.
For documentation, the paper compares generated and final notes using token-level longest-common-subsequence alignment after formatting normalization. Content edits reduce ER. Formatting-only edits are classified separately. Material added by the clinician does not reduce retention directly; instead, additions belong to a separate completeness question.
That last distinction is essential. A short draft could be retained almost perfectly while the clinician adds substantial missing content. Retention can therefore be very high without showing that the system produced most of the final artifact.
The 97.99% result is a large deployment measurement with a narrow meaning
The reported documentation deployment covers more than one million signed encounters over more than six contiguous months ending in the first half of 2026.
Aggregate content-weighted ER is 97.99%. A stricter variant that counts formatting edits against retention is reported at at least 95.7%, while formatting-only changes touched 2.3% of generated tokens.
The specialty results are also relatively compressed. Aggregate ER ranges from 96.8% to 98.9% across thirteen specialties, with nephrology at 98.9%, general surgery at 98.8%, and psychology and primary care at 96.8%.
These are descriptive deployment measurements, not causal estimates of time saved, burnout reduced, productivity increased, or patient outcomes improved. The specialty differences likewise do not establish that specialty characteristics caused the observed variation.
One design choice strengthens the interpretation of the deployment stream: every generated draft during the measurement window was presented to the clinician and delivered to the EHR. The system was therefore not reporting retention only on a hidden subset of high-confidence drafts.
That still leaves a different denominator question unresolved. The paper does not disclose what fraction of delivered drafts ultimately became signed notes.
KnowBench makes the missing numbers part of the benchmark
KnowBench does not define a credible ER claim as one percentage. Its six-item reporting protocol requires enough surrounding information to understand what entered the denominator and what the retention statistic leaves out.
| Reporting element | What it helps establish |
|---|---|
| Measurement window and encounter count | Scale and observation period |
| Presentation and finalization rates | Whether generated artifacts reached review and how many completed the workflow |
| Completeness | How much of the final artifact the system actually supplied |
| Formatting treatment | Whether cosmetic changes affect the headline metric |
| Alignment method | How generated and final artifacts are compared |
| Encounter-level distributions | Whether the aggregate hides a long tail of poor cases |
This is where the paper becomes more interesting than its headline result.
The reported deployment satisfies the presentation component: generated drafts were shown to clinicians and delivered to the EHR. It also specifies alignment and formatting treatment. But the release withholds the signed-note rate over delivered drafts, completeness, and encounter-level ER distributions.
That limits several interpretations.
Without completeness, 97.99% retention cannot distinguish a substantively complete draft from one in which nearly all generated material survives but clinicians add substantial clinical content. Without finalization rate, readers cannot tell how often delivered drafts were abandoned before signature. Without encounter-level distributions, an excellent aggregate can conceal a subset of encounters requiring much heavier correction.
In other words, KnowBench’s disclosure requirements identify exactly why ER should be read as a measurement package rather than a standalone score.
Retained work and safe work are different measurements
The paper is also explicit that clinician acceptance does not establish factual correctness.
A clinician can retain incorrect generated content. Review itself can be imperfect, particularly when automation bias affects attention or verification. Knowtex describes continuing internal factual-consistency audits, but quantitative results from those audits are not published in the paper.
The practical separation is therefore clean:
ER measures what generated work survives review. Safety evaluation measures whether that surviving work is clinically reliable.
The paper also distinguishes artifact retention from actual review burden. ER-text can be calculated from draft-to-final artifacts, but validating whether it represents real effort reduction requires measures closer to active review time and cognitive effort.
For a healthcare organization evaluating a documentation product, this means a high ER number can answer a useful question—how much generated text clinicians leave intact—while leaving the more financially consequential question of how much clinician time was actually released only partially answered.
The framework is broader than clinical notes, but the evidence is not yet
KnowBench extends the same construction beyond documentation by changing the unit of generated work.
Notes and summaries use tokens. Coding uses codes. Orders use fields. Decision support uses recommendations. The benchmark also specifies applications to EHR chart summaries and patient after-visit summaries.
This creates a potentially useful common vocabulary for organizations managing several clinical-AI products. A portfolio team could ask each system what proportion of its generated work survives accountable human review while retaining task-specific definitions of content and completion.
The current empirical evidence, however, is narrower than that framework. The large reported deployment result is for documentation. The paper does not report equivalent production measurements for coding, orders, summaries, or decision support, and those tasks introduce their own confounds. A billing code may be added for reasons unrelated to model failure, for example, while a recommendation may be dismissed because of review fatigue rather than substantive disagreement.
Cross-task comparability therefore depends on the companion measures, not merely on applying the same ER formula.
For buyers, the reporting request matters more than the threshold
A practical Cognaptus inference follows from the benchmark design.
A healthcare buyer does not need to decide whether 97%, 95%, or any other ER level is universally “good.” The more defensible procurement question is whether the vendor can supply the full measurement context required to interpret its number.
That means asking for retention together with completeness, finalization rate, encounter-level distributions, active review effort, and independent factual-consistency evidence. It also means checking whether the reported result belongs to the deployed system as a whole or is being attributed too narrowly to an underlying model.
The 97.99% result in this paper belongs to Knowtex’s proprietary model family operating inside a closed-loop architecture with clinician customization, feedback from edit patterns, and longitudinal reuse of previously attested content. It should not be read as an isolated property of a generic foundation model.
The longitudinal component introduces another unresolved boundary. Retention can include carried-forward material from prior attested notes. That continuity may remove real work, but it can also reflect copy-forward behavior or note bloat. The reported aggregate does not separate those mechanisms.
A benchmark for work removed needs evidence about work remaining
KnowBench’s strongest idea is not that clinical AI can achieve a 97.99% retention rate. It is that routine clinician review can become part of the evaluation infrastructure for deployed AI.
Draft-to-final comparison is attractive because it is scalable, tied to actual use, and based on the artifact for which a clinician ultimately takes responsibility. The benchmark then adds the necessary discipline: define the denominator, measure what was missing as well as what was retained, report the distribution rather than only the aggregate, and keep safety evaluation separate from acceptance.
The current release demonstrates the first half of that proposition at production scale. Its undisclosed companion statistics show why the second half cannot be optional.
For vendors, buyers, and governance teams, ER is therefore most informative when it becomes the start of the measurement conversation—not the number that ends it.
Cognaptus: Automate the Present, Incubate the Future.
-
Jocelyn Kang and Caroline Zhang (2026). KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI. arXiv:2609.15794. https://arxiv.org/abs/2609.15794 ↩︎