TL;DR for operators
In Can Agents Design Libraries for Agents?1, the strongest evaluated agent-designed library condition scores 48.9, compared with 46.6 for production libraries and 34.4 with no library. That result is less a story about agents implementing more functionality than about whether a reusable abstraction actually removes work from the programs written later.
Across the main conditions, test-pass rates sit in a relatively narrow band of roughly 84%–87%. Simplicity separates them much more sharply. In an audit of 810 partially passing cases with excess code, Rigidity plus Verbosity account for 64% of primary classifications, while missing Coverage accounts for about 14%. The library often contains relevant capability; the interface simply makes that capability awkward to reuse.
The operational lesson is to test shared agent-facing libraries through downstream consumers. Unit tests remain necessary, but an internal SDK intended for coding agents should also be evaluated on how much task code it eliminates, whether agents discover its higher-level operations, and where its interface forces glue code or reconstruction outside the library.
A reusable library should remove downstream work
Suppose a company builds one internal SDK so dozens of coding agents do not repeatedly reimplement authentication, data transformations, workflow orchestration, or domain rules.
The natural acceptance test is whether the SDK works: does it install, expose the expected capabilities, and pass its own tests?
That is necessary, but it does not answer the operational question. A library can implement every required capability and still leave downstream agents writing large amounts of custom code because its abstractions are difficult to discover, too rigid for the task, or verbose to compose.
LibraryDesignBench is designed around that distinction. A designer agent first creates a library from a capability specification without being given a fixed interface. Fresh downstream agents then solve programming problems using that library. The benchmark covers 15 library-design tasks, 242 expert-validated downstream problems, and four languages: Python, TypeScript, Rust, and Haskell.
The object being evaluated is therefore not just the library artifact. It is the work that later agents can accomplish with it.
The score rewards task success and code elimination
LibraryDesignBench combines downstream correctness with simplicity. Correctness is squared before being multiplied by a simplicity measure, so a short but substantially incorrect program cannot earn a strong score merely by being small.
Simplicity is measured relative to expert reference programs using four static metrics: cyclomatic complexity, cognitive complexity, Halstead volume, and source lines of code. Ratios are capped at one, so being shorter than the reference does not generate bonus credit.
That construction explains why the headline comparison needs care.
| Library condition | Score | Test pass | Simplicity |
|---|---|---|---|
| Opus 5.5 designer | 48.9 | 86.6% | 64.5 |
| Production library | 46.6 | 85.4% | 61.5 |
| No library | 34.4 | 86.4% | 46.3 |
The no-library condition actually has a test-pass rate comparable with the other two. Its much lower simplicity score shows what the reusable abstraction is buying: agents can often solve the task either way, but without an effective library they must reconstruct more of the domain logic themselves.
The 48.9 versus 46.6 comparison should therefore not be read as evidence that agent-generated libraries are generally superior to mature human libraries. It says that, under this benchmark’s particular tasks, implementers, prompts, budgets, references, and scoring rule, the strongest evaluated designer produced abstractions that supported slightly better measured downstream performance.
Most audited excess code comes from interface friction
A straightforward explanation for poor library reuse would be missing functionality. If the needed capability is absent, downstream agents have little choice but to implement it themselves.
The paper’s failure audit points elsewhere more often.
Among 810 partially passing cells examined for excess-code failures, Rigidity and Verbosity together account for 64% of primary classifications. Missing Coverage accounts for about 14%.
These categories capture a different engineering problem. A library may expose the necessary capability but package it in operations that are too monolithic, impose an unsuitable shape, encode policy that cannot easily be changed, or require several low-level calls where a task-level operation would have absorbed the work.
For an organization building agent-facing infrastructure, this changes what an API review should ask. Feature coverage answers, “Can the library do this?” The more discriminating question is, “How much task-specific reconstruction remains after an agent uses it?”
The 64% figure should stay within its evidentiary boundary: it describes primary labels in the sampled partially passing audit, not the distribution of every failure across the full benchmark.
Even a good library fails when the agent never really looks at it
Library quality is only half of the system.
LibraryUseBench fixes the library side and examines how downstream agents interact with production libraries. Under a minimal prompt, 24% of GPT-5.6 Luna trials never open the library, while 36% of symbols used by the expert reference solutions are never seen.
More prescriptive instructions increase inspection and use. The experiments also separate two levers that are easy to conflate.
Higher reasoning effort mainly improves correctness and search depth. Stronger prompt prescription more directly improves library exploitation and solution simplicity. At high reasoning effort, moving from the minimal prompt to the high-prescription default raises simplicity from 48.7 to 60.4, while pass rates remain in the mid-80s.
This means an API can be reasonably designed and still produce poor reuse if the consumer agent begins coding before examining what already exists.
For agent platforms, documentation access and library-search behavior belong inside the interface contract. A reusable component that is technically available but routinely undiscovered has limited operational value.
Agent-facing APIs may not need to imitate human APIs
The paper also tests a bundled design intervention: consumer-first API sketches, runnable examples, and subagent testing.
Applied to GPT-6 Astra in the Codex harness, that guidance raises the benchmark score from 44.1 to 46.4 and improves three of four languages, mainly through shorter downstream solutions. At the same time, exported-name overlap with corresponding production libraries falls from 19.4% to 13.4%.
That combination is informative. Better downstream performance did not require making the generated API more lexically similar to the existing production interface.
It does not establish a general theory of “agent-native APIs.” The intervention bundles several practices, so the benchmark cannot identify whether API sketching, examples, subagent testing, or their interaction caused the gain. But it weakens the assumption that the safest path for an agent-facing interface is always to reproduce conventions optimized for human developers.
Add consumer-agent trials to SDK acceptance
The paper directly shows benchmark effects. Translating those results into an engineering workflow requires an additional inference.
For an internal SDK intended primarily for coding agents, Cognaptus would treat downstream consumer trials as a second acceptance layer after conventional library testing:
| Acceptance question | Suggested operational check |
|---|---|
| Does the library work? | Installation, library tests, integration tests |
| Do agents discover it? | Documentation opens, searches, files inspected |
| Do they use the intended abstractions? | High-level operations invoked; unnecessary reimplementation tracked |
| Does it remove downstream work? | Compare task code size and complexity against no-library or prior-library baselines |
| Where does reuse break? | Classify failures as coverage, correctness, rigidity, verbosity, or deliverability |
| Is extra design investment worthwhile? | Compare one-time library-generation cost with repeated downstream implementation cost |
This is especially relevant when reusable components serve a large population of automated consumers. Spending more compute or engineering effort on the shared abstraction can be economical if that cost reduces repeated work across many later tasks.
The benchmark demonstrates the measurement logic, not the ROI threshold. Organizations would still need their own workload frequencies, agent costs, failure costs, and maintenance requirements to make that decision.
The benchmark measures one slice of library quality
Several boundaries matter before these results become an internal standard.
The score measures tested correctness and static downstream-program simplicity. It does not comprehensively evaluate library security, runtime efficiency, maintainability, human readability, or complete behavioral correctness. Expert reference programs provide a normalization target but are not guaranteed to be uniquely optimal.
The reported confidence intervals measure reruns of the fixed benchmark rather than generalization to all software-library domains. Scores also depend on the downstream consumer configuration—including model, harness, prompt, task set, and execution budget.
Finally, the failure taxonomy is produced by one model auditor per cell without human labels or an inter-rater agreement study. It is useful diagnostic evidence, but its category shares should not be treated as universally calibrated failure frequencies.
Evaluate the interface through its consumers
Reusable software for coding agents creates a different measurement problem from standalone code generation.
A library can be correct while still being expensive to consume. It can contain the right capability while forcing agents to rebuild it through glue code. And it can be well designed while downstream agents fail to inspect it deeply enough to benefit.
LibraryDesignBench makes those failures observable by moving evaluation one step downstream.
For organizations deploying coding agents, that is the durable idea: approve shared abstractions not only because they work, but because repeated consumers demonstrably write less code, reconstruct less domain logic, and find the capabilities intended for them.
Cognaptus: Automate the Present, Incubate the Future.
-
Gabriel Orlanski and Alex L. Zhang and Avi Trost and Vincent Sunn Chen and Frederic Sala and Aws Albarghouthi and Ludwig Schmidt (2026). Can Agents Design Libraries for Agents?. arXiv:2609.36730. https://arxiv.org/abs/2609.36730 ↩︎