Skip to content
Go back

Can Agents Design Libraries for Agents?

G. Orlanski, A. Zhang, A. Trost, V. Chen, F. Sala, A. Albarghouthi, and L. Schmidt

We are fast approaching a point where the primary users of software libraries are agents, not human engineers. If agents are the users, then a library should be judged by how little code agents need to write correct programs with it, not by how it reads to a human. That is why I am excited to release LibraryDesignBench, a two-phase benchmark that scores an agent-written library solely by how much it helps the future agents that use it.

Asking agents to write libraries for other agents

LibraryDesignBench asks agents to write effective libraries, and we intentionally give them design flexibility. We never prescribe signatures their library must expose. Nor do we ever guide them to specific abstractions or patterns they should use. We intentionally give them underspecified, vague, and ambiguous library specifications because frontier agents need to envision how future agents will use their library, what they will need, and what design will work best. LibraryDesignBench gives agents this vast freedom because the evaluation needs to serve as a test bed for understanding what patterns agents actually prefer, not what we think they will prefer.

Evaluating A Library Based on How Agents Use It

A library that implements something correctly does not immediately provide any value – it needs to be correct and make future code simpler. Thus, the only way to measure a library’s quality is to observe how much future agents benefit from its design decisions. LibraryDesignBench does this in two phases:

  1. Design Phase: The agent implements a full library from an intentionally non-prescriptive specification.
  2. Evaluation Phase: We evaluate the library by observing multiple different implementer agents attempt to solve problems using it.

The only aspect we check in the design phase is if the library is installable in their internet-restricted environment, as we pre-install all libraries for evaluation since trying to install a library is not part of the signal we care about. We then task three agents (GPT-5.6 Luna on Codex, DeepSeek v4.1 Flash on mini-SWE-agent, and GLM 5.3 Flash on mini-SWE-agent) with solving programming problems with the library in as little code as possible. Each problem ×\times agent ×\times library is run in its own environment without internet access, with high reasoning, and with the library pre-installed. We use a highly prescriptive prompt to ensure agents attempt to fully exploit the library.

We now take these solutions and score each with:

score=pass rate(y)2⋅simplicity(y)\text{score} = \text{pass rate}(y)^2 \cdot \text{simplicity}(y)

Pass rate is the percentage of tests passed, and we square it to penalize incorrectness more harshly than simplicity. Simplicity is defined as:

simplicity(y)=1∣M∣∑m∈Mmin⁡{m(y∗)m(y), 1}\text{simplicity}(y) = \frac{1}{|\mathcal{M}|} \sum_{m \in \mathcal{M}} \min\left\{\frac{m(y^*)}{m(y)},\, 1\right\}

M\mathcal{M} is the set of static metrics we use to compare how far off the solution written with the library is from the golden reference, y∗y^*, written with the real production library. Each ratio is capped at 1, so a solution that beats the reference gets no extra credit. The four metrics are:

Section 2 of the paper covers the scoring and standard errors in more detail.

Agents Copy Human Designs But Worse

We evaluate 11 designer setups, covering 9 frontier models with some run in more than one harness, all at high reasoning. Each setup attempts each of the 15 library design tasks three times, which produces 45 libraries. The three implementer agents then use each library to solve every problem in its task. Across all 242 problems, that comes to 242 problems × 3 libraries × 3 implementers = 2,178 trials per setup.

We compare against two baselines that use the same 2,178 trials. In the first, implementers have no library and use a separate prompt. In the second, they have the human-written production library pre-installed. Neither baseline has three designed libraries to vary, so we instead run each problem-implementer pair three times.

Design Pareto
3035404550Score
↑ Beats production library
↓ Actively impedes agents
$0.5$1$2$5$10$20Cost to design one library ($, log) →
Evaluation Pareto
3035404550Score
$0.1$0.15$0.2Implementer cost per problem ($, log) →
Legend and notes
OpenAIAnthropicZ.aiDeepSeekMoonshot AIxAI
NL
No library. The same implementers solve each problem without any library.
PL
Production library. The same implementers use the task's real production library.
Reference score. Dashed at the NL and PL scores in both panels. A designed library scoring below NL actively impedes the agents using it; one scoring above PL beats the production library.
Pareto frontier. The best score at or below each cost. Points off it are faded.

Scores run 0–100: pass rate² × simplicity, averaged over every problem and implementer with tasks weighted equally. Hover or tap a point for exact values.

Opus 5.5 designs libraries that help downstream agents more than the human-written production libraries do. With them, implementers pass just as many tests while writing simpler code. Fable 5.1 roughly matches production. Correctness barely separates designers, since every setup passes about the same share of tests, so the ranking comes down to how much code implementers still have to write. At the other end, DeepSeek V4 Pro’s library actively hurts: implementers do worse with it than with no library at all. The harness also matters. Fable’s libraries are more useful when designed in mini-SWE-agent than in Claude Code, and Astra’s are slightly more useful in mini-SWE-agent than in Codex.

abc In all six libraries abc Astra only abc Fable only

GPT-6 Astra

Library 1
Command::new("inspect")
  .arg(Arg::new("verbose")
    .short('v')
    .action(ArgAction::Count))
  .arg(Arg::new("pids")
    .value_name("PID")
    .num_args(1..)
    .value_parser(
      ValueParser::u64()));
Library 2
Command::new("inspect")
  .arg(Arg::new("verbose")
    .short('v')
    .action(ArgAction::Count))
  .arg(Arg::new("pid")
    .short('p').long("pid")
    .value_parser(
      ValueParser::from_str::<u32>())
    .required(true));
Library 3
Command::new("inspect")
  .arg(Arg::new("verbose")
    .short('v')
    .action(ArgAction::Count))
  .arg(Arg::new("pid")
    .long("pid")
    .action(ArgAction::Append)
    .value_parser(
      ValueParser::u64()));

Fable 5.1

Library 1
Command::new("procs")
  .arg(Arg::count("verbose")
    .short('v'))
  .arg(Arg::option("pid")
    .short('p')
    .uint()
    .multiple());
Library 2
Command::new("proc")
  .arg(Arg::count("verbose")
    .short('v').global(true))
  .subcommand(Command::new(
      "inspect")
    .arg(Arg::positional("pid")
      .int().required(true));
Library 3
Command::new("procinfo")
  .arg(Arg::positional("pid")
    .int()
    .required(true))
  .arg(Arg::count("verbose")
    .short('v'));
Agents converge on the same designs. README quick-starts from three clirs libraries each by GPT-6 Astra and Fable 5.1, with unrelated arguments omitted. Both copy clap's builder API rather than its shorter derive macro. Hover a key entry to spotlight it.

Beyond harness peculiarities, the biggest trend we observed is that on 11/15 tasks, agents copied a design pattern from the corresponding production library. Surprisingly, these patterns are the exact ones the implementer agents used when given the production library. Our failure analysis indicates that the implementer agents write more complex code than needed because interfaces are either too rigid to use or require verbose code. In the former case, agents reimplement functionality the libraries provide, adding bugs in the process.

Most excess code traces back to the library. Primary failure category per designer. Failed tests classifies why at least one test failed; excess code classifies why the solution was longer than the reference.
Categories and method

Limited by the library: no path through the library as shipped does better.

Coverage
Nothing close to the needed capability exists; the fix is a new operation.
Correctness
The capability exists but has a bug: wrong output on a legitimate input, a misleading error, or a path too slow for the problem.
Rigidity
It nearly fits but cannot be adapted: a hard-coded policy, a data model that cannot hold the value, a monolithic operation, or a scope that excludes the input.
Verbosity
It fits, but using it takes excess code: a verbose interface, unpackaged wiring, shape conversion, or repeated declarations.

Library not fully exploited: a path through the library as shipped removes the code or fixes the failure.

Underused library
A simpler path through the library exists, but the implementer never found it, did not recognize it, or missed a fact needed to use it.

How. For each library from six mini-SWE-agent designers, we sampled three solutions that failed at least one test and wrote more code than the reference. A GPT-5.6 Luna auditor read the library, the implementer's trajectory, the tests, and the reference, then classified each symptom by the smallest library change that would have prevented it. The bars show the primary cause: 810 classifications per symptom, pooled across designers.

Takeaway. 82% of excess-code causes are limits of the library itself. Rigidity and Verbosity alone account for 64%, against 14% for missing capabilities. For failed tests, an underused library leads at 43%, and over half of those are an incomplete contract: the implementer found the capability but not a precondition or default needed to use it.

Each solution was audited once by one model, without running the proposed path. Shares describe partially passing solutions, not every solution.

Can Agents Use Libraries Without Instruction?

Our main evaluation prompt is rather heavy-handed because we need agents to fully exploit the library to measure its upper bound. But we also want to understand how models will behave given minimal additional instructions. Thus, we also release LibraryUseBench, our evaluation phase in which agents must use the production library to write as little code as possible. All agents use mini-SWE-agent and use High thinking. Opus 5.5, unsurprisingly, is the best model evaluated, but Sonnet 5.5 is close behind at just over half the cost per problem. The gap between them is small, but it holds up when we compare the two problem by problem. GPT-6 Luna is by far the cheapest model, but it passes only 71% of tests, and GPT-5.6 Luna rounds out the lower end.

LibraryUseBench
20253035404550556065707580Score
$0.01$0.02$0.05$0.1$0.2$0.5$1Cost per problem ($, log) →

Using a real library well scales with the model. LibraryUseBench score against cost per problem. Every model gets the task's production library and one instruction: use it, and write as little code as possible.

Legend and notes
OpenAIAnthropicZ.aiDeepSeek
Pareto frontier. The best score at or below each cost. Points off it are faded.
95% interval. The vertical bar through each point.

There is no design phase: the library is fixed to the production library, and the same problems are scored with the same formula, pass rate² × simplicity averaged over every problem with problems weighted equally. All models run in mini-SWE-agent. Hover or tap a point for exact values.

Do Agent-First Designs Help?

Designers tend to copy the production library, so we tried steering them away from it. We append a guidance prompt to the spec. It says the library will be used only by coding agents scored on how little code they write, and it asks the designer to:

Guided libraries copy fewer names from the production library and fold multi-step patterns into single calls. In clirs, for example, builder chains become one task-level call. Downstream solutions get shorter, and the score improves for every implementer, but it still lands just below the production library at roughly twice the design cost. Moving away from human designs helps, but we still haven’t found what the design agents prefer.

Designing explicitly for agents helps, modestly. Left: score gain per language when GPT-6 Astra designs with explicit agent-first guidance, split into simplicity and correctness. Right: clirs examples under each prompt.
Setup and results

Setup. GPT-6 Astra (Codex, high reasoning) designs each library with and without a guidance prompt appended to the unchanged specification. The prompt says only coding agents, scored on how little code they write, will use the library, and asks for consumer-first API sketches, runnable usage examples, and testing with subagents. We test the combined intervention, not each piece.

Results. The score rises from 44.1 to 46.4, just below the production library (46.6) and standard-prompt Opus 5.5 (48.9). It improves for all three implementers and three of four languages, mostly through shorter solutions: simplicity contributes +2.5 points and correctness −0.9. Guided libraries also look less like the production library, reproducing 13.4% of its exported names versus 19.4%; in clirs, builder chains become single task-level calls.

Cost. Guidance nearly doubles the cost of designing a library ($8.24 vs. $4.51) and leaves the cost per problem unchanged.

Conclusion

The full leaderboards and tasks are at ldbench.com. You can view the full repo here and the task repo here. You can read the full technical report here. If you have a task you would like to see, open an issue on the repo or get involved in the Discord.

This would not be possible without my amazing collaborators: Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, and Ludwig Schmidt. I also want to thank John Yang, Parth Asawa, Xavier Garcia, Ryan Carelli, Arun Kumar, Floriad Brand, and Nick Roberts for their helpful feedback and discussions. LDB is supported by DARPA, NSF, Prime Intellect, and Snorkel AI through the Open Benchmarks Grant.



Next Post
GPT-5.4 Writes Clean Code That Fails More Tests