Computer Science editorial
The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior
The core problem
As large language models (LLMs) improve, agentic systems are becoming increasingly important, and tools are a key design dimension because they determine how agents access information and take action in their environments. Prior work on agent tooling has primarily focused on expanding what agents can do—adding new capabilities, APIs, and actions—but has paid less systematic attention to how those capabilities are organized and exposed to the model. The authors refer to this latter design dimension as **tool architecture**.
The paper studies tool architecture in coding agents through controlled experiments on repository-level issue fixing. It compares six tool architectures that hold the underlying information and actions similar while varying how they are organized and exposed to the model, across three actors and a total of 11,700 trajectories. The central question is whether tool architecture alone—independent of raw capability—changes agent behavior. The results show that it does: even when tools provide similar capabilities, tool architecture changes agent behavior in measurable, sometimes dramatic ways.
Innovation
The experiments show that, even when tools provide similar capabilities, tool architecture changes agent behavior. Three headline findings stand out:
1. **Consistency:** Compared to a basic architecture where the agent has only the `bash` tool, more structured low-level interfaces improve consistency across repeated attempts by up to **4.7×**.
2. **Exploration:** Natural-language search broadens repository exploration and increases access to relevant files by more than **11%**.
3. **Efficiency:** Python CodeAct-style interfaces achieve similar task performance with **41.6% fewer steps** and **56.3% lower token usage**.
By contrast, lightweight text-based cognitive-scaffolding tools, such as tools that let the agent record intermediate reasoning, have **limited effect** on actor behavior. This asymmetry is notable: architectural changes that reshape how information and actions are exposed to the model produce large behavioral shifts, while tools that merely add a reasoning scratchpad do not.
The consistency result can be summarized as a relative improvement:
\text{Consistency Gain} = \frac{\text{Consistency}_{\text{structured low-level}}}{\text{Consistency}_{\text{bash-only}}}
Why it matters
The results reframe tool design for coding agents. The dominant question in prior work has been *what can the agent do?*—expanding the action space. This paper argues that *how* those actions and information are organized and exposed—the tool architecture—is an equally important, and previously under-systematized, design dimension.
The findings suggest several mechanisms. Structured low-level interfaces may reduce ambiguity in how the agent invokes operations, leading to more repeatable behavior across attempts (up to 4.7× consistency). Natural-language search may lower the cognitive barrier to exploring the repository, increasing access to relevant files by more than 11%. Python CodeAct-style interfaces may allow the agent to batch and compose actions more efficiently, yielding similar task performance with 41.6% fewer steps and 56.3% lower token usage.
The null result for lightweight text-based cognitive-scaffolding tools is equally informative: merely giving the agent a place to record intermediate reasoning does not, by itself, change actor behavior. This suggests that the leverage lies in the interface between the agent and its environment—how capabilities are surfaced—rather than in auxiliary reasoning aids.
A conceptual flow of the experimental comparison can be represented as:
Practically, the paper implies that agent builders should treat tool architecture as a first-class design variable, comparable to model choice or prompt design. The large effect sizes—4.7× consistency, >11% exploration gain, 41.6% fewer steps, 56.3% lower tokens—suggest that architectural choices can deliver substantial gains without changing the underlying capabilities available to the agent. The limited effect of cognitive-scaffolding tools further cautions against assuming that any added tooling improves behavior; the interface, not the mere presence of a tool, is what shapes the agent.
Who should read this
Opening member content…