LISTEN TO THIS ARTICLE

The paper When Tools Get in the Way isolates a failure mode that is easy to miss in agent design: a model can answer worse when an unnecessary but related tool is merely available. The authors tested whether tool availability changes answers to closed-domain questions that should be answered from the model's own knowledge, not from an external call Paturi et al..

Evidence base: the primary When Tools Get in the Way paper, OpenAI's function-calling guidance, Berkeley Function Calling Leaderboard context, and related Swarm Signal coverage on tool-use datasets, tool-call waits and tool-use regressions.

Key takeaways

  • The study pairs tool-requiring questions with closed-domain questions that should not need the tool Paturi et al..
  • Answer quality fell when an unnecessary related tool was available, even though the question should have been answerable directly Paturi et al..
  • A short scope-aware system instruction recovered much of the lost performance, so tool routing policy belongs in the release gate rather than only in prompt polish Paturi et al..

The earlier question is whether those tools make ordinary answers less reliable when no call is needed.

The failure happens before the call

Most tool-use evaluation asks whether the model calls the right function, fills arguments correctly or handles tool output. That is still necessary. The Berkeley Function Calling Leaderboard describes its current evaluation as measuring large language models' ability to call functions, or tools, accurately using real-world data BFCL.

This paper asks a narrower and uncomfortable question: what happens when the tool is present but not needed? The closed-domain query should be answered directly. The tool is related to the domain, so the model has a plausible reason to think about it, but the answer does not depend on calling it.

The result affects more than latency. The paper reports that accuracy drops even when the model rarely calls the unnecessary tool, which means the schema can change the model's answering behaviour before any external system is invoked Paturi et al..

Why production teams should care

Tool menus are often treated as harmless context. If a function is safe and potentially useful, teams leave it exposed and trust the model to ignore it. OpenAI's function-calling guide frames functions as tools defined by JSON schema and notes that the model can decide whether and how many functions to call by default OpenAI.

That default creates a design obligation. If a customer-support assistant always sees billing, policy, search, CRM and refund tools, the question is not only whether it picks the correct tool. The earlier question is whether those tools make ordinary answers less reliable when no call is needed.

The practical failure is subtle. A model may hedge, defer, over-explain, abstain or restructure the answer because the available tool implies that the current question might require external lookup. The visible symptom is not necessarily a bad API call; it can be a worse natural-language answer.

This is adjacent to ToolGrad's answer-first training pattern.

Scope instructions are a release control

The useful part of the paper is not that every tool list is dangerous. It is that a simple scope instruction recovered much of the lost performance in the reported setup Paturi et al.. That points to a cheap release test:

  • run closed-domain questions with tools hidden;
  • run the same questions with the production tool set visible;
  • run them again after a required tool interaction;
  • fail the release if answer quality drops without a justified call.

This is adjacent to ToolGrad's answer-first training pattern. ToolGrad treats executable traces as training material. This paper says the absence of a trace can also be evidence: the model should preserve direct-answer behaviour when the tool boundary says the tool is out of scope.

It also connects to progress-aware tool-call serving. Runtime plumbing can reduce wasted waits, but it cannot repair a model that has already let irrelevant tool availability distort the answer.

What to change in tool design

First, make tool exposure conditional. Do not show every safe tool on every turn. If a question can be answered from policy text already in context, retrieval or account tools should not be visible unless the user asks for account-specific action.

Second, test negative scope explicitly. A tool evaluation set should include examples where a tool looks relevant but must not be used. Without that slice, teams only learn whether the model can call tools, not whether it can ignore them.

Third, record no-call decisions. A trace that correctly avoids a tool is as important as a successful call. It tells you that the model understood the boundary between internal knowledge, retrieved context and side-effecting action.

The release question is simple: does adding the tool make the answer better, or does it merely make the model less sure of what it already knows?

Source trail

Research and technical sources:

Related Swarm Signal analysis: