LISTEN TO THIS ARTICLE

Tool-calling systems already know more about long-running work than the model server can see. A new paper, Ask the Tool, Don't Guess, argues that the missing interface is a progress signal from the running tool to the serving layer, so cache policy is driven by current work rather than a pre-call estimate Liu et al..

Evidence base: a systems paper on tool-call progress signals, current vLLM cache documentation, OpenAI's tool-calling guide, and related Swarm Signal coverage on tool-use reliability and stalled tool behaviour.

Key takeaways

  • Main change: tool calls can expose runtime progress without changing what the model sees.
  • Practical implication: serving systems can make better key-value cache decisions during tool waits.
  • Caveat or risk: this is a systems result, not proof that the tool output is correct or safe.

Progress hints do not make a tool trustworthy.

The waiting state is now part of serving

Function calling is usually described as an application loop: the model emits a tool call, the application runs it, then the result is passed back into the conversation OpenAI. That framing is correct, but it hides a serving problem. While the tool is running, the request's key-value cache may still occupy scarce accelerator memory.

The serving layer then has to decide whether to keep that cache warm, evict it, or recover it later. vLLM's automatic prefix caching documentation shows why these blocks matter: reused cache can avoid recomputing a shared prefix vLLM. In a tool-using workload, the difficult question is not only whether a prefix might be reused. It is whether the current tool wait is nearly over.

The paper's useful claim is narrow. Instead of asking the server to infer duration from a tool name, historical averages or declared estimates, the tool can report progress while it runs Liu et al.. That turns an invisible waiting period into an explicit scheduling signal.

What the progress hint changes

The authors test two kinds of signal: a rough fraction of work remaining and a near-end indication that the call is close to finishing Liu et al.. They also report that the harness recovers these signals without changing the information exposed to the agent, which matters because a progress channel should not become a hidden reasoning aid Liu et al..

The reported payoff is a serving metric, not a new reasoning benchmark. Plugged into a production engine, the approach reduced p90 time to first token after a tool call against an LRU baseline Liu et al.. That is the concrete reason builders should care: the delay after a tool returns is part of the user experience, and it is partly governed by cache policy.

There is an important boundary here. Progress hints do not make a tool trustworthy. They do not prove the tool result is fresh, authorised or complete. They only help the serving system decide how to allocate memory while waiting.

The running tool is the surface with the freshest information.

Where this fits with cache engineering

The broader cache story is moving away from one simple rule. Prefix caching saves prefill work when requests share context vLLM. KV offloading extends the same idea by moving completed cache blocks to larger, slower tiers and promoting them back on demand vLLM. The progress-hint paper adds another input: whether a paused request is likely to need its cache soon.

That matters most for agentic workloads with uneven tool latency. A search call, a browser action, a code run and a database report do not wait in the same way. A static guess made before the call starts can be wrong even if the average duration looked sensible. The running tool is the surface with the freshest information.

For teams already tracking tool traces, the natural implementation question is whether progress can be reported through the existing harness rather than bolted onto each application. The paper's answer is encouraging but still early: a harness can recover usable progress signals in public corpora, but production teams would need to check their own tool mix, side effects and security boundaries before exposing a new reporting channel Liu et al..

The operational test

The practical check is simple. Identify the tool calls that dominate wall-clock waiting time, then ask whether the serving system currently knows when each call is halfway through, stalled, or close to completion. If the answer is no, cache policy is probably guessing.

This connects to Swarm Signal's earlier coverage of tool-use agents needing failure labels and agents that stop using tools. Those pieces focus on whether the model used the tool correctly. This result sits one layer lower: once the tool is running, the runtime still needs a truthful account of progress.

The caution is that progress should remain separate from control authority. A tool should be able to say "nearly done" without being allowed to smuggle new instructions into the model context or trigger an unsafe retry. Treat the hint as scheduling metadata, not as part of the answer.

What to measure next

A good adoption test would report a small scorecard beside ordinary task success: tool wait time, post-tool time to first token, cache eviction or recovery rate during tool waits, and error rate after long-running calls. The latency measures show whether the user actually feels the improvement. The reliability measures catch the risk of making the cache smarter while leaving the workflow brittle.

For builders, the near-term decision is not whether every tool needs a progress API. It is whether the slowest, most common and most expensive tools should stop being silent while the serving layer manages memory around them.

Source trail

Research:

Technical context:

Related Swarm Signal analysis: