Benchmark Watch
Evaluation notes, benchmark interpretation, leaderboard skepticism, and measurement failures.
Guides and explainers
Detailed guides and practical technical analysis.
No guides are published for this topic yet.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
Agent Test-Time Scaling Needs Reuse, Not More Rollouts
General AgentBench reports that running agents for more interaction steps or more sampled trajectories did not reliably improve ten leading agents,...
Multi-Agent Finance Workflows Need Cost Curves, Not More Agents
A March 2026 benchmark on financial-document processing makes the uncomfortable point: the most accurate multi-agent architecture was not the obvious...
Self-Improving Agents Have an Evaluator Problem
Anthropic's June 2026 update on recursive self-improvement is not a distant sci-fi warning. The company says its engineers now ship 8x as much code per...
Coding Agents Need Trajectory Reviews, Not Pass Bits
Most coding-agent benchmarks still compress a whole run into one bit: did the task pass? AgentLens argues that users experience the whole trajectory...
Tool-Use Agents Need Failure Labels, Not Pass Rates
Tool-use agents can fail in ways a final accuracy score hides, because the same wrong answer can come from skipped tools, ignored outputs, fabricated...
Agent Leaderboards Can Be Cheaper Without Being Safer
A March 2026 paper on efficient agent benchmarking found that mid-difficulty task subsets can remove large parts of an agent benchmark while preserving...
Multimodal Memory Tests Expose the Personal-Agent Gap
Product teams are turning memory into the selling point for personal agents. The hard question is no longer whether they can remember a preference; it is...
Power Grid Agents Need Constraint Tests, Not Chat Scores
A June 2026 power-systems benchmark argues that language-model agents can solve grid-engineering tasks, but the useful signal is narrower: the agent must...
TerminalWorld Makes Agent Benchmarks Harder to Fake
TerminalWorld turns public terminal recordings into validated agent tasks. The signal is not a higher leaderboard score. It is a harder benchmark supply chain.
Tool Agents Need State Diffs, Not API Call Scores
Tool-use agents are moving from choosing the right API to changing live product state. That makes a clean function-call score too small a test because...