Benchmark Watch

Evaluation notes, benchmark interpretation, leaderboard skepticism, and measurement failures.

Guides and explainers

Detailed guides and practical technical analysis.

No guides are published for this topic yet.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Multilingual Agents Need Workflow Tests, Not Translation Scores

Multilingual Agents Need Workflow Tests, Not Translation Scores

PolyWorkBench landed on arXiv on 7 July 2026 with a useful correction to enterprise-agent hype: a global workflow is not a translated English task...

5 min read
Agent Benchmarks Need Runtime Receipts, Not Model Labels

Agent Benchmarks Need Runtime Receipts, Not Model Labels

RuBench's revised 19 July 2026 release contains a small but important warning for coding-agent buyers: one audited product configuration silently...

4 min read
Data Agents Need Exploration Budgets, Not SQL Magic

Data Agents Need Exploration Budgets, Not SQL Magic

Data Agent Benchmark landed on arXiv on 21 March 2026 with a result that should make enterprise analytics teams pause: the best tested frontier model...

4 min read
Assistant Agents Need Reminder Tests, Not Recall Scores

Assistant Agents Need Reminder Tests, Not Recall Scores

Most agent-memory benchmarks ask whether a model can recover old information. PM-Bench asks a harsher question: can an agent remember to do the right...

4 min read
The 12-to-72 Problem: Computer-Use Agents Hit Human Scores but Miss the Point

The 12-to-72 Problem: Computer-Use Agents Hit Human Scores but Miss the Point

Computer-use agents jumped from 12% to 72% on OSWorld in 18 months. The scores look like progress. The latency and efficiency numbers tell a different story.

4 min read
Why Multi-Agent Papers Don't Replicate in Production

Why Multi-Agent Papers Don't Replicate in Production

A paper from Tran and Kiela tested 28 multi-agent configurations across four architectures: Sequential, Parallel, Debate, and Ensemble. Every single one...

7 min read
Multimodal Agents Score 40% Where Humans Score 72%

Multimodal Agents Score 40% Where Humans Score 72%

Every frontier lab now ships models that see, hear, and read. The assumption is that more modalities mean more capable agents. The benchmarks tell a...

6 min read
Computer-Use Agents Fail Long Workflows, Not Mouse Clicks

Computer-Use Agents Fail Long Workflows, Not Mouse Clicks

Computer-use agents are clearing more short benchmark tasks, but the new failure line is workflow length. A June 2026 benchmark called OSWorld 2.0 tests...

5 min read
How to Build Agent Evals That Catch Real Failures

How to Build Agent Evals That Catch Real Failures

Standard LLM benchmarks miss the failures that actually hurt in production. Here's how to build an evaluation system for agents that catches cascading errors, trajectory drift, and policy violations before they reach users.

9 min read
Multi-Agent Systems Are Booming — But Real-Work Benchmarks Still Bite

Multi-Agent Systems Are Booming — But Real-Work Benchmarks Still Bite

Multi-agent workflows are growing fast, but APEX-Agents, AgentRx, Databricks, and Gartner show a gap between adoption, task success, and production readiness.

6 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.