Benchmark Watch
Evaluation notes, benchmark interpretation, leaderboard skepticism, and measurement failures.
Guides and explainers
Detailed guides and practical technical analysis.
No guides are published for this topic yet.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
PM-Bench Tests Delayed Agent Memory
PM-Bench turns prospective memory into an agent evaluation problem: can a system remember to act on a future cue while ordinary work continues? The...
Can Security Agents Work Under Human Custody?
Hack The Box's 2026 Global Cyber Skills Benchmark is a useful adoption signal because it measures where security practitioners chose to use AI agents...
Bazaar Shows Pricing Agents Miss Profit
[Bazaar](https://arxiv.org/abs/2608.00102) is a benchmark submitted on 30 July 2026 for testing whether a language-model merchant can learn customer...
SWE-Bench ProMax Tests Refactoring Depth
SWE-Bench ProMax moves coding-agent evaluation from single-issue repair towards large, behaviour-preserving refactors. The paper was submitted on 10...
Tool Agents Need Taint Boundaries, Not Bigger Scopes
APPA landed on arXiv on 27 July 2026 with a blunt result for tool-agent security: permission prompts are too late once poisoned or confidential data has...
Chip-Design Agents Need Token ROI, Not Demo Flows
FluxBench landed on arXiv on 20 July 2026 with a useful warning for chip-design teams: two agents can start from the same foundation model and still...
Terminal Agents Need Progress Curves, Not Victory Screens
Long-Horizon-Terminal-Bench landed on arXiv in July 2026 with an awkward result for terminal-agent buyers: even the strongest tested model still failed...
Million-Token Context Still Fails the Workload Test
Anthropic reported on February 5, 2026 that Claude Opus 4.6 scored 76% on the 8-needle 1M-token MRCR v2 test while Claude Sonnet 4.5 scored 18.5% on the...
Agent Evals Need Harnesses, Not More Scoreboards
AgentCompass first landed on arXiv on 15 July 2026 with a practical complaint: agent evaluation is fragmented, tightly coupled, and hard to reproduce...
Coding Agent Benchmarks Hit the Generalization Wall
Scale's SWE-Bench Pro public leaderboard reports that top models scoring above 70% on SWE-Bench Verified fall to 23.3% for OpenAI GPT-5 and 23.1% for...