Benchmark Watch

Evaluation notes, benchmark interpretation, leaderboard skepticism, and measurement failures.

Guides and explainers

Detailed guides and practical technical analysis.

No guides are published for this topic yet.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
PM-Bench Tests Delayed Agent Memory

PM-Bench Tests Delayed Agent Memory

PM-Bench turns prospective memory into an agent evaluation problem: can a system remember to act on a future cue while ordinary work continues? The...

4 min read
Can Security Agents Work Under Human Custody?

Can Security Agents Work Under Human Custody?

Hack The Box's 2026 Global Cyber Skills Benchmark is a useful adoption signal because it measures where security practitioners chose to use AI agents...

6 min read
Bazaar Shows Pricing Agents Miss Profit

Bazaar Shows Pricing Agents Miss Profit

[Bazaar](https://arxiv.org/abs/2608.00102) is a benchmark submitted on 30 July 2026 for testing whether a language-model merchant can learn customer...

5 min read
SWE-Bench ProMax Tests Refactoring Depth

SWE-Bench ProMax Tests Refactoring Depth

SWE-Bench ProMax moves coding-agent evaluation from single-issue repair towards large, behaviour-preserving refactors. The paper was submitted on 10...

4 min read
Tool Agents Need Taint Boundaries, Not Bigger Scopes

Tool Agents Need Taint Boundaries, Not Bigger Scopes

APPA landed on arXiv on 27 July 2026 with a blunt result for tool-agent security: permission prompts are too late once poisoned or confidential data has...

4 min read
Chip-Design Agents Need Token ROI, Not Demo Flows

Chip-Design Agents Need Token ROI, Not Demo Flows

FluxBench landed on arXiv on 20 July 2026 with a useful warning for chip-design teams: two agents can start from the same foundation model and still...

5 min read
Terminal Agents Need Progress Curves, Not Victory Screens

Terminal Agents Need Progress Curves, Not Victory Screens

Long-Horizon-Terminal-Bench landed on arXiv in July 2026 with an awkward result for terminal-agent buyers: even the strongest tested model still failed...

4 min read
Million-Token Context Still Fails the Workload Test

Million-Token Context Still Fails the Workload Test

Anthropic reported on February 5, 2026 that Claude Opus 4.6 scored 76% on the 8-needle 1M-token MRCR v2 test while Claude Sonnet 4.5 scored 18.5% on the...

7 min read
Agent Evals Need Harnesses, Not More Scoreboards

Agent Evals Need Harnesses, Not More Scoreboards

AgentCompass first landed on arXiv on 15 July 2026 with a practical complaint: agent evaluation is fragmented, tightly coupled, and hard to reproduce...

3 min read
Coding Agent Benchmarks Hit the Generalization Wall

Coding Agent Benchmarks Hit the Generalization Wall

Scale's SWE-Bench Pro public leaderboard reports that top models scoring above 70% on SWE-Bench Verified fall to 23.3% for OpenAI GPT-5 and 23.1% for...

6 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.