evals
Guides and explainers
Detailed guides and practical technical analysis.
No guides are published for this topic yet.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
TerminalWorld Makes Agent Benchmarks Harder to Fake
TerminalWorld turns public terminal recordings into validated agent tasks. The signal is not a higher leaderboard score. It is a harder benchmark supply chain.
Agent Tool Menus Are a Safety Surface
New agent benchmarks suggest the visible tool menu is not a neutral implementation detail. It changes success, cost, wrong-tool calls, and risk exposure.
Agent Memory Fails on Relationships, Not Recall
New June 2026 memory benchmarks show why long-running agents fail when facts conflict, evolve, or depend on hidden relationships.
Agent Benchmarking Doesn't Need Every Task
Efficient agent benchmarking points to a cheaper way to compare agents: run the tasks that still separate systems, not every task in the suite.
Evaluation-Aware Memory: How Agents Should Remember What They Can Prove
Agent memory should promote facts only after evals prove they improve task outcomes, not just because retrieval found them.