Agent Design
How you actually build AI agents that work. Architectures, tool use, memory patterns, and the frameworks worth paying attention to.
Guides and explainers
Detailed guides and practical technical analysis.
Latest analysis
Recent research, benchmark reviews and technical updates.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
Messier Maps Agent Benchmark Drift
Messier, revised on 30 August 2026, treats agent evaluation as a data-integration problem rather than another leaderboard race...
Tool Parallelism Has A Scheduling Problem
PeakBench is a useful benchmark because it tests a failure that rarely appears in tool-use leaderboards: an agent can choose the right calls, respect the...
VAKRA Shows API Reasoning Decay
VAKRA is an August 2026 benchmark for checking whether systems can reason across APIs, retrieved documents and natural-language tool policies in one...
ScrambleToolBench Finds Tool-Map Inertia
ScrambleToolBench is a 3 August 2026 benchmark for a practical tool-use failure: a system can learn what hidden tools do, then fail to revise that map...
ParEvalLayer Makes Partial Evals Decidable
ParEvalLayer is a 3 August 2026 proposal for a practical agent-evaluation problem: teams often see early task outcomes before a full benchmark run...
Tool Agents Need Taint Boundaries, Not Bigger Scopes
APPA landed on arXiv on 27 July 2026 with a blunt result for tool-agent security: permission prompts are too late once poisoned or confidential data has...
Terminal Agents Need Progress Curves, Not Victory Screens
Long-Horizon-Terminal-Bench landed on arXiv in July 2026 with an awkward result for terminal-agent buyers: even the strongest tested model still failed...
Agent Observability Is Escaping the Dashboard
Agent observability is moving from vendor dashboards into trace contracts that make every model call, tool call, handoff, guardrail, and evaluator step inspectable.
Agent Evals Need Harnesses, Not More Scoreboards
AgentCompass first landed on arXiv on 15 July 2026 with a practical complaint: agent evaluation is fragmented, tightly coupled, and hard to reproduce...
Self-Improving Agents Have an Evaluator Problem
Anthropic's June 2026 update on recursive self-improvement is not a distant sci-fi warning. The company says its engineers now ship 8x as much code per...