Agent Design

How you actually build AI agents that work. Architectures, tool use, memory patterns, and the frameworks worth paying attention to.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
VAKRA Shows API Reasoning Decay

VAKRA Shows API Reasoning Decay

VAKRA is an August 2026 benchmark for checking whether systems can reason across APIs, retrieved documents and natural-language tool policies in one...

4 min read
ScrambleToolBench Finds Tool-Map Inertia

ScrambleToolBench Finds Tool-Map Inertia

ScrambleToolBench is a 3 August 2026 benchmark for a practical tool-use failure: a system can learn what hidden tools do, then fail to revise that map...

4 min read
ParEvalLayer Makes Partial Evals Decidable

ParEvalLayer Makes Partial Evals Decidable

ParEvalLayer is a 3 August 2026 proposal for a practical agent-evaluation problem: teams often see early task outcomes before a full benchmark run...

4 min read
Tool Agents Need Taint Boundaries, Not Bigger Scopes

Tool Agents Need Taint Boundaries, Not Bigger Scopes

APPA landed on arXiv on 27 July 2026 with a blunt result for tool-agent security: permission prompts are too late once poisoned or confidential data has...

4 min read
Terminal Agents Need Progress Curves, Not Victory Screens

Terminal Agents Need Progress Curves, Not Victory Screens

Long-Horizon-Terminal-Bench landed on arXiv in July 2026 with an awkward result for terminal-agent buyers: even the strongest tested model still failed...

4 min read
Agent Observability Is Escaping the Dashboard

Agent Observability Is Escaping the Dashboard

Agent observability is moving from vendor dashboards into trace contracts that make every model call, tool call, handoff, guardrail, and evaluator step inspectable.

3 min read
Agent Evals Need Harnesses, Not More Scoreboards

Agent Evals Need Harnesses, Not More Scoreboards

AgentCompass first landed on arXiv on 15 July 2026 with a practical complaint: agent evaluation is fragmented, tightly coupled, and hard to reproduce...

3 min read
Self-Improving Agents Have an Evaluator Problem

Self-Improving Agents Have an Evaluator Problem

Anthropic's June 2026 update on recursive self-improvement is not a distant sci-fi warning. The company says its engineers now ship 8x as much code per...

3 min read
The Agent Project That Should Have Been One LLM Call

The Agent Project That Should Have Been One LLM Call

Some enterprise agent projects fail because autonomy was added where a bounded single-call LLM design would have delivered cleaner behavior and lower operational risk.

10 min read
Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-use agents can fail in ways a final accuracy score hides, because the same wrong answer can come from skipped tools, ignored outputs, fabricated...

4 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.