Practical tools
Templates for planning work and money
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
Power Grid Agents Need Constraint Tests, Not Chat Scores
A June 2026 power-systems benchmark argues that language-model agents can solve grid-engineering tasks, but the useful signal is narrower: the agent must...
TerminalWorld Makes Agent Benchmarks Harder to Fake
TerminalWorld turns public terminal recordings into validated agent tasks. The signal is not a higher leaderboard score. It is a harder benchmark supply chain.
Agent Tool Menus Are a Safety Surface
New agent benchmarks suggest the visible tool menu is not a neutral implementation detail. It changes success, cost, wrong-tool calls, and risk exposure.
Agent Memory Fails on Relationships, Not Recall
New June 2026 memory benchmarks show why long-running agents fail when facts conflict, evolve, or depend on hidden relationships.
Agent Benchmarking Doesn't Need Every Task
Efficient agent benchmarking points to a cheaper way to compare agents: run the tasks that still separate systems, not every task in the suite.
SMAC-Talk Shows Agent Chat Is Not Coordination
SMAC-Talk adds natural-language communication and deception to StarCraft-style multi-agent evaluation. The result is a useful warning: agent chat can expose coordination failure as easily as it fixes
Tool Agents Need State Diffs, Not API Call Scores
Tool-use agents are moving from choosing the right API to changing live product state. That makes a clean function-call score too small a test because...
Healthcare AI Agents Move Beyond Drug Discovery
Healthcare AI agents are moving into admin, triage and prior-authorisation workflows. The real gate is safety, evidence and accountable handoff.
Multilingual Agents Need Workflow Tests, Not Translation Scores
PolyWorkBench landed on arXiv on 7 July 2026 with a useful correction to enterprise-agent hype: a global workflow is not a translated English task...
Self-Improving Agents Need Hard Boundaries
Self-improving agents can rewrite code, prompts and memory. Production teams need rollback, approval gates and evaluator change control.