Briefings
Short, decision-oriented analysis of a specific AI systems development.
Guides and explainers
Detailed guides and practical technical analysis.
Latest analysis
Recent research, benchmark reviews and technical updates.
No recent analysis is published for this topic yet.
Practical tools
Templates for budgets and project planning
BoredTools offers practical spreadsheets for budgets, freelance work and small projects.
Can SheetCompass Read Real Workbooks?
SheetCompass tackles a production problem that ordinary table prompts often hide: spreadsheets are spatial workbooks, not flat text files. The paper was...
MasDrift Measures Authorisation Drift
MasDrift tests a quiet failure in multi-agent systems: the task gets delegated, but the user's boundary does not. The August 2026 paper matters because it...
SWE-Bench Mobile Exposes Agent Variance
SWE-Bench Mobile tests coding tools against industry mobile-development work rather than public Python bug fixes. Its uncomfortable result is that the...
ParEvalLayer Makes Partial Evals Decidable
ParEvalLayer is a 3 August 2026 proposal for a practical agent-evaluation problem: teams often see early task outcomes before a full benchmark run...
IFCMemoryBench Finds the Missing Project Facts
IFCMemoryBench turns agent memory into an engineering test: can a building-information assistant remember project facts from earlier sessions and combine...
The Agent Project That Should Have Been One LLM Call
Some enterprise agent projects fail because autonomy was added where a bounded single-call LLM design would have delivered cleaner behavior and lower operational risk.
Open Source AI Impact: Who Wins When Models Get Cheap
Open source AI used to be the cheaper substitute. In 2026, that is too small.
AI Coding Agents: What Actually Works in Production
GitHub reports that 46% of all new code is now AI-generated. Ninety-two percent of US developers use AI coding tools daily. Claude Code hit $2.5 billion...
Seven Protocols, 1% Adoption: The Agent Economy's Infrastructure-Reality Gap
Visa, Mastercard, PayPal, Stripe, Coinbase, Google, and Shopify all shipped agent payment protocols in the last sixteen months. Seven competing standards...
Anthropic's 186-Deal Experiment Shows What the Agent Economy Actually Looks Like
In December 2025, Anthropic gave 69 employees $100 each and told them to let Claude agents trade on their behalf. The agents bought and sold real items (services, digital goods, subscriptions) listed by other employees in a controlled marketplace. The experiment ran for several weeks. When it end...