Briefings

Short, decision-oriented analysis of a specific AI systems development.

Latest analysis

Recent research, benchmark reviews and technical updates.

No recent analysis is published for this topic yet.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Can SheetCompass Read Real Workbooks?

Can SheetCompass Read Real Workbooks?

SheetCompass tackles a production problem that ordinary table prompts often hide: spreadsheets are spatial workbooks, not flat text files. The paper was...

5 min read
MasDrift Measures Authorisation Drift

MasDrift Measures Authorisation Drift

MasDrift tests a quiet failure in multi-agent systems: the task gets delegated, but the user's boundary does not. The August 2026 paper matters because it...

5 min read
SWE-Bench Mobile Exposes Agent Variance

SWE-Bench Mobile Exposes Agent Variance

SWE-Bench Mobile tests coding tools against industry mobile-development work rather than public Python bug fixes. Its uncomfortable result is that the...

5 min read
ParEvalLayer Makes Partial Evals Decidable

ParEvalLayer Makes Partial Evals Decidable

ParEvalLayer is a 3 August 2026 proposal for a practical agent-evaluation problem: teams often see early task outcomes before a full benchmark run...

4 min read
IFCMemoryBench Finds the Missing Project Facts

IFCMemoryBench Finds the Missing Project Facts

IFCMemoryBench turns agent memory into an engineering test: can a building-information assistant remember project facts from earlier sessions and combine...

5 min read
The Agent Project That Should Have Been One LLM Call

The Agent Project That Should Have Been One LLM Call

Some enterprise agent projects fail because autonomy was added where a bounded single-call LLM design would have delivered cleaner behavior and lower operational risk.

10 min read
Open Source AI Impact: Who Wins When Models Get Cheap

Open Source AI Impact: Who Wins When Models Get Cheap

Open source AI used to be the cheaper substitute. In 2026, that is too small.

11 min read
AI Coding Agents: What Actually Works in Production

AI Coding Agents: What Actually Works in Production

GitHub reports that 46% of all new code is now AI-generated. Ninety-two percent of US developers use AI coding tools daily. Claude Code hit $2.5 billion...

16 min read
Seven Protocols, 1% Adoption: The Agent Economy's Infrastructure-Reality Gap

Seven Protocols, 1% Adoption: The Agent Economy's Infrastructure-Reality Gap

Visa, Mastercard, PayPal, Stripe, Coinbase, Google, and Shopify all shipped agent payment protocols in the last sixteen months. Seven competing standards...

6 min read
Anthropic's 186-Deal Experiment Shows What the Agent Economy Actually Looks Like

Anthropic's 186-Deal Experiment Shows What the Agent Economy Actually Looks Like

In December 2025, Anthropic gave 69 employees $100 each and told them to let Claude agents trade on their behalf. The agents bought and sold real items (services, digital goods, subscriptions) listed by other employees in a controlled marketplace. The experiment ran for several weeks. When it end...

4 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.