Failure Briefs

Postmortem-style analysis of AI system failures, fragility, and production risk.

Practical tools

Templates for budgets and project planning

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Your Agent's System Prompt Is Fighting Itself

Your Agent's System Prompt Is Fighting Itself

A framework called Arbiter treats agent system prompts as auditable code. Applied to Claude Code, Codex CLI, and Gemini CLI, it found 152 interference patterns — including critical contradictions and a structural data loss bug — for a total cost of $0.27.

3 min read
Alignment Works in English. In Japanese, It Backfires.

Alignment Works in English. In Japanese, It Backfires.

A new study shows the same alignment intervention that produces strong safety effects in English reverses direction in Japanese, increasing harmful outputs. Tested across 1,584 simulations, 16 languages, and three model families.

3 min read
Most AI Agents Don't Know When They're Wrong

Most AI Agents Don't Know When They're Wrong

A 4B parameter model just matched GPT-4o on tool-use tasks by learning to verify its own actions. The CoVe paper shows verification-first training beats the retry-and-pray approach plaguing production

6 min read
One Fake Source Broke Every Agent

One Fake Source Broke Every Agent

A single misinformation article injected into search rankings crashed GPT-5's accuracy from 65.1% to 18.2%. The agents had unlimited access to truthful sources and couldn't be bothered to look.

3 min read
AI Agent Security Checklist

AI Agent Security Checklist

AI agents don't just have a security problem. They have a fundamentally different security problem than the systems they're replacing. Five attack surfaces and the defense patterns that actually work.

3 min read
Chain-of-Thought Prompting: When It Works, When It Fails, and Why

Chain-of-Thought Prompting: When It Works, When It Fails, and Why

Chain-of-thought is the most studied prompting technique in AI, and the most misapplied. A decision framework for when it helps, when it hurts, and what it costs.

9 min read
Your Multi-Agent System Is Colliding

Your Multi-Agent System Is Colliding

Most production agent systems don't fail because individual agents are stupid. They fail because three agents tried to solve the same problem...

6 min read
Computer-Use Agents Can't Stop Breaking Things

Computer-Use Agents Can't Stop Breaking Things

Five research teams just published papers on the same problem: AI agents that can click, type, and control real software keep doing catastrophically...

7 min read
Synthetic Data Won't Save You From Model Collapse

Synthetic Data Won't Save You From Model Collapse

The AI industry's running out of internet. Every major lab's already scraped the same corpus, and the easy gains from scaling data are tapering. The...

19 min read
AI Agents Are Security's Newest Nightmare

AI Agents Are Security's Newest Nightmare

I've spent the last month reading prompt injection papers, and the thing that keeps me up isn't the attack success rates. It's how many production systems...

16 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.