Tyler

X

Practical tools

Templates for planning work and money

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-Use Agents Need Failure Labels, Not Pass Rates

Tool-use agents can fail in ways a final accuracy score hides, because the same wrong answer can come from skipped tools, ignored outputs, fabricated...

4 min read
Agent Cost Optimization: How to Track and Reduce LLM Spend

Agent Cost Optimization: How to Track and Reduce LLM Spend

Token prices dropped 280x over two years. Enterprise AI budgets rose 320% in the same period. That's not a paradox. It's what happens when agentic...

18 min read
Agent Sandboxes Need Egress Budgets, Not Trust Prompts

Agent Sandboxes Need Egress Budgets, Not Trust Prompts

The live risk in agent security is shifting from "will the model say something unsafe?" to "what can the harness actually touch after the model decides?"...

5 min read
Inference Optimization: From 10x Cost to 10x Speed

Inference Optimization: From 10x Cost to 10x Speed

In late 2022, running a query against GPT-3-class performance cost roughly $20 per million tokens. By March 2026, multiple models exceed that same...

10 min read
RAG Cost Attacks Turn Retrieval Into a Budget Risk

RAG Cost Attacks Turn Retrieval Into a Budget Risk

A June 2026 paper on retrieval-augmented inference cost attacks reports a failure mode that many RAG teams are not testing: poisoned external documents...

4 min read
Agent Observability Needs Provenance, Not More Logs

Agent Observability Needs Provenance, Not More Logs

Agent observability is drifting toward a familiar trap: capture every trace, then ask an engineer to work out why the agent did the wrong thing. A June...

4 min read
Scaling Laws Explained for Practitioners: What Actually Matters in 2026

Scaling Laws Explained for Practitioners: What Actually Matters in 2026

Scaling laws promised a simple deal: spend more compute, get better models. For three years, that deal held. Kaplan et al. drew the first power-law curves...

18 min read
Agent Leaderboards Can Be Cheaper Without Being Safer

Agent Leaderboards Can Be Cheaper Without Being Safer

A March 2026 paper on efficient agent benchmarking found that mid-difficulty task subsets can remove large parts of an agent benchmark while preserving...

5 min read
Multi-Agent Systems Need Specs Before More Agents

Multi-Agent Systems Need Specs Before More Agents

Multi-agent systems are getting easier to assemble and harder to trust. A new June 2026 paper from Cisco researchers argues that the missing layer is not...

5 min read
Multimodal Memory Tests Expose the Personal-Agent Gap

Multimodal Memory Tests Expose the Personal-Agent Gap

Product teams are turning memory into the selling point for personal agents. The hard question is no longer whether they can remember a preference; it is...

5 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.