Tyler

X

Practical tools

Templates for planning work and money

BoredTools offers practical spreadsheets for budgets, freelance work and small projects.

Browse templates Free budget tracker
VAKRA Shows API Reasoning Decay

VAKRA Shows API Reasoning Decay

VAKRA is an August 2026 benchmark for checking whether systems can reason across APIs, retrieved documents and natural-language tool policies in one...

4 min read
Memory Scores Can Inflate Agent Rewards

Memory Scores Can Inflate Agent Rewards

Memory Reward Inflation names a failure in self-improving systems that learn from stored episodes: the score attached to a memory can become a reward...

5 min read
Can SheetCompass Read Real Workbooks?

Can SheetCompass Read Real Workbooks?

SheetCompass tackles a production problem that ordinary table prompts often hide: spreadsheets are spatial workbooks, not flat text files. The paper was...

5 min read
SWE-Bench ProMax Tests Refactoring Depth

SWE-Bench ProMax Tests Refactoring Depth

SWE-Bench ProMax moves coding-agent evaluation from single-issue repair towards large, behaviour-preserving refactors. The paper was submitted on 10...

4 min read
MasDrift Measures Authorisation Drift

MasDrift Measures Authorisation Drift

MasDrift tests a quiet failure in multi-agent systems: the task gets delegated, but the user's boundary does not. The August 2026 paper matters because it...

5 min read
SWE-Bench Mobile Exposes Agent Variance

SWE-Bench Mobile Exposes Agent Variance

SWE-Bench Mobile tests coding tools against industry mobile-development work rather than public Python bug fixes. Its uncomfortable result is that the...

5 min read
ScrambleToolBench Finds Tool-Map Inertia

ScrambleToolBench Finds Tool-Map Inertia

ScrambleToolBench is a 3 August 2026 benchmark for a practical tool-use failure: a system can learn what hidden tools do, then fail to revise that map...

4 min read
EcoAgent-Bench Prices Escalation Choices

EcoAgent-Bench Prices Escalation Choices

EcoAgent-Bench is a 6 August 2026 benchmark for a deployment question that ordinary task-success scores blur: when should a tool-using system spend more,...

5 min read
ParEvalLayer Makes Partial Evals Decidable

ParEvalLayer Makes Partial Evals Decidable

ParEvalLayer is a 3 August 2026 proposal for a practical agent-evaluation problem: teams often see early task outcomes before a full benchmark run...

4 min read
ToolPrivacyBench Exposes Purpose Drift

ToolPrivacyBench Exposes Purpose Drift

ToolPrivacyBench turns agent privacy into a tool-call audit: when a workflow succeeds, did each tool receive only the private facts it needed? The June...

5 min read
Swarm Signal
0:00
0:00
Up Next

Queue is empty. Click "+ Queue" on any article to add it.