Autoprompt Skill Review: Cutting Agentic Failures by 45%, Measured

Autoprompt Skill Review: Cutting Agentic Failures by 45%, Measured

Autoprompt is an open-source agent skill that wraps a coding agent in a plan, build, check, verify and finish loop. Its headline result — 45% fewer failures — is 29 down to 16 failures on 89 Terminal-Bench 2.1 tasks, a +14.61 percentage-point pass-rate gain whose 95% confidence interval is +2.02pp to +27.20pp. The direction is real; the precision is not. This review is deliberately narrow. We recap the product in one paragraph and then spend the rest of the article on the three things no catalog page does: compute the error bar, price the trade-off, and test the claim against category-level evidence. If you want the feature tour, see our earlier Autoprompt walkthrough. ...

October 1, 2026 · 19 min · baeseokjae
Autoprompt Skill Review: The Coding Agent Prompt Skill That Cuts Failures 45%

Autoprompt Skill Review: The Coding Agent Prompt Skill That Cuts Failures 45%

Autoprompt is a coding agent prompt skill that injects an orchestration procedure — plan, delegate, implement, verify — into 11 host agents, including Claude Code, Codex, and OpenCode. Its headline claim of 45% fewer failures comes from a single 89-task Terminal-Bench 2.1 run where failures fell from 29 to 16. That result is real arithmetic but version 1 evidence, and it costs roughly 3x wall-clock time and 2x tokens. What the 45% Figure Actually Measures The number everyone quotes is not a percentage of tasks solved and it is not a percentage reduction in tokens. It is a reduction in failures on one benchmark, measured once. ...

September 28, 2026 · 14 min · baeseokjae