Autoprompt Skill Review: Cutting Agentic Failures by 45%, Measured

Autoprompt Skill Review: Cutting Agentic Failures by 45%, Measured

Autoprompt is an open-source agent skill that wraps a coding agent in a plan, build, check, verify and finish loop. Its headline result — 45% fewer failures — is 29 down to 16 failures on 89 Terminal-Bench 2.1 tasks, a +14.61 percentage-point pass-rate gain whose 95% confidence interval is +2.02pp to +27.20pp. The direction is real; the precision is not. This review is deliberately narrow. We recap the product in one paragraph and then spend the rest of the article on the three things no catalog page does: compute the error bar, price the trade-off, and test the claim against category-level evidence. If you want the feature tour, see our earlier Autoprompt walkthrough. ...

October 1, 2026 · 19 min · baeseokjae