<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Coding Agent Self-Verification Loop on RockB</title><link>https://baeseokjae.github.io/tags/coding-agent-self-verification-loop/</link><description>Recent content in Coding Agent Self-Verification Loop on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:51:12 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/coding-agent-self-verification-loop/index.xml" rel="self" type="application/rss+xml"/><item><title>Autoprompt Skill Review: Cutting Agentic Failures by 45%, Measured</title><link>https://baeseokjae.github.io/posts/autoprompt-coding-agent-skill/</link><pubDate>Thu, 01 Oct 2026 00:51:12 +0000</pubDate><guid>https://baeseokjae.github.io/posts/autoprompt-coding-agent-skill/</guid><description>Autoprompt&amp;#39;s 45% claim recomputed: +14.61pp gain, 95% CI +2.02pp to +27.20pp, and roughly 1.5 pass-rate points per extra 10% of tokens.</description><content:encoded><![CDATA[<p>Autoprompt is an open-source agent skill that wraps a coding agent in a plan, build, check, verify and finish loop. Its headline result — 45% fewer failures — is 29 down to 16 failures on 89 Terminal-Bench 2.1 tasks, a +14.61 percentage-point pass-rate gain whose 95% confidence interval is +2.02pp to +27.20pp. The direction is real; the precision is not.</p>
<p>This review is deliberately narrow. We recap the product in one paragraph and then spend the rest of the article on the three things no catalog page does: compute the error bar, price the trade-off, and test the claim against category-level evidence. If you want the feature tour, see our <a href="https://baeseokjae.github.io/posts/autoprompt-skill-coding-agent-2026/">earlier Autoprompt walkthrough</a>.</p>
<h2 id="what-is-the-autoprompt-skill-in-one-paragraph">What Is the Autoprompt Skill in One Paragraph?</h2>
<p>Autoprompt is a coordination layer, not another coding agent. It installs into a host tool — Claude Code, Codex, OpenCode and eight others — and, when explicitly invoked, runs a fixed loop: choose a route, plan, build, check, independently verify, then finish. Roles are split across levels L0 to L4 so that no single agent plans a change, approves it and certifies its own verification. It requires Node.js 20+, Python 3.11+ with PyYAML and Bash 4.3+, and v2.0.0 ships a single CLI that activates the skill per provider. That is the whole product; everything below is about whether its number survives scrutiny.</p>
<h2 id="what-does-45-fewer-failures-actually-measure">What Does &ldquo;45% Fewer Failures&rdquo; Actually Measure?</h2>
<p>It measures one metric on one benchmark, once. Baseline OpenCode 1.18.7 solved 60 of 89 Terminal-Bench 2.1 tasks. With Autoprompt active, the same harness on the same model solved 73 of 89. Failures fell from 29 to 16.</p>
<p>That is a defensible statement, and it is also the limit of what the published evidence supports. The figure is a reduction in the failure <strong>rate</strong> on a single 89-task run, not a reduction in defects you will see in your repository, and not a comparison against any other tool.</p>
<table>
  <thead>
      <tr>
          <th>Read the same result a different way</th>
          <th>Baseline</th>
          <th>With Autoprompt</th>
          <th>Change</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tasks solved</td>
          <td>60 / 89</td>
          <td>73 / 89</td>
          <td>+13 tasks</td>
      </tr>
      <tr>
          <td>Pass rate</td>
          <td>67.42%</td>
          <td>82.02%</td>
          <td>+14.61 pp</td>
      </tr>
      <tr>
          <td>Failure rate</td>
          <td>32.58%</td>
          <td>17.98%</td>
          <td>−14.61 pp</td>
      </tr>
      <tr>
          <td>Failures (count)</td>
          <td>29</td>
          <td>16</td>
          <td>−44.8% (~45%)</td>
      </tr>
      <tr>
          <td>Relative pass-rate gain</td>
          <td>—</td>
          <td>—</td>
          <td>+21.67%</td>
      </tr>
      <tr>
          <td>Failures per solved task</td>
          <td>0.483</td>
          <td>0.219</td>
          <td>−54.7%</td>
      </tr>
  </tbody>
</table>
<p>Every row is the same data. The &ldquo;45%&rdquo; headline is the most flattering row that is still arithmetically true, which is exactly why it propagated and the rest did not.</p>
<h2 id="why-is-a-45-reduction-really-a-181x-claim">Why Is a 45% Reduction Really a 1.81x Claim?</h2>
<p>Because a percentage reduction in failures is not a percentage improvement in capability, and the difference is large enough to change how you budget for it.</p>
<p>Two useful translations exist. In odds terms, 29 failures becoming 16 is <strong>1.81x fewer failures</strong> — the treatment arm fails about 55% as often. In rate terms, 32.58% becoming 17.98% is a <strong>44.8% reduction in the failure rate</strong>, rounded to 45%. Both are honest. &ldquo;45% better agent&rdquo; is not, and it is the phrase most readers hear.</p>
<p>For a working developer the odds reading is the operative one: on a hard task where the plain agent has a 33% chance of failing, the wrapped agent has roughly an 18% chance, so you still expect to repair one task in six. That is a meaningful improvement and it is not a transformation.</p>
<h2 id="what-is-the-confidence-interval-on-the-45-claim">What Is the Confidence Interval on the 45% Claim?</h2>
<p>The error bar nobody publishes is +2.02pp to +27.20pp at 95% confidence. That is the single most useful number in this review.</p>
<table>
  <thead>
      <tr>
          <th>Statistic (60/89 vs 73/89, n = 89 per arm)</th>
          <th>Value</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Absolute difference in pass rate</td>
          <td>+14.61 pp</td>
      </tr>
      <tr>
          <td>95% CI, two-proportion interval</td>
          <td><strong>+2.02 pp to +27.20 pp</strong></td>
      </tr>
      <tr>
          <td>Two-proportion z</td>
          <td>2.24</td>
      </tr>
      <tr>
          <td>Two-sided p (normal)</td>
          <td>0.025</td>
      </tr>
      <tr>
          <td>Fisher exact, one-sided p</td>
          <td>0.019</td>
      </tr>
      <tr>
          <td>Fisher exact, two-sided p</td>
          <td>0.038</td>
      </tr>
      <tr>
          <td>Statistical power to detect this effect at n=89</td>
          <td>≈61%</td>
      </tr>
      <tr>
          <td>n per arm to detect +5pp at 80% power</td>
          <td>≈1,319</td>
      </tr>
      <tr>
          <td>n per arm to detect +3pp at 80% power</td>
          <td>≈3,735</td>
      </tr>
  </tbody>
</table>
<p>Read the interval, not the point estimate. The result is statistically significant — the lower bound clears zero — but the CI spans from &ldquo;a marginal benefit you might not notice&rdquo; to &ldquo;a near doubling of the failure reduction you were promised.&rdquo; A 14.61-point change on 89 tasks cannot distinguish those two worlds, and it cannot rule out an effect as small as two points.</p>
<p>The lower bound is worth translating. At +2.02pp, the wrapped agent would pass 69.44% instead of 67.42% — about 61.8 solves instead of 60, meaning roughly one extra task solved out of 89 for the same ~2x token spend. That is the pessimistic-but-consistent version of the claim. Nobody selling this skill prints it.</p>
<p>The power figure matters too. With n=89, a study has only about a 61% chance of detecting an effect the size of the one actually observed. This is a single run reporting a positive result, not a replicate — and single runs in this size class routinely produce point estimates that shrink on retest.</p>
<h2 id="is-the-evidence-asymmetric--did-the-treatment-arm-keep-its-receipts">Is the Evidence Asymmetric — Did the Treatment Arm Keep Its Receipts?</h2>
<p>No. The baseline retained all 89 per-task verdicts; the Autoprompt arm did not, and that asymmetry is more serious than a dead link.</p>
<p>The benchmark document links to a per-task verdict file that returns HTTP 404, and the top-level <code>benchmark/</code> directory it lived in is absent from the repository&rsquo;s main branch — verified by walking the full git tree on 2026-10-01, not inferred from the broken link. The document itself concedes that the Autoprompt arm &ldquo;cannot be rebuilt task by task&rdquo; because &ldquo;its original per-task map was not retained.&rdquo; The baseline&rsquo;s ledger exists; the treatment arm&rsquo;s does not.</p>
<p>This inverts the usual storage pattern. When a benchmark claims an improvement, the arm whose result is surprising is the arm whose raw output you most need to inspect — to check for task contamination, harness leakage, retries, or selective reporting. Here the ordinary arm is fully auditable and the extraordinary one is a summary table.</p>
<p>None of that is proof of error, and it should not be read as an accusation. It is a statement about what a reader can verify: the published counts can be re-analysed (the interval above is that analysis), but the underlying per-task result cannot be reproduced or falsified without rerunning the whole benchmark at your own expense. A result you cannot rebuild is a result you must trust, and trust is the one input a benchmark is supposed to remove.</p>
<h2 id="does-the-v2-release-invalidate-the-benchmark">Does the v2 Release Invalidate the Benchmark?</h2>
<p>For the claim as currently marketed, yes — the 45% is version 1 evidence attached to a version 2 product.</p>
<p>v2.0.0 shipped on 2026-09-09 with a new CLI activation path, eleven provider adapters, <code>path=</code> controls (auto, direct, light, roadmap) and new concurrency controls. The benchmark was run with OpenCode 1.18.7 in the v1 line; the current repository lists OpenCode 1.18.29 among its tested hosts. As of 2026-10-01 the <code>docs/benchmarks/</code> directory contains only the Terminal-Bench 2.1 file, a Codex canary note and a low-compute mechanism note — <strong>there is no v2 benchmark</strong>.</p>
<p>That distinction is not pedantry. A new activation path, new routing controls and eleven adapter changes alter the very execution path the benchmark measured. Every article published after 2026-09-09 that presents 45% as the current, measured performance of the software you are about to install is citing a superseded artifact. The honest framing is: &ldquo;v1 measured +14.61pp on one run; v2 is unmeasured.&rdquo;</p>
<h2 id="how-does-autoprompt-compare-to-other-agent-skills">How Does Autoprompt Compare to Other Agent Skills?</h2>
<p>Autoprompt sits in a small minority that works at all, and about in the middle of the curated-skill field — and the second half of that sentence is the part nobody likes.</p>
<p>Two independent benchmarks bracket this category. SWE-Skills-Bench evaluated 49 skills across roughly 565 task instances with deterministic pytest verification and found that <strong>39 of 49 skills produced zero pass-rate improvement</strong>, with an average gain of just +1.2%. Token overhead ranged from −78% to +451% while pass rates stayed flat, and three skills actively degraded performance by up to −10%. Anyone assuming a popular skill helps has roughly an 80% chance of being wrong about a random one.</p>
<p>SkillsBench is the kinder baseline: 87 tasks across 8 domains and 18 model-harness configurations, with matched no-skill and with-skill conditions. Curated skills lifted average pass rate from 33.9% to 50.5% — <strong>+16.6pp</strong>, a 25.5% normalised gain, with per-configuration gains from +4.1pp to +25.7pp.</p>
<p>Set Autoprompt&rsquo;s +14.61pp against those numbers:</p>
<table>
  <thead>
      <tr>
          <th>Evidence set</th>
          <th>Scope</th>
          <th>Effect on pass rate</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>SWE-Skills-Bench average</td>
          <td>49 skills, ~565 instances</td>
          <td>+1.2 pp</td>
      </tr>
      <tr>
          <td>SWE-Skills-Bench null result</td>
          <td>39 of 49 skills</td>
          <td>0 pp</td>
      </tr>
      <tr>
          <td>SkillsBench fleet average</td>
          <td>87 tasks, 18 configs</td>
          <td>+16.6 pp</td>
      </tr>
      <tr>
          <td><strong>Autoprompt, Terminal-Bench 2.1</strong></td>
          <td><strong>1 run, 89 tasks</strong></td>
          <td><strong>+14.61 pp</strong></td>
      </tr>
  </tbody>
</table>
<p>The headline that looks enormous in isolation is slightly below the fleet average of curated skills in context. Both statements are true, and together they give the honest verdict: Autoprompt is in the minority of skills that measurably help, and it is not an outlier among them. If your baseline assumption was &ldquo;agent skills do nothing&rdquo; (a 1.2% expectation), this is a large win. If your baseline was &ldquo;curated skills give about 16 points,&rdquo; this is ordinary.</p>
<h2 id="what-is-the-trap-next-door-for-self-generated-skills">What Is the Trap Next Door for Self-Generated Skills?</h2>
<p>The same SkillsBench work contains the finding most relevant to anyone tempted to copy the pattern: when agents were asked to author their own skills before starting a task, they performed <strong>worse than baseline, dropping 8.1 to 11.5 percentage points</strong>.</p>
<p>That is a direct warning for teams that read Autoprompt&rsquo;s architecture and conclude they can have their orchestrator auto-generate a project-specific checklist each run. The evidence says the opposite. Curated skills help; self-generated skills hurt, plausibly because an agent writing its own procedure optimises for plausibility rather than for the verification steps it is least likely to perform unaided. Autoprompt&rsquo;s own design leans the right way here — it is explicit, invoked by name (<code>/autoprompt</code>, <code>$autoprompt</code>, <code>/skill:autoprompt</code>) rather than authored on the fly — but the temptation to bolt on auto-generation is real, and the literature is not neutral about it.</p>
<h2 id="what-does-the-45-cost-in-tokens-and-time">What Does the 45% Cost in Tokens and Time?</h2>
<p>It costs roughly 3x wall-clock time and 2x tokens, and nobody has published the exchange rate between the two. We can compute an approximate one.</p>
<table>
  <thead>
      <tr>
          <th>Cost component</th>
          <th>Published value</th>
          <th>Evidence quality</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Wall-clock time</td>
          <td>~3x</td>
          <td>Project states it is a planning estimate from user reports</td>
      </tr>
      <tr>
          <td>Tokens</td>
          <td>~2x</td>
          <td>Same — timing and token logs were not retained</td>
      </tr>
      <tr>
          <td>Pass-rate gain</td>
          <td>+14.61 pp (CI +2.02 to +27.20)</td>
          <td>Measured once, n=89</td>
      </tr>
      <tr>
          <td>Implied efficiency</td>
          <td>≈1.5–1.6 pp of pass rate per extra 10% of tokens</td>
          <td>Derived by this review</td>
      </tr>
  </tbody>
</table>
<p>The project&rsquo;s own README labels the 3x/2x figures &ldquo;planning estimates based on user experience reports, not measured benchmark results,&rdquo; which is unusually candid — most projects would print them as fact. But it also means the one number a buyer needs does not exist. At an assumed 1.9x token multiplier, the +14.61pp gain comes at roughly 90% more tokens than baseline; spread across the 13 extra solves, that is about <strong>7% of a full baseline run&rsquo;s token budget per additional task solved</strong>. Or, in the form that survives a business case: you are buying roughly 1.5 percentage points of pass rate for every extra 10% of tokens you spend.</p>
<p>Whether that is worth it depends entirely on the alternative use of the same tokens — a second opinion from a stronger model, a broader test suite, or simply a retry with a different seed. Autoprompt&rsquo;s honest framing is a price, not a percentage, and the price is currently an estimate of an estimate.</p>
<h2 id="why-do-failure-reduction-tools-exist-at-all">Why Do Failure-Reduction Tools Exist at All?</h2>
<p>Because long-horizon reliability, not raw capability, is the binding constraint — and the arithmetic of chained steps explains why a verification layer can help at all.</p>
<p>METR&rsquo;s time-horizon research found that frontier agents succeed on nearly 100% of tasks a human would finish in under about four minutes, but on under 10% of tasks taking a human over roughly four hours, with the 50%-reliability horizon doubling approximately every seven months since 2019. Capability is not the problem at long horizons; the ability to keep 20 or 50 steps correct is.</p>
<p>Compounding makes that concrete. At 95% per-step reliability, a 20-step task succeeds about 36% of the time and a 50-step task about 8%.</p>
<table>
  <thead>
      <tr>
          <th>Per-step reliability</th>
          <th>20-step task success</th>
          <th>50-step task success</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>90%</td>
          <td>12.2%</td>
          <td>0.5%</td>
      </tr>
      <tr>
          <td>95%</td>
          <td>35.9%</td>
          <td>7.7%</td>
      </tr>
      <tr>
          <td>97%</td>
          <td>54.4%</td>
          <td>21.8%</td>
      </tr>
      <tr>
          <td>99%</td>
          <td>81.8%</td>
          <td>60.5%</td>
      </tr>
  </tbody>
</table>
<p>This is why a loop that adds independent verification is a plausible intervention rather than a gimmick: it targets step-level error, which is where the compounding lives. It is also why the <em>size</em> of Autoprompt&rsquo;s measured gain is plausible — moving an agent from roughly 67% to roughly 82% on a hard terminal benchmark is exactly the scale of improvement you would expect from catching a fraction of mid-trajectory mistakes, not from making the model smarter. And it is why the measurement is hard: the same compounding that makes the intervention valuable makes single-run estimates volatile.</p>
<h2 id="is-8202-actually-a-good-score">Is 82.02% Actually a Good Score?</h2>
<p>It is a mid-table score. Public Terminal-Bench 2.1 leaderboards put top model-harness combinations at roughly 87–91% on the same 89 tasks, so Autoprompt&rsquo;s 82.02% — with an 89-task, OpenCode 1.18.7, DeepSeek V4 Flash harness — sits below the frontier, not at it.</p>
<p>The comparability warning matters as much as the number. Terminal-Bench scores are not interchangeable across boards: the official board measures agent-plus-model combinations, Artificial Analysis tests models directly with labelled effort tiers, and vals.ai retests everything under a unified Terminus 2 harness. Citing &ldquo;82.02%&rdquo; without naming the harness and model invites a wrong read, in either direction — as a disappointing score when the harness was never frontier, or as a frontier score when the board is measuring something else. The safe sentence is the specific one: OpenCode 1.18.7 with DeepSeek V4 Flash and Autoprompt scored 82.02% on Terminal-Bench 2.1.</p>
<h2 id="why-do-the-stars-and-npm-downloads-point-the-other-way">Why Do the Stars and npm Downloads Point the Other Way?</h2>
<p>Adoption is decelerating while the headline number compounds, and the open-issue queue tells you where maintainer attention is going.</p>
<table>
  <thead>
      <tr>
          <th>Signal (as of 2026-10-01)</th>
          <th>Value</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GitHub stars</td>
          <td>1,298</td>
      </tr>
      <tr>
          <td>Forks</td>
          <td>88</td>
      </tr>
      <tr>
          <td>Open issues</td>
          <td>3 — all Windows activation blockers</td>
      </tr>
      <tr>
          <td>Created / last push</td>
          <td>2026-08-17 / 2026-09-28</td>
      </tr>
      <tr>
          <td>Licence</td>
          <td>MIT</td>
      </tr>
      <tr>
          <td>Stars/day, first 3 days</td>
          <td>≈146/day</td>
      </tr>
      <tr>
          <td>Stars/day, following ~38 days</td>
          <td>≈23/day</td>
      </tr>
      <tr>
          <td>npm <code>latest</code></td>
          <td>1.0.4 (published 2026-08-21)</td>
      </tr>
      <tr>
          <td>npm downloads, week to 2026-09-29</td>
          <td>32</td>
      </tr>
      <tr>
          <td>npm downloads, month to 2026-09-29</td>
          <td>144</td>
      </tr>
  </tbody>
</table>
<p>The star curve is a launch burst into a slow tail, not compounding growth: roughly 146 stars a day for three days, then about 23 a day for the next five weeks. npm tells the same story more sharply — the published package is still 1.0.4, two minor lines behind GitHub&rsquo;s v2.0.0, and therefore still carries nine providers and the old <code>mode=</code>/<code>max_subs=</code> interface. Weekly downloads are in the tens.</p>
<p>Meanwhile the three open issues are all Windows activation failures (#27 native Windows refused, #28 OpenCode <code>cmd-shim</code> EINVAL, #29 npm shims for native package bins), and the README&rsquo;s own requirement line says &ldquo;these versions passed Linux runs.&rdquo; The practical reading: the eleven-provider claim is a <strong>code-path</strong> claim, the <strong>verified execution surface is Linux</strong>, and if you are on Windows you are currently the maintainer&rsquo;s queue. If you want the tool to keep working, file issues rather than stars.</p>
<h2 id="how-portable-is-autoprompt-across-its-11-host-agents">How Portable Is Autoprompt Across Its 11 Host Agents?</h2>
<p>The adapters are real, and portability is the dimension no agent-skill framework has fully solved.</p>
<p>The tested-version table lists Claude Code 2.1.263, Codex 0.148.0, OpenCode 1.18.29, Kilo Code 7.5.15, VS Code 1.136.1, Prime Agent 0.7.2, Oh My Pi 18.1.14, DeepSeek Harness 0.1.2-rc.1, Reasonix 1.30.0, Hermes Agent 0.21.1 and Grok Build 1.0.13 — eleven hosts, each pinned to a version that passed on Linux. The v2 CLI unifies activation across them, which is genuine progress over the v1 invocation-per-host mess.</p>
<p>Two gaps remain. First, one independent review rates the project&rsquo;s portability &ldquo;portable with changes&rdquo; rather than portable, because the v2 CLI activation path is provider-specific — the unification is at the command surface, not at the execution semantics. Second, custom model routing does not exist on all hosts, so a team that needs to pin a specific model per role will find the coverage uneven. Independent taxonomy work reaches the same conclusion at the framework level: no agent-skill system currently covers specification, context, roles, execution, validation and portability simultaneously. Budget for a portability tax on any host other than Linux plus a mainstream agent.</p>
<h2 id="who-should-use-autoprompt-in-2026-and-who-should-not">Who Should Use Autoprompt in 2026, and Who Should Not?</h2>
<p>Buy it for hard, ambiguous tasks on a Linux workstation with an established host agent — and skip it everywhere else.</p>
<p><strong>A good fit if you:</strong></p>
<ul>
<li>Run Claude Code, Codex or OpenCode on Linux or macOS and already have a working test suite.</li>
<li>Spend your time on multi-step, underspecified tasks where the agent &ldquo;gets stuck mid-task&rdquo; rather than failing instantly.</li>
<li>Value the verification structure more than the speed — the independent-check discipline is the durable part of the design, and it is transferable even if you later drop the tool.</li>
<li>Need evidence you can point at: the +14.61pp gain is statistically significant, and you now know its interval.</li>
</ul>
<p><strong>A poor fit if you:</strong></p>
<ul>
<li>Do very small tasks. A 2x token multiplier on a one-file change is pure loss, and the project itself notes gains may vary significantly between small and large tasks.</li>
<li>Work primarily on Windows today. All three open issues are Windows activation failures.</li>
<li>Need custom model routing on every host.</li>
<li>Have no local terminal runtime for the agent to execute in — without that, there is nothing for the verification loop to verify.</li>
<li>Are choosing between Autoprompt and simply spending the same 2x tokens on a stronger model or a larger test suite. That comparison has never been published, and it is the one that decides most budgets.</li>
</ul>
<p>Read against the third-party score available — FollowAgents rates the project 69/100, &ldquo;use with care,&rdquo; with praise focused specifically on disclosure quality and a caveat that the review is stale on provider count and version — the fair summary is that the engineering and the honesty are both above average while the evidence remains one run.</p>
<h2 id="how-do-you-reproduce-or-reject-the-claim-yourself">How Do You Reproduce or Reject the Claim Yourself?</h2>
<p>The claim is falsifiable, and the cheapest path is a paired test on your own backlog.</p>
<ol>
<li><strong>Fix the harness.</strong> Pin one host agent and one model. Mixed hosts invalidate every comparison, as the three-board divergence on Terminal-Bench 2.1 shows.</li>
<li><strong>Select tasks with headroom.</strong> Pick 30–50 tasks you already fail. Testing on tasks you pass measures nothing.</li>
<li><strong>Run both arms on the same tasks.</strong> Baseline run, then an Autoprompt run, same commit, same environment, ideally same day to reduce drift.</li>
<li><strong>Score deterministically.</strong> Use your test suite or an external grader, not agent self-assessment. This is the step most teams skip, and it is the step that matters — an agent judging its own work is precisely the failure mode the skill exists to prevent.</li>
<li><strong>Compute the interval, not just the difference.</strong> At 30 paired tasks you have almost no power; a 15-point difference on 30 tasks will not be distinguishable from noise. If your observed difference is small, treat the result as inconclusive rather than negative.</li>
<li><strong>Log tokens and wall clock.</strong> The project could not, and that is why its cost figure is an estimate. Your run should not repeat that mistake.</li>
<li><strong>Run <code>autoprompt doctor --strict</code></strong> before measuring, so a partial install is not scored as a skill failure.</li>
</ol>
<p>Expect a real but unglamorous outcome. The published effect is +14.61pp with a lower bound of +2.02pp; on a 40-task set, a plausible result is three to six extra solves at roughly double the tokens.</p>
<h2 id="what-is-the-verdict-on-the-autoprompt-skill">What Is the Verdict on the Autoprompt Skill?</h2>
<p>It works, the number is real, and the number is smaller and less certain than the headline implies.</p>
<p>The defensible statement of the evidence is this: on a single 89-task Terminal-Bench 2.1 run, wrapping OpenCode 1.18.7 with Autoprompt raised pass rate from 67.42% to 82.02%, a +14.61pp gain (95% CI +2.02pp to +27.20pp, p=0.025) that reduces the failure rate 44.8% — 1.81x fewer failures — at an estimated 3x time and 2x tokens, with no v2 benchmark yet published.</p>
<p>Everything past the comma is why this article exists. Most surfaces that rank for &ldquo;autoprompt skill&rdquo; have already lost it: catalog pages still publish nine providers, v1.0.4 and 940 stars while the repository says eleven providers, v2.0.0 and 1,298 stars. The 45% survived intact while every verifiable detail around it decayed — which is the normal behaviour of a good headline and the reason it deserves an error bar. Use the skill for hard tasks, budget the tokens honestly, and quote the interval.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Does the Autoprompt skill actually work?</strong>
The evidence says yes, within limits. Autoprompt&rsquo;s own Terminal-Bench 2.1 run moved OpenCode 1.18.7 from 60/89 to 73/89 solves (+14.61pp, p=0.025), and independent category benchmarks show most agent skills produce no measurable gain at all, so a skill that clears the bar is a minority case. However, the result is a single 89-task run, the treatment arm&rsquo;s per-task ledger was not retained, and the 95% confidence interval spans +2.02pp to +27.20pp. Treat it as a real but imprecisely measured improvement, not a guarantee.</p>
<p><strong>Where does the 45% figure come from and is it accurate?</strong>
It comes from failures falling from 29 to 16 on 89 Terminal-Bench 2.1 tasks. That is a 44.8% reduction in the failure rate, so 45% is accurate as a description of that specific metric. It is not a 45% improvement in capability: the pass rate rose 14.61 points (67.42% to 82.02%), the relative pass-rate gain is 21.67%, and in odds terms the wrapped agent fails 1.81x less often. Any source describing it as &ldquo;45% better&rdquo; has changed the meaning of the number.</p>
<p><strong>What is the confidence interval on the Autoprompt benchmark, and why does it matter?</strong>
The +14.61pp difference corresponds to a 95% confidence interval of roughly +2.02pp to +27.20pp at n=89 per arm (two-proportion z=2.24; Fisher exact two-sided p=0.038). No competitor publishes this. It matters because the interval is wide: the same data is consistent with a benefit as large as 27 points and as small as 2 points, and at 2 points the tool would win about one extra task out of 89. With n=89, statistical power to detect the observed effect is only about 61%, so this should be read as a positive single run, not a settled effect size.</p>
<p><strong>Does Autoprompt work on Windows, and is the npm package current?</strong>
Windows is currently the hard gate: all three open issues on the repository as of 2026-10-01 are Windows activation failures (#27 native Windows refused, #28 OpenCode cmd-shim EINVAL, #29 npm shims for native package bins), and the project states that its tested versions &ldquo;passed Linux runs.&rdquo; Separately, npm&rsquo;s <code>latest</code> tag is still 1.0.4 (published 2026-08-21), two minor lines behind the GitHub v2.0.0 release of 2026-09-09, so the npm package still ships nine providers and the old <code>mode=</code>/<code>max_subs=</code> interface. Install from the v2.0.0 release if you need the CLI and the eleven adapters.</p>
<p><strong>How much does Autoprompt cost in tokens and time, and is it worth it?</strong>
The project estimates roughly 3x wall-clock time and 2x tokens, and explicitly labels both as planning estimates from user experience reports because timing and token logs were not retained. Pair that with the +14.61pp gain and you get approximately 1.5 to 1.6 percentage points of pass rate per extra 10% of tokens — a real but modest exchange rate. It is worth it on hard, multi-step tasks where your agent tends to get stuck and you have a deterministic test suite to verify against. It is a clear loss on small single-file changes, where the multiplier buys nothing.</p>
]]></content:encoded></item></channel></rss>