<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Coding Agent Orchestration on RockB</title><link>https://baeseokjae.github.io/tags/coding-agent-orchestration/</link><description>Recent content in Coding Agent Orchestration on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 28 Sep 2026 12:55:34 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/coding-agent-orchestration/index.xml" rel="self" type="application/rss+xml"/><item><title>Autoprompt Skill Review: The Coding Agent Prompt Skill That Cuts Failures 45%</title><link>https://baeseokjae.github.io/posts/autoprompt-skill-coding-agent-failures/</link><pubDate>Mon, 28 Sep 2026 12:55:34 +0000</pubDate><guid>https://baeseokjae.github.io/posts/autoprompt-skill-coding-agent-failures/</guid><description>Autoprompt claims 45% fewer coding-agent failures. We trace the number to its run, verify the repo state on 2026-09-28, and price the token cost.</description><content:encoded><![CDATA[<p>Autoprompt is a coding agent prompt skill that injects an orchestration procedure — plan, delegate, implement, verify — into 11 host agents, including Claude Code, Codex, and OpenCode. Its headline claim of 45% fewer failures comes from a single 89-task Terminal-Bench 2.1 run where failures fell from 29 to 16. That result is real arithmetic but version 1 evidence, and it costs roughly 3x wall-clock time and 2x tokens.</p>
<h2 id="what-the-45-figure-actually-measures">What the 45% Figure Actually Measures</h2>
<p>The number everyone quotes is not a percentage of tasks solved and it is not a percentage reduction in tokens. It is a reduction in failures on one benchmark, measured once.</p>
<p>The run is documented as Terminal-Bench 2.1 with 89 tasks, OpenCode 1.18.7 as the harness, and DeepSeek V4 Flash as the model. Baseline solves were 60 of 89. With Autoprompt active, solves rose to 73 of 89. Failures went from 29 to 16, which is a 44.8% reduction — rounded to 45% in every headline since.</p>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>Baseline</th>
          <th>With Autoprompt</th>
          <th>Delta</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tasks solved</td>
          <td>60 / 89</td>
          <td>73 / 89</td>
          <td>+13 tasks</td>
      </tr>
      <tr>
          <td>Solve rate</td>
          <td>67.42%</td>
          <td>82.02%</td>
          <td>+14.61 pp</td>
      </tr>
      <tr>
          <td>Failures</td>
          <td>29</td>
          <td>16</td>
          <td>-13 (-44.8%)</td>
      </tr>
      <tr>
          <td>Wall-clock time</td>
          <td>1x</td>
          <td>~3x (estimated)</td>
          <td>not measured</td>
      </tr>
      <tr>
          <td>Token usage</td>
          <td>1x</td>
          <td>~2x (estimated)</td>
          <td>not measured</td>
      </tr>
  </tbody>
</table>
<p>Two things make that table worth reading slowly. First, the denominator: 13 tasks out of 89 is a meaningful but bounded result, and 44.8% is a property of the failure slice (13 of 29), not of the workload. The same result is also described in the project&rsquo;s own README as &ldquo;about 2x fewer mistakes,&rdquo; which is the identical measurement expressed against the smaller surviving set. Neither framing is dishonest; the choice of denominator is what makes a 13-task improvement look like a near-halving.</p>
<p>Second, and more important: cost was not instrumented. The project discloses the ~3x time and ~2x token trade-off as &ldquo;planning estimates based on user experience reports, not measured benchmark results,&rdquo; and explains that timing and token logs were not retained. So the accuracy half of the claim is a measured benchmark; the cost half is not a measurement at all. If you are budgeting an unattended run, treat 3x as a floor and verify it on your own task distribution before trusting it in a CI budget.</p>
<h2 id="autoprompt-v2-what-changed-since-the-original-review">Autoprompt v2: What Changed Since the Original Review</h2>
<p>The repository has moved past the benchmark. Autoprompt shipped v2.0.0 on 2026-09-09, and the README now labels the 45% result as &ldquo;version 1 benchmarks,&rdquo; with a note that version 2 benchmarks will follow. No v2 efficacy number exists as of 2026-09-28. That matters for anyone citing the headline today: the marketing line is unchanged, but the artifact it refers to is two releases old.</p>
<p>What v2 actually changed is scope and control, not the accuracy thesis:</p>
<table>
  <thead>
      <tr>
          <th>Area</th>
          <th>v1</th>
          <th>v2.0.0</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Interface</td>
          <td>Slash-command modes</td>
          <td>CLI: <code>autoprompt activate PROVIDER --target /abs/project -- '&lt;goal&gt;'</code></td>
      </tr>
      <tr>
          <td>Providers</td>
          <td>Codex, OpenCode, Prime Agent</td>
          <td>11, adding Hermes Agent and Grok Build</td>
      </tr>
      <tr>
          <td>Controls</td>
          <td>Mode presets</td>
          <td>Path selection plus model, effort, and concurrency flags</td>
      </tr>
      <tr>
          <td>Subagent bounds</td>
          <td>Not configurable</td>
          <td><code>--concurrency tokensaver|wide|custom --max-subs N</code></td>
      </tr>
      <tr>
          <td>Installer lifecycle</td>
          <td>Manual</td>
          <td>Managed install/update lifecycle</td>
      </tr>
  </tbody>
</table>
<p>The v2 workflow is more explicit about where control sits. Instead of a mode switch, you select a work path — <code>path=auto|direct|light|roadmap</code> — and a concurrency profile. The default is <code>tokensaver</code>, which caps fan-out at six subagents. The documentation is clear that choosing a path does not waive authorization, capability, budget, or verification requirements, and that invalid combinations error out rather than silently falling back to a different route.</p>
<h2 id="the-evidence-boundary-moved-and-it-got-thinner">The Evidence Boundary Moved, and It Got Thinner</h2>
<p>The single most important fact for a 2026 buyer is not in the README: the benchmark evidence is no longer fully checkable.</p>
<p>The per-task verdict file that the benchmark documentation points to for all 89 retained verdicts now returns HTTP 404. The <code>benchmark/</code> directory that once held the raw artifacts is absent from the main branch — verified by walking the recursive git tree listing of 1,381 paths, not inferred from a link. The methodology summary survives in <code>docs/benchmarks/terminal-bench-2.1.md</code>, but the per-task data a skeptical reader would use to audit the run does not.</p>
<p>That is a reproducibility regression, and it is the correct thing to hold against this project, because prompt-skill vendors have an obvious incentive to publish the summary and not the ledger. A benchmark you cannot audit is a marketing claim with a methodology section. Independent verification does not become impossible — you can rerun Terminal-Bench yourself — but the cheapest possible audit path (read the verdicts) is closed, and that cost falls on the buyer.</p>
<p>The v2.0.0 release notes show a similar honesty pattern in miniature. They concede that &ldquo;full live acceptance of the final Codex artifact is not established,&rdquo; and that Windows and macOS builds are installer kits with native runtime verification remaining Linux-only. That is unusually candid disclosure for a release announcement, and it is the reason this review reads the project as careful rather than reckless: it publishes caveats that hurt it. It simply has not published version 2 evidence yet.</p>
<h2 id="eleven-providers-one-verified-platform">Eleven Providers, One Verified Platform</h2>
<p>Provider count is the feature most third-party articles get stale. Autoprompt now claims 11 supported hosts. The v2.0.0 work extended the Codex workflow across all of them and added Hermes Agent and Grok Build to a list that previously stopped at Codex, OpenCode, and Prime Agent.</p>
<p>The platform support story is narrower than the provider story. All three open items on the repository concern Windows:</p>
<ul>
<li>Issue #27 — activation is refused on Windows with <code>PROVIDER_UNSUPPORTED</code> because all 10 reviewed local records are <code>linux/x64</code> and <code>trusted-public-keys.json</code> lists no Windows keys.</li>
<li>Issue #28 — OpenCode activation fails on the npm <code>cmd</code> shim.</li>
<li>PR #29 — an in-flight fix for Windows npm shims for native package bins.</li>
</ul>
<p>Read together with the release notes&rsquo; Linux-only verification line, the practical rule is simple: the 11-provider claim is a code-path claim, and the verified execution surface is Linux. If your team codes on Windows, this is not yet a drop-in tool, and the open issues say so more precisely than any review can.</p>
<h2 id="paths-and-run-controls-where-v2-spends-your-tokens">Paths and Run Controls: Where v2 Spends Your Tokens</h2>
<p>Cost in this category is not a fixed multiplier; it is a budget you set. The mechanism is straightforward and is the same one the independent literature describes: every phase is another model call over the same code, the judge is a separate model call with its own context, and parallel lanes re-read the same files. Fan-out is therefore the dominant cost lever, more than the model choice per phase.</p>
<p>The v2 controls map to that reality directly:</p>
<ul>
<li><code>--concurrency tokensaver</code> — at most six subagents, the default. Sensible for a laptop or a metered API key.</li>
<li><code>--concurrency wide</code> — maximum parallel lanes. Fast, and the most expensive way to run it.</li>
<li><code>--concurrency custom --max-subs N</code> — the setting most teams should actually use, because it lets you align fan-out with the size of the change rather than the optimism of the moment.</li>
<li><code>path=direct</code> — minimal routing for a small, well-specified change.</li>
<li><code>path=light</code> / <code>roadmap</code> — progressive planning depth for larger or ambiguity-heavy work.</li>
</ul>
<p>One operational caution worth repeating from independent walkthroughs: six lanes can touch deployment credentials, so run unattended sessions in a disposable VM or a container with scoped keys rather than on your working machine. And there is a real quality floor below which orchestration is pure overhead — if the task has nothing to execute or test, the verification phase produces one more opinion at full token price.</p>
<p>The default <code>tokensaver</code> profile is a reasonable acknowledgment of that floor. It is also the setting to keep when you are running agents on a VPS budget, where a wide fan-out can turn a ten-minute change into a twenty-dollar one.</p>
<h2 id="what-independent-skill-benchmarks-say-about-this-category">What Independent Skill Benchmarks Say About This Category</h2>
<p>Autoprompt is one instance of a category that has been measured independently, and the category averages are not kind. Any review that quotes 45% without quoting the surrounding literature is selling, not informing.</p>
<p>The closest comparable study, SWE-Skills-Bench, ran 565 task instances across 49 public SWE skills with GitHub repositories pinned at fixed commits and execution-based verification. Thirty-nine of the 49 skills produced zero pass-rate improvement. The average gain was +1.2%. Token overhead ranged from -78% to +451% while pass rates stayed flat. Seven specialized skills gained up to +30%, and three actively degraded performance by up to -10% through version-mismatched guidance.</p>
<p>SkillsBench, testing 87 tasks across 18 model-harness configurations with matched no-skill controls, found that curated skills raise the average pass rate from 33.9% to 50.5% — a +16.6 percentage point gain — but with gains ranging from +4.1 to +25.7 points, and 13 of 87 tasks showing negative deltas. Its most useful structural finding is that skill value is a property of the specific stack, not of skills in general. Compact focused skills beat exhaustive bundles by a wide margin, and loading four or more skills together yielded only +10.1 points.</p>
<p>A third line of work explains the failure mode that matters most for a verification-heavy tool. The paper &ldquo;Agent Skills Can Be Harmful&rdquo; confirmed 307 skill-induced failures across those same benchmarks: 125 functional and 182 efficiency regressions. Efficiency damage was dominated by what the authors call Excessive Procedure at 62.6% — not prompt length — with excessive verification (67 cases) and heavy implementation pipelines (30 cases) as the largest drivers. Functional failures rarely came from irrelevant skills; seemingly relevant skills caused wrong or omitted implementation elements in 68.8% of cases. That is Autoprompt&rsquo;s exact design surface: it turns verification checklists into mandatory work, which is simultaneously the mechanism of its accuracy claim and the largest documented source of token regression in the literature.</p>
<p>Two further results bound the realistic ceiling. Under progressively realistic conditions, where the agent must retrieve skills itself instead of receiving a hand-picked one, skill benefits decay toward no-skill baselines — and query-specific refinement recovers performance, lifting Claude Opus 4.6 from 57.7% to 65.5% on Terminal-Bench 2.0. A separate 500-skill, roughly 38,000-trajectory evaluation found relevant skills lifting scores by 5.5 to 22 points, with open-weight GLM 5.1 plus a skill scoring 91.1 at about $0.89 per scenario against Opus 4.8 at 92.7 for $3.26.</p>
<p>That last number is the under-discussed lever. If you pair a cheaper model with a strong procedure, you can approach frontier quality at a third to a quarter of the cost — which changes the economics of Autoprompt&rsquo;s ~2x token multiplier far more than shaving one subagent will.</p>
<p>For context on why this market exists at all: Anthropic&rsquo;s 2026 survey of more than 500 technical leaders found about 90% of organizations use AI for coding, 86% deploy coding agents for production code, and 57% run agents on multi-stage workflows. Time savings were reported at 59% across code generation, review and testing, and research alike. Integration with existing systems (46%) is the top adoption blocker, ahead of data quality (42%) and cost (43%).</p>
<h2 id="the-npm-versus-github-distribution-gap">The npm Versus GitHub Distribution Gap</h2>
<p>If you install this tool with the command most articles still publish, you get the old version.</p>
<table>
  <thead>
      <tr>
          <th>Channel</th>
          <th>Version</th>
          <th>Published</th>
          <th>Notes</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GitHub</td>
          <td>v2.0.0</td>
          <td>2026-09-09</td>
          <td>11 providers, CLI, path and concurrency controls</td>
      </tr>
      <tr>
          <td>npm (<code>autoprompt-skill</code>)</td>
          <td>1.0.4</td>
          <td>2026-08-21</td>
          <td>Zero runtime dependencies, two majors behind</td>
      </tr>
  </tbody>
</table>
<p>npm weekly downloads were 38 for the week of 2026-09-21 to 2026-09-27, against third-party trackers still publishing the older figure of 535 per week — a 14x discrepancy that mostly reflects how slowly catalog data propagates.</p>
<p>Repository state on the date this review was verified, 2026-09-28: 1,295 stars, 88 forks, 3 open issues, MIT license, 44 commits, created 2026-08-17, latest release v2.0.0 on 2026-09-09. Star velocity at launch was cited at 437 stars in three days, roughly 153 per day.</p>
<p>The version divergence explains most of the factual disagreements between reviews of this tool. Published comparisons disagree on provider count (9 versus 11) and star count (590 versus 1,300) because they were written against different points in a fast-moving repository. Some also repeat unattributed anecdotes — debugging-time reductions at an investment firm, e-commerce reliability gains — with no source at all. The 45% figure is exactly the kind of claim that survives this propagation intact while every verifiable detail around it goes stale; the version data in this review is dated deliberately so you can see how far it has drifted since.</p>
<h2 id="where-autoprompt-earns-its-cost-and-where-it-does-not">Where Autoprompt Earns Its Cost, and Where It Does Not</h2>
<p>The category literature says a minority of skills genuinely help and the winners are concrete workflows, not general best-practice documents. Autoprompt is positioned in the subcategory that Anthropic reports as delivering the largest quality improvement: verification-type skills. That is the strongest structural argument in its favor, and it does not depend on the 45% at all.</p>
<p>The strongest practical argument is role separation. A reviewer that runs the code and can disagree on evidence is a different thing from a model approving its own output, and that architectural difference is what makes the benchmark result plausible — plan, implement, and verify as distinct calls with distinct context. Autonomy here is bounded by design: the workflow stops for choices that change the result or actions that require authority. If you were sold &ldquo;fully autonomous coding,&rdquo; that is not what this is, and the stopping behavior is a feature.</p>
<p>Where it does not earn its cost:</p>
<ul>
<li><strong>Small, well-specified changes.</strong> The project&rsquo;s own caveat is that small tasks &ldquo;may differ significantly,&rdquo; and the efficiency-regression literature is dominated by exactly this pattern. Use <code>path=direct</code> or skip it.</li>
<li><strong>Tasks with nothing to verify.</strong> No executable check means the verification phase buys an opinion, not evidence.</li>
<li><strong>Windows development teams.</strong> The verified surface is Linux; three open items say so.</li>
<li><strong>Workflows needing an auditable number.</strong> The per-task evidence is 404 and v2 has no benchmark. You are trusting a methodology document.</li>
<li><strong>Multi-skill setups.</strong> If you already load four or more skills, the incremental value measured by SkillsBench drops sharply, and skill metadata competes for context — Claude Code caps skill metadata at roughly 1% of the context window (~2,000 tokens, about 15-25 skills on a 200K model) before descriptions get truncated.</li>
</ul>
<h2 id="verdict-who-should-use-it-in-2026">Verdict: Who Should Use It in 2026</h2>
<p>Autoprompt is a well-engineered orchestration skill with a real, single-run accuracy result, unusually candid disclosure about its own gaps, and a version 2 that has not been benchmarked. The 45% is defensible only with its denominator attached: 13 tasks out of 89, on one harness, with one model, costing an unmeasured ~3x time and ~2x tokens.</p>
<p>Adopt it if you run long agentic coding sessions on Linux, you already have test suites the verifier can execute, and you want structural role separation between implementation and review. Start with the default <code>tokensaver</code> concurrency, pin <code>--max-subs</code> to something honest for your change size, and measure your own token delta in week one.</p>
<p>Do not adopt it on the strength of the headline. Do not cite it as an established result without noting that the evidence directory is gone and v2 is untested. And if you need independent confirmation before standardizing, rerun Terminal-Bench on your own repositories — which is what the project&rsquo;s own disclosure effectively invites you to do.</p>
<p>The honest one-line summary: this is the exception the skeptical literature allows for, priced at 3x the wall clock, sold with a number that is two releases old.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Does Autoprompt really reduce coding agent failures by 45%?</strong></p>
<p>On one measured run, yes: failures fell from 29 to 16 across 89 Terminal-Bench 2.1 tasks using OpenCode 1.18.7 and DeepSeek V4 Flash, a 44.8% reduction. That is a single benchmark on a single harness and model combination, and the project itself labels it a version 1 result. The same 13-task improvement is also described as &ldquo;about 2x fewer mistakes.&rdquo;</p>
<p><strong>Is the 45% benchmark independently verifiable?</strong></p>
<p>Not from the project&rsquo;s own artifacts. The per-task verdict file the benchmark docs link to returns HTTP 404, and the <code>benchmark/</code> directory is absent from the main branch among 1,381 paths in the recursive tree. The methodology document remains, so you can read how the run was done, but the raw verdicts needed to audit it are no longer published. Independent verification means rerunning the benchmark yourself.</p>
<p><strong>How much more expensive is Autoprompt than a plain coding agent?</strong></p>
<p>The disclosed figure is roughly 3x wall-clock time and 2x tokens — and the project states plainly that these are planning estimates from user experience reports, not measured benchmark results, because timing and token logs were not retained. Actual overhead depends mostly on fan-out, which you control via the <code>tokensaver</code> (six subagents maximum), <code>wide</code>, or <code>custom --max-subs N</code> concurrency profiles.</p>
<p><strong>Which coding agents does Autoprompt support, and on which platforms?</strong></p>
<p>v2.0.0 claims 11 providers, extending the original Codex, OpenCode, and Prime Agent list and adding Hermes Agent and Grok Build. Platform verification is narrower: the release notes concede native runtime verification remains Linux-only, and the three open repository items all concern Windows — <code>PROVIDER_UNSUPPORTED</code> activation failures caused by linux-only trusted keys, an OpenCode activation failure on the npm <code>cmd</code> shim, and an in-flight PR fixing Windows npm shims.</p>
<p><strong>Should I install Autoprompt from npm or GitHub?</strong></p>
<p>GitHub. npm still ships 1.0.4, published 2026-08-21, while the repository is at v2.0.0 from 2026-09-09 — a two-major-version gap that means the npm package lacks the CLI, the path controls, and the expanded provider support. npm weekly downloads were 38 for the week ending 2026-09-27, so the stale channel is also the less-used one.</p>
]]></content:encoded></item></channel></rss>