<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Agent Skill Token Overhead on RockB</title><link>https://baeseokjae.github.io/tags/agent-skill-token-overhead/</link><description>Recent content in Agent Skill Token Overhead on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:51:12 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/agent-skill-token-overhead/index.xml" rel="self" type="application/rss+xml"/><item><title>Autoprompt Skill Review: Cutting Agentic Failures by 45%, Measured</title><link>https://baeseokjae.github.io/posts/autoprompt-coding-agent-skill/</link><pubDate>Thu, 01 Oct 2026 00:51:12 +0000</pubDate><guid>https://baeseokjae.github.io/posts/autoprompt-coding-agent-skill/</guid><description>Autoprompt&amp;#39;s 45% claim recomputed: +14.61pp gain, 95% CI +2.02pp to +27.20pp, and roughly 1.5 pass-rate points per extra 10% of tokens.</description><content:encoded><![CDATA[<p>Autoprompt is an open-source agent skill that wraps a coding agent in a plan, build, check, verify and finish loop. Its headline result — 45% fewer failures — is 29 down to 16 failures on 89 Terminal-Bench 2.1 tasks, a +14.61 percentage-point pass-rate gain whose 95% confidence interval is +2.02pp to +27.20pp. The direction is real; the precision is not.</p>
<p>This review is deliberately narrow. We recap the product in one paragraph and then spend the rest of the article on the three things no catalog page does: compute the error bar, price the trade-off, and test the claim against category-level evidence. If you want the feature tour, see our <a href="https://baeseokjae.github.io/posts/autoprompt-skill-coding-agent-2026/">earlier Autoprompt walkthrough</a>.</p>
<h2 id="what-is-the-autoprompt-skill-in-one-paragraph">What Is the Autoprompt Skill in One Paragraph?</h2>
<p>Autoprompt is a coordination layer, not another coding agent. It installs into a host tool — Claude Code, Codex, OpenCode and eight others — and, when explicitly invoked, runs a fixed loop: choose a route, plan, build, check, independently verify, then finish. Roles are split across levels L0 to L4 so that no single agent plans a change, approves it and certifies its own verification. It requires Node.js 20+, Python 3.11+ with PyYAML and Bash 4.3+, and v2.0.0 ships a single CLI that activates the skill per provider. That is the whole product; everything below is about whether its number survives scrutiny.</p>
<h2 id="what-does-45-fewer-failures-actually-measure">What Does &ldquo;45% Fewer Failures&rdquo; Actually Measure?</h2>
<p>It measures one metric on one benchmark, once. Baseline OpenCode 1.18.7 solved 60 of 89 Terminal-Bench 2.1 tasks. With Autoprompt active, the same harness on the same model solved 73 of 89. Failures fell from 29 to 16.</p>
<p>That is a defensible statement, and it is also the limit of what the published evidence supports. The figure is a reduction in the failure <strong>rate</strong> on a single 89-task run, not a reduction in defects you will see in your repository, and not a comparison against any other tool.</p>
<table>
  <thead>
      <tr>
          <th>Read the same result a different way</th>
          <th>Baseline</th>
          <th>With Autoprompt</th>
          <th>Change</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tasks solved</td>
          <td>60 / 89</td>
          <td>73 / 89</td>
          <td>+13 tasks</td>
      </tr>
      <tr>
          <td>Pass rate</td>
          <td>67.42%</td>
          <td>82.02%</td>
          <td>+14.61 pp</td>
      </tr>
      <tr>
          <td>Failure rate</td>
          <td>32.58%</td>
          <td>17.98%</td>
          <td>−14.61 pp</td>
      </tr>
      <tr>
          <td>Failures (count)</td>
          <td>29</td>
          <td>16</td>
          <td>−44.8% (~45%)</td>
      </tr>
      <tr>
          <td>Relative pass-rate gain</td>
          <td>—</td>
          <td>—</td>
          <td>+21.67%</td>
      </tr>
      <tr>
          <td>Failures per solved task</td>
          <td>0.483</td>
          <td>0.219</td>
          <td>−54.7%</td>
      </tr>
  </tbody>
</table>
<p>Every row is the same data. The &ldquo;45%&rdquo; headline is the most flattering row that is still arithmetically true, which is exactly why it propagated and the rest did not.</p>
<h2 id="why-is-a-45-reduction-really-a-181x-claim">Why Is a 45% Reduction Really a 1.81x Claim?</h2>
<p>Because a percentage reduction in failures is not a percentage improvement in capability, and the difference is large enough to change how you budget for it.</p>
<p>Two useful translations exist. In odds terms, 29 failures becoming 16 is <strong>1.81x fewer failures</strong> — the treatment arm fails about 55% as often. In rate terms, 32.58% becoming 17.98% is a <strong>44.8% reduction in the failure rate</strong>, rounded to 45%. Both are honest. &ldquo;45% better agent&rdquo; is not, and it is the phrase most readers hear.</p>
<p>For a working developer the odds reading is the operative one: on a hard task where the plain agent has a 33% chance of failing, the wrapped agent has roughly an 18% chance, so you still expect to repair one task in six. That is a meaningful improvement and it is not a transformation.</p>
<h2 id="what-is-the-confidence-interval-on-the-45-claim">What Is the Confidence Interval on the 45% Claim?</h2>
<p>The error bar nobody publishes is +2.02pp to +27.20pp at 95% confidence. That is the single most useful number in this review.</p>
<table>
  <thead>
      <tr>
          <th>Statistic (60/89 vs 73/89, n = 89 per arm)</th>
          <th>Value</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Absolute difference in pass rate</td>
          <td>+14.61 pp</td>
      </tr>
      <tr>
          <td>95% CI, two-proportion interval</td>
          <td><strong>+2.02 pp to +27.20 pp</strong></td>
      </tr>
      <tr>
          <td>Two-proportion z</td>
          <td>2.24</td>
      </tr>
      <tr>
          <td>Two-sided p (normal)</td>
          <td>0.025</td>
      </tr>
      <tr>
          <td>Fisher exact, one-sided p</td>
          <td>0.019</td>
      </tr>
      <tr>
          <td>Fisher exact, two-sided p</td>
          <td>0.038</td>
      </tr>
      <tr>
          <td>Statistical power to detect this effect at n=89</td>
          <td>≈61%</td>
      </tr>
      <tr>
          <td>n per arm to detect +5pp at 80% power</td>
          <td>≈1,319</td>
      </tr>
      <tr>
          <td>n per arm to detect +3pp at 80% power</td>
          <td>≈3,735</td>
      </tr>
  </tbody>
</table>
<p>Read the interval, not the point estimate. The result is statistically significant — the lower bound clears zero — but the CI spans from &ldquo;a marginal benefit you might not notice&rdquo; to &ldquo;a near doubling of the failure reduction you were promised.&rdquo; A 14.61-point change on 89 tasks cannot distinguish those two worlds, and it cannot rule out an effect as small as two points.</p>
<p>The lower bound is worth translating. At +2.02pp, the wrapped agent would pass 69.44% instead of 67.42% — about 61.8 solves instead of 60, meaning roughly one extra task solved out of 89 for the same ~2x token spend. That is the pessimistic-but-consistent version of the claim. Nobody selling this skill prints it.</p>
<p>The power figure matters too. With n=89, a study has only about a 61% chance of detecting an effect the size of the one actually observed. This is a single run reporting a positive result, not a replicate — and single runs in this size class routinely produce point estimates that shrink on retest.</p>
<h2 id="is-the-evidence-asymmetric--did-the-treatment-arm-keep-its-receipts">Is the Evidence Asymmetric — Did the Treatment Arm Keep Its Receipts?</h2>
<p>No. The baseline retained all 89 per-task verdicts; the Autoprompt arm did not, and that asymmetry is more serious than a dead link.</p>
<p>The benchmark document links to a per-task verdict file that returns HTTP 404, and the top-level <code>benchmark/</code> directory it lived in is absent from the repository&rsquo;s main branch — verified by walking the full git tree on 2026-10-01, not inferred from the broken link. The document itself concedes that the Autoprompt arm &ldquo;cannot be rebuilt task by task&rdquo; because &ldquo;its original per-task map was not retained.&rdquo; The baseline&rsquo;s ledger exists; the treatment arm&rsquo;s does not.</p>
<p>This inverts the usual storage pattern. When a benchmark claims an improvement, the arm whose result is surprising is the arm whose raw output you most need to inspect — to check for task contamination, harness leakage, retries, or selective reporting. Here the ordinary arm is fully auditable and the extraordinary one is a summary table.</p>
<p>None of that is proof of error, and it should not be read as an accusation. It is a statement about what a reader can verify: the published counts can be re-analysed (the interval above is that analysis), but the underlying per-task result cannot be reproduced or falsified without rerunning the whole benchmark at your own expense. A result you cannot rebuild is a result you must trust, and trust is the one input a benchmark is supposed to remove.</p>
<h2 id="does-the-v2-release-invalidate-the-benchmark">Does the v2 Release Invalidate the Benchmark?</h2>
<p>For the claim as currently marketed, yes — the 45% is version 1 evidence attached to a version 2 product.</p>
<p>v2.0.0 shipped on 2026-09-09 with a new CLI activation path, eleven provider adapters, <code>path=</code> controls (auto, direct, light, roadmap) and new concurrency controls. The benchmark was run with OpenCode 1.18.7 in the v1 line; the current repository lists OpenCode 1.18.29 among its tested hosts. As of 2026-10-01 the <code>docs/benchmarks/</code> directory contains only the Terminal-Bench 2.1 file, a Codex canary note and a low-compute mechanism note — <strong>there is no v2 benchmark</strong>.</p>
<p>That distinction is not pedantry. A new activation path, new routing controls and eleven adapter changes alter the very execution path the benchmark measured. Every article published after 2026-09-09 that presents 45% as the current, measured performance of the software you are about to install is citing a superseded artifact. The honest framing is: &ldquo;v1 measured +14.61pp on one run; v2 is unmeasured.&rdquo;</p>
<h2 id="how-does-autoprompt-compare-to-other-agent-skills">How Does Autoprompt Compare to Other Agent Skills?</h2>
<p>Autoprompt sits in a small minority that works at all, and about in the middle of the curated-skill field — and the second half of that sentence is the part nobody likes.</p>
<p>Two independent benchmarks bracket this category. SWE-Skills-Bench evaluated 49 skills across roughly 565 task instances with deterministic pytest verification and found that <strong>39 of 49 skills produced zero pass-rate improvement</strong>, with an average gain of just +1.2%. Token overhead ranged from −78% to +451% while pass rates stayed flat, and three skills actively degraded performance by up to −10%. Anyone assuming a popular skill helps has roughly an 80% chance of being wrong about a random one.</p>
<p>SkillsBench is the kinder baseline: 87 tasks across 8 domains and 18 model-harness configurations, with matched no-skill and with-skill conditions. Curated skills lifted average pass rate from 33.9% to 50.5% — <strong>+16.6pp</strong>, a 25.5% normalised gain, with per-configuration gains from +4.1pp to +25.7pp.</p>
<p>Set Autoprompt&rsquo;s +14.61pp against those numbers:</p>
<table>
  <thead>
      <tr>
          <th>Evidence set</th>
          <th>Scope</th>
          <th>Effect on pass rate</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>SWE-Skills-Bench average</td>
          <td>49 skills, ~565 instances</td>
          <td>+1.2 pp</td>
      </tr>
      <tr>
          <td>SWE-Skills-Bench null result</td>
          <td>39 of 49 skills</td>
          <td>0 pp</td>
      </tr>
      <tr>
          <td>SkillsBench fleet average</td>
          <td>87 tasks, 18 configs</td>
          <td>+16.6 pp</td>
      </tr>
      <tr>
          <td><strong>Autoprompt, Terminal-Bench 2.1</strong></td>
          <td><strong>1 run, 89 tasks</strong></td>
          <td><strong>+14.61 pp</strong></td>
      </tr>
  </tbody>
</table>
<p>The headline that looks enormous in isolation is slightly below the fleet average of curated skills in context. Both statements are true, and together they give the honest verdict: Autoprompt is in the minority of skills that measurably help, and it is not an outlier among them. If your baseline assumption was &ldquo;agent skills do nothing&rdquo; (a 1.2% expectation), this is a large win. If your baseline was &ldquo;curated skills give about 16 points,&rdquo; this is ordinary.</p>
<h2 id="what-is-the-trap-next-door-for-self-generated-skills">What Is the Trap Next Door for Self-Generated Skills?</h2>
<p>The same SkillsBench work contains the finding most relevant to anyone tempted to copy the pattern: when agents were asked to author their own skills before starting a task, they performed <strong>worse than baseline, dropping 8.1 to 11.5 percentage points</strong>.</p>
<p>That is a direct warning for teams that read Autoprompt&rsquo;s architecture and conclude they can have their orchestrator auto-generate a project-specific checklist each run. The evidence says the opposite. Curated skills help; self-generated skills hurt, plausibly because an agent writing its own procedure optimises for plausibility rather than for the verification steps it is least likely to perform unaided. Autoprompt&rsquo;s own design leans the right way here — it is explicit, invoked by name (<code>/autoprompt</code>, <code>$autoprompt</code>, <code>/skill:autoprompt</code>) rather than authored on the fly — but the temptation to bolt on auto-generation is real, and the literature is not neutral about it.</p>
<h2 id="what-does-the-45-cost-in-tokens-and-time">What Does the 45% Cost in Tokens and Time?</h2>
<p>It costs roughly 3x wall-clock time and 2x tokens, and nobody has published the exchange rate between the two. We can compute an approximate one.</p>
<table>
  <thead>
      <tr>
          <th>Cost component</th>
          <th>Published value</th>
          <th>Evidence quality</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Wall-clock time</td>
          <td>~3x</td>
          <td>Project states it is a planning estimate from user reports</td>
      </tr>
      <tr>
          <td>Tokens</td>
          <td>~2x</td>
          <td>Same — timing and token logs were not retained</td>
      </tr>
      <tr>
          <td>Pass-rate gain</td>
          <td>+14.61 pp (CI +2.02 to +27.20)</td>
          <td>Measured once, n=89</td>
      </tr>
      <tr>
          <td>Implied efficiency</td>
          <td>≈1.5–1.6 pp of pass rate per extra 10% of tokens</td>
          <td>Derived by this review</td>
      </tr>
  </tbody>
</table>
<p>The project&rsquo;s own README labels the 3x/2x figures &ldquo;planning estimates based on user experience reports, not measured benchmark results,&rdquo; which is unusually candid — most projects would print them as fact. But it also means the one number a buyer needs does not exist. At an assumed 1.9x token multiplier, the +14.61pp gain comes at roughly 90% more tokens than baseline; spread across the 13 extra solves, that is about <strong>7% of a full baseline run&rsquo;s token budget per additional task solved</strong>. Or, in the form that survives a business case: you are buying roughly 1.5 percentage points of pass rate for every extra 10% of tokens you spend.</p>
<p>Whether that is worth it depends entirely on the alternative use of the same tokens — a second opinion from a stronger model, a broader test suite, or simply a retry with a different seed. Autoprompt&rsquo;s honest framing is a price, not a percentage, and the price is currently an estimate of an estimate.</p>
<h2 id="why-do-failure-reduction-tools-exist-at-all">Why Do Failure-Reduction Tools Exist at All?</h2>
<p>Because long-horizon reliability, not raw capability, is the binding constraint — and the arithmetic of chained steps explains why a verification layer can help at all.</p>
<p>METR&rsquo;s time-horizon research found that frontier agents succeed on nearly 100% of tasks a human would finish in under about four minutes, but on under 10% of tasks taking a human over roughly four hours, with the 50%-reliability horizon doubling approximately every seven months since 2019. Capability is not the problem at long horizons; the ability to keep 20 or 50 steps correct is.</p>
<p>Compounding makes that concrete. At 95% per-step reliability, a 20-step task succeeds about 36% of the time and a 50-step task about 8%.</p>
<table>
  <thead>
      <tr>
          <th>Per-step reliability</th>
          <th>20-step task success</th>
          <th>50-step task success</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>90%</td>
          <td>12.2%</td>
          <td>0.5%</td>
      </tr>
      <tr>
          <td>95%</td>
          <td>35.9%</td>
          <td>7.7%</td>
      </tr>
      <tr>
          <td>97%</td>
          <td>54.4%</td>
          <td>21.8%</td>
      </tr>
      <tr>
          <td>99%</td>
          <td>81.8%</td>
          <td>60.5%</td>
      </tr>
  </tbody>
</table>
<p>This is why a loop that adds independent verification is a plausible intervention rather than a gimmick: it targets step-level error, which is where the compounding lives. It is also why the <em>size</em> of Autoprompt&rsquo;s measured gain is plausible — moving an agent from roughly 67% to roughly 82% on a hard terminal benchmark is exactly the scale of improvement you would expect from catching a fraction of mid-trajectory mistakes, not from making the model smarter. And it is why the measurement is hard: the same compounding that makes the intervention valuable makes single-run estimates volatile.</p>
<h2 id="is-8202-actually-a-good-score">Is 82.02% Actually a Good Score?</h2>
<p>It is a mid-table score. Public Terminal-Bench 2.1 leaderboards put top model-harness combinations at roughly 87–91% on the same 89 tasks, so Autoprompt&rsquo;s 82.02% — with an 89-task, OpenCode 1.18.7, DeepSeek V4 Flash harness — sits below the frontier, not at it.</p>
<p>The comparability warning matters as much as the number. Terminal-Bench scores are not interchangeable across boards: the official board measures agent-plus-model combinations, Artificial Analysis tests models directly with labelled effort tiers, and vals.ai retests everything under a unified Terminus 2 harness. Citing &ldquo;82.02%&rdquo; without naming the harness and model invites a wrong read, in either direction — as a disappointing score when the harness was never frontier, or as a frontier score when the board is measuring something else. The safe sentence is the specific one: OpenCode 1.18.7 with DeepSeek V4 Flash and Autoprompt scored 82.02% on Terminal-Bench 2.1.</p>
<h2 id="why-do-the-stars-and-npm-downloads-point-the-other-way">Why Do the Stars and npm Downloads Point the Other Way?</h2>
<p>Adoption is decelerating while the headline number compounds, and the open-issue queue tells you where maintainer attention is going.</p>
<table>
  <thead>
      <tr>
          <th>Signal (as of 2026-10-01)</th>
          <th>Value</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GitHub stars</td>
          <td>1,298</td>
      </tr>
      <tr>
          <td>Forks</td>
          <td>88</td>
      </tr>
      <tr>
          <td>Open issues</td>
          <td>3 — all Windows activation blockers</td>
      </tr>
      <tr>
          <td>Created / last push</td>
          <td>2026-08-17 / 2026-09-28</td>
      </tr>
      <tr>
          <td>Licence</td>
          <td>MIT</td>
      </tr>
      <tr>
          <td>Stars/day, first 3 days</td>
          <td>≈146/day</td>
      </tr>
      <tr>
          <td>Stars/day, following ~38 days</td>
          <td>≈23/day</td>
      </tr>
      <tr>
          <td>npm <code>latest</code></td>
          <td>1.0.4 (published 2026-08-21)</td>
      </tr>
      <tr>
          <td>npm downloads, week to 2026-09-29</td>
          <td>32</td>
      </tr>
      <tr>
          <td>npm downloads, month to 2026-09-29</td>
          <td>144</td>
      </tr>
  </tbody>
</table>
<p>The star curve is a launch burst into a slow tail, not compounding growth: roughly 146 stars a day for three days, then about 23 a day for the next five weeks. npm tells the same story more sharply — the published package is still 1.0.4, two minor lines behind GitHub&rsquo;s v2.0.0, and therefore still carries nine providers and the old <code>mode=</code>/<code>max_subs=</code> interface. Weekly downloads are in the tens.</p>
<p>Meanwhile the three open issues are all Windows activation failures (#27 native Windows refused, #28 OpenCode <code>cmd-shim</code> EINVAL, #29 npm shims for native package bins), and the README&rsquo;s own requirement line says &ldquo;these versions passed Linux runs.&rdquo; The practical reading: the eleven-provider claim is a <strong>code-path</strong> claim, the <strong>verified execution surface is Linux</strong>, and if you are on Windows you are currently the maintainer&rsquo;s queue. If you want the tool to keep working, file issues rather than stars.</p>
<h2 id="how-portable-is-autoprompt-across-its-11-host-agents">How Portable Is Autoprompt Across Its 11 Host Agents?</h2>
<p>The adapters are real, and portability is the dimension no agent-skill framework has fully solved.</p>
<p>The tested-version table lists Claude Code 2.1.263, Codex 0.148.0, OpenCode 1.18.29, Kilo Code 7.5.15, VS Code 1.136.1, Prime Agent 0.7.2, Oh My Pi 18.1.14, DeepSeek Harness 0.1.2-rc.1, Reasonix 1.30.0, Hermes Agent 0.21.1 and Grok Build 1.0.13 — eleven hosts, each pinned to a version that passed on Linux. The v2 CLI unifies activation across them, which is genuine progress over the v1 invocation-per-host mess.</p>
<p>Two gaps remain. First, one independent review rates the project&rsquo;s portability &ldquo;portable with changes&rdquo; rather than portable, because the v2 CLI activation path is provider-specific — the unification is at the command surface, not at the execution semantics. Second, custom model routing does not exist on all hosts, so a team that needs to pin a specific model per role will find the coverage uneven. Independent taxonomy work reaches the same conclusion at the framework level: no agent-skill system currently covers specification, context, roles, execution, validation and portability simultaneously. Budget for a portability tax on any host other than Linux plus a mainstream agent.</p>
<h2 id="who-should-use-autoprompt-in-2026-and-who-should-not">Who Should Use Autoprompt in 2026, and Who Should Not?</h2>
<p>Buy it for hard, ambiguous tasks on a Linux workstation with an established host agent — and skip it everywhere else.</p>
<p><strong>A good fit if you:</strong></p>
<ul>
<li>Run Claude Code, Codex or OpenCode on Linux or macOS and already have a working test suite.</li>
<li>Spend your time on multi-step, underspecified tasks where the agent &ldquo;gets stuck mid-task&rdquo; rather than failing instantly.</li>
<li>Value the verification structure more than the speed — the independent-check discipline is the durable part of the design, and it is transferable even if you later drop the tool.</li>
<li>Need evidence you can point at: the +14.61pp gain is statistically significant, and you now know its interval.</li>
</ul>
<p><strong>A poor fit if you:</strong></p>
<ul>
<li>Do very small tasks. A 2x token multiplier on a one-file change is pure loss, and the project itself notes gains may vary significantly between small and large tasks.</li>
<li>Work primarily on Windows today. All three open issues are Windows activation failures.</li>
<li>Need custom model routing on every host.</li>
<li>Have no local terminal runtime for the agent to execute in — without that, there is nothing for the verification loop to verify.</li>
<li>Are choosing between Autoprompt and simply spending the same 2x tokens on a stronger model or a larger test suite. That comparison has never been published, and it is the one that decides most budgets.</li>
</ul>
<p>Read against the third-party score available — FollowAgents rates the project 69/100, &ldquo;use with care,&rdquo; with praise focused specifically on disclosure quality and a caveat that the review is stale on provider count and version — the fair summary is that the engineering and the honesty are both above average while the evidence remains one run.</p>
<h2 id="how-do-you-reproduce-or-reject-the-claim-yourself">How Do You Reproduce or Reject the Claim Yourself?</h2>
<p>The claim is falsifiable, and the cheapest path is a paired test on your own backlog.</p>
<ol>
<li><strong>Fix the harness.</strong> Pin one host agent and one model. Mixed hosts invalidate every comparison, as the three-board divergence on Terminal-Bench 2.1 shows.</li>
<li><strong>Select tasks with headroom.</strong> Pick 30–50 tasks you already fail. Testing on tasks you pass measures nothing.</li>
<li><strong>Run both arms on the same tasks.</strong> Baseline run, then an Autoprompt run, same commit, same environment, ideally same day to reduce drift.</li>
<li><strong>Score deterministically.</strong> Use your test suite or an external grader, not agent self-assessment. This is the step most teams skip, and it is the step that matters — an agent judging its own work is precisely the failure mode the skill exists to prevent.</li>
<li><strong>Compute the interval, not just the difference.</strong> At 30 paired tasks you have almost no power; a 15-point difference on 30 tasks will not be distinguishable from noise. If your observed difference is small, treat the result as inconclusive rather than negative.</li>
<li><strong>Log tokens and wall clock.</strong> The project could not, and that is why its cost figure is an estimate. Your run should not repeat that mistake.</li>
<li><strong>Run <code>autoprompt doctor --strict</code></strong> before measuring, so a partial install is not scored as a skill failure.</li>
</ol>
<p>Expect a real but unglamorous outcome. The published effect is +14.61pp with a lower bound of +2.02pp; on a 40-task set, a plausible result is three to six extra solves at roughly double the tokens.</p>
<h2 id="what-is-the-verdict-on-the-autoprompt-skill">What Is the Verdict on the Autoprompt Skill?</h2>
<p>It works, the number is real, and the number is smaller and less certain than the headline implies.</p>
<p>The defensible statement of the evidence is this: on a single 89-task Terminal-Bench 2.1 run, wrapping OpenCode 1.18.7 with Autoprompt raised pass rate from 67.42% to 82.02%, a +14.61pp gain (95% CI +2.02pp to +27.20pp, p=0.025) that reduces the failure rate 44.8% — 1.81x fewer failures — at an estimated 3x time and 2x tokens, with no v2 benchmark yet published.</p>
<p>Everything past the comma is why this article exists. Most surfaces that rank for &ldquo;autoprompt skill&rdquo; have already lost it: catalog pages still publish nine providers, v1.0.4 and 940 stars while the repository says eleven providers, v2.0.0 and 1,298 stars. The 45% survived intact while every verifiable detail around it decayed — which is the normal behaviour of a good headline and the reason it deserves an error bar. Use the skill for hard tasks, budget the tokens honestly, and quote the interval.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Does the Autoprompt skill actually work?</strong>
The evidence says yes, within limits. Autoprompt&rsquo;s own Terminal-Bench 2.1 run moved OpenCode 1.18.7 from 60/89 to 73/89 solves (+14.61pp, p=0.025), and independent category benchmarks show most agent skills produce no measurable gain at all, so a skill that clears the bar is a minority case. However, the result is a single 89-task run, the treatment arm&rsquo;s per-task ledger was not retained, and the 95% confidence interval spans +2.02pp to +27.20pp. Treat it as a real but imprecisely measured improvement, not a guarantee.</p>
<p><strong>Where does the 45% figure come from and is it accurate?</strong>
It comes from failures falling from 29 to 16 on 89 Terminal-Bench 2.1 tasks. That is a 44.8% reduction in the failure rate, so 45% is accurate as a description of that specific metric. It is not a 45% improvement in capability: the pass rate rose 14.61 points (67.42% to 82.02%), the relative pass-rate gain is 21.67%, and in odds terms the wrapped agent fails 1.81x less often. Any source describing it as &ldquo;45% better&rdquo; has changed the meaning of the number.</p>
<p><strong>What is the confidence interval on the Autoprompt benchmark, and why does it matter?</strong>
The +14.61pp difference corresponds to a 95% confidence interval of roughly +2.02pp to +27.20pp at n=89 per arm (two-proportion z=2.24; Fisher exact two-sided p=0.038). No competitor publishes this. It matters because the interval is wide: the same data is consistent with a benefit as large as 27 points and as small as 2 points, and at 2 points the tool would win about one extra task out of 89. With n=89, statistical power to detect the observed effect is only about 61%, so this should be read as a positive single run, not a settled effect size.</p>
<p><strong>Does Autoprompt work on Windows, and is the npm package current?</strong>
Windows is currently the hard gate: all three open issues on the repository as of 2026-10-01 are Windows activation failures (#27 native Windows refused, #28 OpenCode cmd-shim EINVAL, #29 npm shims for native package bins), and the project states that its tested versions &ldquo;passed Linux runs.&rdquo; Separately, npm&rsquo;s <code>latest</code> tag is still 1.0.4 (published 2026-08-21), two minor lines behind the GitHub v2.0.0 release of 2026-09-09, so the npm package still ships nine providers and the old <code>mode=</code>/<code>max_subs=</code> interface. Install from the v2.0.0 release if you need the CLI and the eleven adapters.</p>
<p><strong>How much does Autoprompt cost in tokens and time, and is it worth it?</strong>
The project estimates roughly 3x wall-clock time and 2x tokens, and explicitly labels both as planning estimates from user experience reports because timing and token logs were not retained. Pair that with the +14.61pp gain and you get approximately 1.5 to 1.6 percentage points of pass rate per extra 10% of tokens — a real but modest exchange rate. It is worth it on hard, multi-step tasks where your agent tends to get stuck and you have a deterministic test suite to verify against. It is a clear loss on small single-file changes, where the multiplier buys nothing.</p>
]]></content:encoded></item><item><title>Autoprompt Skill Review: The Coding Agent Prompt Skill That Cuts Failures 45%</title><link>https://baeseokjae.github.io/posts/autoprompt-skill-coding-agent-failures/</link><pubDate>Mon, 28 Sep 2026 12:55:34 +0000</pubDate><guid>https://baeseokjae.github.io/posts/autoprompt-skill-coding-agent-failures/</guid><description>Autoprompt claims 45% fewer coding-agent failures. We trace the number to its run, verify the repo state on 2026-09-28, and price the token cost.</description><content:encoded><![CDATA[<p>Autoprompt is a coding agent prompt skill that injects an orchestration procedure — plan, delegate, implement, verify — into 11 host agents, including Claude Code, Codex, and OpenCode. Its headline claim of 45% fewer failures comes from a single 89-task Terminal-Bench 2.1 run where failures fell from 29 to 16. That result is real arithmetic but version 1 evidence, and it costs roughly 3x wall-clock time and 2x tokens.</p>
<h2 id="what-the-45-figure-actually-measures">What the 45% Figure Actually Measures</h2>
<p>The number everyone quotes is not a percentage of tasks solved and it is not a percentage reduction in tokens. It is a reduction in failures on one benchmark, measured once.</p>
<p>The run is documented as Terminal-Bench 2.1 with 89 tasks, OpenCode 1.18.7 as the harness, and DeepSeek V4 Flash as the model. Baseline solves were 60 of 89. With Autoprompt active, solves rose to 73 of 89. Failures went from 29 to 16, which is a 44.8% reduction — rounded to 45% in every headline since.</p>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>Baseline</th>
          <th>With Autoprompt</th>
          <th>Delta</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tasks solved</td>
          <td>60 / 89</td>
          <td>73 / 89</td>
          <td>+13 tasks</td>
      </tr>
      <tr>
          <td>Solve rate</td>
          <td>67.42%</td>
          <td>82.02%</td>
          <td>+14.61 pp</td>
      </tr>
      <tr>
          <td>Failures</td>
          <td>29</td>
          <td>16</td>
          <td>-13 (-44.8%)</td>
      </tr>
      <tr>
          <td>Wall-clock time</td>
          <td>1x</td>
          <td>~3x (estimated)</td>
          <td>not measured</td>
      </tr>
      <tr>
          <td>Token usage</td>
          <td>1x</td>
          <td>~2x (estimated)</td>
          <td>not measured</td>
      </tr>
  </tbody>
</table>
<p>Two things make that table worth reading slowly. First, the denominator: 13 tasks out of 89 is a meaningful but bounded result, and 44.8% is a property of the failure slice (13 of 29), not of the workload. The same result is also described in the project&rsquo;s own README as &ldquo;about 2x fewer mistakes,&rdquo; which is the identical measurement expressed against the smaller surviving set. Neither framing is dishonest; the choice of denominator is what makes a 13-task improvement look like a near-halving.</p>
<p>Second, and more important: cost was not instrumented. The project discloses the ~3x time and ~2x token trade-off as &ldquo;planning estimates based on user experience reports, not measured benchmark results,&rdquo; and explains that timing and token logs were not retained. So the accuracy half of the claim is a measured benchmark; the cost half is not a measurement at all. If you are budgeting an unattended run, treat 3x as a floor and verify it on your own task distribution before trusting it in a CI budget.</p>
<h2 id="autoprompt-v2-what-changed-since-the-original-review">Autoprompt v2: What Changed Since the Original Review</h2>
<p>The repository has moved past the benchmark. Autoprompt shipped v2.0.0 on 2026-09-09, and the README now labels the 45% result as &ldquo;version 1 benchmarks,&rdquo; with a note that version 2 benchmarks will follow. No v2 efficacy number exists as of 2026-09-28. That matters for anyone citing the headline today: the marketing line is unchanged, but the artifact it refers to is two releases old.</p>
<p>What v2 actually changed is scope and control, not the accuracy thesis:</p>
<table>
  <thead>
      <tr>
          <th>Area</th>
          <th>v1</th>
          <th>v2.0.0</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Interface</td>
          <td>Slash-command modes</td>
          <td>CLI: <code>autoprompt activate PROVIDER --target /abs/project -- '&lt;goal&gt;'</code></td>
      </tr>
      <tr>
          <td>Providers</td>
          <td>Codex, OpenCode, Prime Agent</td>
          <td>11, adding Hermes Agent and Grok Build</td>
      </tr>
      <tr>
          <td>Controls</td>
          <td>Mode presets</td>
          <td>Path selection plus model, effort, and concurrency flags</td>
      </tr>
      <tr>
          <td>Subagent bounds</td>
          <td>Not configurable</td>
          <td><code>--concurrency tokensaver|wide|custom --max-subs N</code></td>
      </tr>
      <tr>
          <td>Installer lifecycle</td>
          <td>Manual</td>
          <td>Managed install/update lifecycle</td>
      </tr>
  </tbody>
</table>
<p>The v2 workflow is more explicit about where control sits. Instead of a mode switch, you select a work path — <code>path=auto|direct|light|roadmap</code> — and a concurrency profile. The default is <code>tokensaver</code>, which caps fan-out at six subagents. The documentation is clear that choosing a path does not waive authorization, capability, budget, or verification requirements, and that invalid combinations error out rather than silently falling back to a different route.</p>
<h2 id="the-evidence-boundary-moved-and-it-got-thinner">The Evidence Boundary Moved, and It Got Thinner</h2>
<p>The single most important fact for a 2026 buyer is not in the README: the benchmark evidence is no longer fully checkable.</p>
<p>The per-task verdict file that the benchmark documentation points to for all 89 retained verdicts now returns HTTP 404. The <code>benchmark/</code> directory that once held the raw artifacts is absent from the main branch — verified by walking the recursive git tree listing of 1,381 paths, not inferred from a link. The methodology summary survives in <code>docs/benchmarks/terminal-bench-2.1.md</code>, but the per-task data a skeptical reader would use to audit the run does not.</p>
<p>That is a reproducibility regression, and it is the correct thing to hold against this project, because prompt-skill vendors have an obvious incentive to publish the summary and not the ledger. A benchmark you cannot audit is a marketing claim with a methodology section. Independent verification does not become impossible — you can rerun Terminal-Bench yourself — but the cheapest possible audit path (read the verdicts) is closed, and that cost falls on the buyer.</p>
<p>The v2.0.0 release notes show a similar honesty pattern in miniature. They concede that &ldquo;full live acceptance of the final Codex artifact is not established,&rdquo; and that Windows and macOS builds are installer kits with native runtime verification remaining Linux-only. That is unusually candid disclosure for a release announcement, and it is the reason this review reads the project as careful rather than reckless: it publishes caveats that hurt it. It simply has not published version 2 evidence yet.</p>
<h2 id="eleven-providers-one-verified-platform">Eleven Providers, One Verified Platform</h2>
<p>Provider count is the feature most third-party articles get stale. Autoprompt now claims 11 supported hosts. The v2.0.0 work extended the Codex workflow across all of them and added Hermes Agent and Grok Build to a list that previously stopped at Codex, OpenCode, and Prime Agent.</p>
<p>The platform support story is narrower than the provider story. All three open items on the repository concern Windows:</p>
<ul>
<li>Issue #27 — activation is refused on Windows with <code>PROVIDER_UNSUPPORTED</code> because all 10 reviewed local records are <code>linux/x64</code> and <code>trusted-public-keys.json</code> lists no Windows keys.</li>
<li>Issue #28 — OpenCode activation fails on the npm <code>cmd</code> shim.</li>
<li>PR #29 — an in-flight fix for Windows npm shims for native package bins.</li>
</ul>
<p>Read together with the release notes&rsquo; Linux-only verification line, the practical rule is simple: the 11-provider claim is a code-path claim, and the verified execution surface is Linux. If your team codes on Windows, this is not yet a drop-in tool, and the open issues say so more precisely than any review can.</p>
<h2 id="paths-and-run-controls-where-v2-spends-your-tokens">Paths and Run Controls: Where v2 Spends Your Tokens</h2>
<p>Cost in this category is not a fixed multiplier; it is a budget you set. The mechanism is straightforward and is the same one the independent literature describes: every phase is another model call over the same code, the judge is a separate model call with its own context, and parallel lanes re-read the same files. Fan-out is therefore the dominant cost lever, more than the model choice per phase.</p>
<p>The v2 controls map to that reality directly:</p>
<ul>
<li><code>--concurrency tokensaver</code> — at most six subagents, the default. Sensible for a laptop or a metered API key.</li>
<li><code>--concurrency wide</code> — maximum parallel lanes. Fast, and the most expensive way to run it.</li>
<li><code>--concurrency custom --max-subs N</code> — the setting most teams should actually use, because it lets you align fan-out with the size of the change rather than the optimism of the moment.</li>
<li><code>path=direct</code> — minimal routing for a small, well-specified change.</li>
<li><code>path=light</code> / <code>roadmap</code> — progressive planning depth for larger or ambiguity-heavy work.</li>
</ul>
<p>One operational caution worth repeating from independent walkthroughs: six lanes can touch deployment credentials, so run unattended sessions in a disposable VM or a container with scoped keys rather than on your working machine. And there is a real quality floor below which orchestration is pure overhead — if the task has nothing to execute or test, the verification phase produces one more opinion at full token price.</p>
<p>The default <code>tokensaver</code> profile is a reasonable acknowledgment of that floor. It is also the setting to keep when you are running agents on a VPS budget, where a wide fan-out can turn a ten-minute change into a twenty-dollar one.</p>
<h2 id="what-independent-skill-benchmarks-say-about-this-category">What Independent Skill Benchmarks Say About This Category</h2>
<p>Autoprompt is one instance of a category that has been measured independently, and the category averages are not kind. Any review that quotes 45% without quoting the surrounding literature is selling, not informing.</p>
<p>The closest comparable study, SWE-Skills-Bench, ran 565 task instances across 49 public SWE skills with GitHub repositories pinned at fixed commits and execution-based verification. Thirty-nine of the 49 skills produced zero pass-rate improvement. The average gain was +1.2%. Token overhead ranged from -78% to +451% while pass rates stayed flat. Seven specialized skills gained up to +30%, and three actively degraded performance by up to -10% through version-mismatched guidance.</p>
<p>SkillsBench, testing 87 tasks across 18 model-harness configurations with matched no-skill controls, found that curated skills raise the average pass rate from 33.9% to 50.5% — a +16.6 percentage point gain — but with gains ranging from +4.1 to +25.7 points, and 13 of 87 tasks showing negative deltas. Its most useful structural finding is that skill value is a property of the specific stack, not of skills in general. Compact focused skills beat exhaustive bundles by a wide margin, and loading four or more skills together yielded only +10.1 points.</p>
<p>A third line of work explains the failure mode that matters most for a verification-heavy tool. The paper &ldquo;Agent Skills Can Be Harmful&rdquo; confirmed 307 skill-induced failures across those same benchmarks: 125 functional and 182 efficiency regressions. Efficiency damage was dominated by what the authors call Excessive Procedure at 62.6% — not prompt length — with excessive verification (67 cases) and heavy implementation pipelines (30 cases) as the largest drivers. Functional failures rarely came from irrelevant skills; seemingly relevant skills caused wrong or omitted implementation elements in 68.8% of cases. That is Autoprompt&rsquo;s exact design surface: it turns verification checklists into mandatory work, which is simultaneously the mechanism of its accuracy claim and the largest documented source of token regression in the literature.</p>
<p>Two further results bound the realistic ceiling. Under progressively realistic conditions, where the agent must retrieve skills itself instead of receiving a hand-picked one, skill benefits decay toward no-skill baselines — and query-specific refinement recovers performance, lifting Claude Opus 4.6 from 57.7% to 65.5% on Terminal-Bench 2.0. A separate 500-skill, roughly 38,000-trajectory evaluation found relevant skills lifting scores by 5.5 to 22 points, with open-weight GLM 5.1 plus a skill scoring 91.1 at about $0.89 per scenario against Opus 4.8 at 92.7 for $3.26.</p>
<p>That last number is the under-discussed lever. If you pair a cheaper model with a strong procedure, you can approach frontier quality at a third to a quarter of the cost — which changes the economics of Autoprompt&rsquo;s ~2x token multiplier far more than shaving one subagent will.</p>
<p>For context on why this market exists at all: Anthropic&rsquo;s 2026 survey of more than 500 technical leaders found about 90% of organizations use AI for coding, 86% deploy coding agents for production code, and 57% run agents on multi-stage workflows. Time savings were reported at 59% across code generation, review and testing, and research alike. Integration with existing systems (46%) is the top adoption blocker, ahead of data quality (42%) and cost (43%).</p>
<h2 id="the-npm-versus-github-distribution-gap">The npm Versus GitHub Distribution Gap</h2>
<p>If you install this tool with the command most articles still publish, you get the old version.</p>
<table>
  <thead>
      <tr>
          <th>Channel</th>
          <th>Version</th>
          <th>Published</th>
          <th>Notes</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GitHub</td>
          <td>v2.0.0</td>
          <td>2026-09-09</td>
          <td>11 providers, CLI, path and concurrency controls</td>
      </tr>
      <tr>
          <td>npm (<code>autoprompt-skill</code>)</td>
          <td>1.0.4</td>
          <td>2026-08-21</td>
          <td>Zero runtime dependencies, two majors behind</td>
      </tr>
  </tbody>
</table>
<p>npm weekly downloads were 38 for the week of 2026-09-21 to 2026-09-27, against third-party trackers still publishing the older figure of 535 per week — a 14x discrepancy that mostly reflects how slowly catalog data propagates.</p>
<p>Repository state on the date this review was verified, 2026-09-28: 1,295 stars, 88 forks, 3 open issues, MIT license, 44 commits, created 2026-08-17, latest release v2.0.0 on 2026-09-09. Star velocity at launch was cited at 437 stars in three days, roughly 153 per day.</p>
<p>The version divergence explains most of the factual disagreements between reviews of this tool. Published comparisons disagree on provider count (9 versus 11) and star count (590 versus 1,300) because they were written against different points in a fast-moving repository. Some also repeat unattributed anecdotes — debugging-time reductions at an investment firm, e-commerce reliability gains — with no source at all. The 45% figure is exactly the kind of claim that survives this propagation intact while every verifiable detail around it goes stale; the version data in this review is dated deliberately so you can see how far it has drifted since.</p>
<h2 id="where-autoprompt-earns-its-cost-and-where-it-does-not">Where Autoprompt Earns Its Cost, and Where It Does Not</h2>
<p>The category literature says a minority of skills genuinely help and the winners are concrete workflows, not general best-practice documents. Autoprompt is positioned in the subcategory that Anthropic reports as delivering the largest quality improvement: verification-type skills. That is the strongest structural argument in its favor, and it does not depend on the 45% at all.</p>
<p>The strongest practical argument is role separation. A reviewer that runs the code and can disagree on evidence is a different thing from a model approving its own output, and that architectural difference is what makes the benchmark result plausible — plan, implement, and verify as distinct calls with distinct context. Autonomy here is bounded by design: the workflow stops for choices that change the result or actions that require authority. If you were sold &ldquo;fully autonomous coding,&rdquo; that is not what this is, and the stopping behavior is a feature.</p>
<p>Where it does not earn its cost:</p>
<ul>
<li><strong>Small, well-specified changes.</strong> The project&rsquo;s own caveat is that small tasks &ldquo;may differ significantly,&rdquo; and the efficiency-regression literature is dominated by exactly this pattern. Use <code>path=direct</code> or skip it.</li>
<li><strong>Tasks with nothing to verify.</strong> No executable check means the verification phase buys an opinion, not evidence.</li>
<li><strong>Windows development teams.</strong> The verified surface is Linux; three open items say so.</li>
<li><strong>Workflows needing an auditable number.</strong> The per-task evidence is 404 and v2 has no benchmark. You are trusting a methodology document.</li>
<li><strong>Multi-skill setups.</strong> If you already load four or more skills, the incremental value measured by SkillsBench drops sharply, and skill metadata competes for context — Claude Code caps skill metadata at roughly 1% of the context window (~2,000 tokens, about 15-25 skills on a 200K model) before descriptions get truncated.</li>
</ul>
<h2 id="verdict-who-should-use-it-in-2026">Verdict: Who Should Use It in 2026</h2>
<p>Autoprompt is a well-engineered orchestration skill with a real, single-run accuracy result, unusually candid disclosure about its own gaps, and a version 2 that has not been benchmarked. The 45% is defensible only with its denominator attached: 13 tasks out of 89, on one harness, with one model, costing an unmeasured ~3x time and ~2x tokens.</p>
<p>Adopt it if you run long agentic coding sessions on Linux, you already have test suites the verifier can execute, and you want structural role separation between implementation and review. Start with the default <code>tokensaver</code> concurrency, pin <code>--max-subs</code> to something honest for your change size, and measure your own token delta in week one.</p>
<p>Do not adopt it on the strength of the headline. Do not cite it as an established result without noting that the evidence directory is gone and v2 is untested. And if you need independent confirmation before standardizing, rerun Terminal-Bench on your own repositories — which is what the project&rsquo;s own disclosure effectively invites you to do.</p>
<p>The honest one-line summary: this is the exception the skeptical literature allows for, priced at 3x the wall clock, sold with a number that is two releases old.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Does Autoprompt really reduce coding agent failures by 45%?</strong></p>
<p>On one measured run, yes: failures fell from 29 to 16 across 89 Terminal-Bench 2.1 tasks using OpenCode 1.18.7 and DeepSeek V4 Flash, a 44.8% reduction. That is a single benchmark on a single harness and model combination, and the project itself labels it a version 1 result. The same 13-task improvement is also described as &ldquo;about 2x fewer mistakes.&rdquo;</p>
<p><strong>Is the 45% benchmark independently verifiable?</strong></p>
<p>Not from the project&rsquo;s own artifacts. The per-task verdict file the benchmark docs link to returns HTTP 404, and the <code>benchmark/</code> directory is absent from the main branch among 1,381 paths in the recursive tree. The methodology document remains, so you can read how the run was done, but the raw verdicts needed to audit it are no longer published. Independent verification means rerunning the benchmark yourself.</p>
<p><strong>How much more expensive is Autoprompt than a plain coding agent?</strong></p>
<p>The disclosed figure is roughly 3x wall-clock time and 2x tokens — and the project states plainly that these are planning estimates from user experience reports, not measured benchmark results, because timing and token logs were not retained. Actual overhead depends mostly on fan-out, which you control via the <code>tokensaver</code> (six subagents maximum), <code>wide</code>, or <code>custom --max-subs N</code> concurrency profiles.</p>
<p><strong>Which coding agents does Autoprompt support, and on which platforms?</strong></p>
<p>v2.0.0 claims 11 providers, extending the original Codex, OpenCode, and Prime Agent list and adding Hermes Agent and Grok Build. Platform verification is narrower: the release notes concede native runtime verification remains Linux-only, and the three open repository items all concern Windows — <code>PROVIDER_UNSUPPORTED</code> activation failures caused by linux-only trusted keys, an OpenCode activation failure on the npm <code>cmd</code> shim, and an in-flight PR fixing Windows npm shims.</p>
<p><strong>Should I install Autoprompt from npm or GitHub?</strong></p>
<p>GitHub. npm still ships 1.0.4, published 2026-08-21, while the repository is at v2.0.0 from 2026-09-09 — a two-major-version gap that means the npm package lacks the CLI, the path controls, and the expanded provider support. npm weekly downloads were 38 for the week ending 2026-09-27, so the stale channel is also the less-used one.</p>
]]></content:encoded></item></channel></rss>