<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Weakened Test Assertion Detection AI on RockB</title><link>https://baeseokjae.github.io/tags/weakened-test-assertion-detection-ai/</link><description>Recent content in Weakened Test Assertion Detection AI on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 02:20:51 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/weakened-test-assertion-detection-ai/index.xml" rel="self" type="application/rss+xml"/><item><title>Prove-It: An Adversarial Verification Agent Skill for Claude Code (2026 Guide)</title><link>https://baeseokjae.github.io/posts/prove-it-adversarial-verification-skill/</link><pubDate>Thu, 01 Oct 2026 02:20:51 +0000</pubDate><guid>https://baeseokjae.github.io/posts/prove-it-adversarial-verification-skill/</guid><description>Prove-It is a single-file Agent Skill that forces a coding agent to falsify its own &amp;#39;done&amp;#39; claim and return one of four verdicts, not reassurance.</description><content:encoded><![CDATA[<p>An adversarial verification agent is a checker with an inverted objective: instead of collecting signals that confirm a claim, it is built to refute it. Prove-It turns that refuter on a coding agent&rsquo;s own &ldquo;done&rdquo; claim, answering with one of four verdicts — PROVEN, FAILED, NOT PROVEN, or BLOCKED — instead of a paragraph of reassurance.</p>
<p>That is the short answer. The rest of this guide explains why the pattern exists, what the four verdicts actually mean, where the skill stops being useful, and how a 12-case self-authored benchmark should be read without overselling it.</p>
<h2 id="what-is-an-adversarial-verification-agent">What Is an Adversarial Verification Agent?</h2>
<p>Most people arrive at this topic assuming &ldquo;adversarial verification&rdquo; is a synonym for code review or a stricter test suite. It is neither, and the distinction matters because it determines what the agent is actually optimizing for.</p>
<table>
  <thead>
      <tr>
          <th>Pattern</th>
          <th>Question it asks</th>
          <th>Who does the work</th>
          <th>Failure mode it catches</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Unit testing</td>
          <td>&ldquo;Does this input produce this output?&rdquo;</td>
          <td>A fixed test suite</td>
          <td>Regressions in covered behaviour</td>
      </tr>
      <tr>
          <td>Self-review</td>
          <td>&ldquo;Does this code look right to me?&rdquo;</td>
          <td>The same agent, same context</td>
          <td>Almost nothing — same premises</td>
      </tr>
      <tr>
          <td>Code review</td>
          <td>&ldquo;What else might be wrong here?&rdquo;</td>
          <td>A human or a peer agent</td>
          <td>Design issues, missing cases</td>
      </tr>
      <tr>
          <td>Adversarial verification</td>
          <td>&ldquo;What would let this pass while the claim is still false?&rdquo;</td>
          <td>A refuter with an inverted objective</td>
          <td>False assurance</td>
      </tr>
  </tbody>
</table>
<p>The adversary&rsquo;s job is not to find more issues. Its job is to answer one question about one frozen claim: is there a world in which the evidence you are holding is true and the claim is still false? That is falsification-first verification, and it is a different search direction from review. Review expands the search space; refutation constrains it to a single proposition and tries to break that proposition specifically.</p>
<p>Prove-It states the reframe directly in its own doctrine: don&rsquo;t try to prove it works, try to prove it doesn&rsquo;t — the same prover-verifier inversion <a href="https://arxiv.org/abs/2407.13692">formalized in OpenAI&rsquo;s prover-verifier games</a>. The skill is deliberately small — four rules and one verdict — and it changes the verification <em>objective</em> rather than the test framework, which is why it ports across stacks, languages, and agent hosts without a runtime dependency.</p>
<h2 id="why-cant-your-ai-agents-done-claim-be-trusted">Why Can&rsquo;t Your AI Agent&rsquo;s &ldquo;Done&rdquo; Claim Be Trusted?</h2>
<p>The case for an external refuter is not philosophical. It is measurable, and the measurement is uncomfortable.</p>
<p>A benchmark of 14 open-source non-reasoning models — <a href="https://arxiv.org/abs/2507.02778">Self-Correction Bench</a>, published as arXiv 2507.02778 and presented at COLM 2026 — found a 64.5% average &ldquo;Self-Correction Blind Spot&rdquo;: these models reliably corrected an error when it was presented as external input, but failed to correct the <em>byte-identical</em> error when it appeared inside their own output. This is not an artifact of artificial test errors. When the models&rsquo; own naturally generated errors were re-presented externally, they caught only 4.3% to 10.8% of them.</p>
<p>The most useful detail in that paper for anyone building agent harnesses is the mitigation. Simply appending the word &ldquo;Wait&rdquo; <a href="https://arxiv.org/pdf/2507.02778">reduced the self-correction blind spot by 89.3%</a>. The capability was already present; it required external activation rather than new capability. A follow-up study, <a href="https://arxiv.org/html/2606.05976v1"><em>The Self-Correction Illusion</em> (arXiv 2606.05976)</a>, put numbers on the same mechanism: relabeling a wrong claim from the agent&rsquo;s own reasoning into an external user message lifted the correction rate by 23 to 93 percentage points across seven model families. A &ldquo;self-distrust&rdquo; prompt that left the claim in place yielded only 0-23% correction, against roughly 70% for the relabel.</p>
<p>The conclusion is structural, not stylistic: a same-session &ldquo;double-check&rdquo; is not a second opinion. It is a consistency check against the agent&rsquo;s own premises. Huang et al. (<a href="https://arxiv.org/abs/2310.01798"><em>Large Language Models Cannot Self-Correct Reasoning Yet</em>, arXiv 2310.01798</a>, ICLR 2024) showed that intrinsic self-correction without external feedback often fails to improve and can actively degrade reasoning accuracy. <a href="https://arxiv.org/abs/2305.11738">CRITIC</a> (arXiv 2305.11738) showed what does work: critiquing that calls external tools rather than re-reading its own prose. And <a href="https://arxiv.org/abs/2407.00215"><em>LLM Critics Help Catch LLM Bugs</em></a> (McAleese et al., OpenAI) established the empirical basis for LLM-as-verifier — purpose-trained critics find bugs in real-world code that human reviewers miss, and their critiques were preferred over human-written ones in evaluation.</p>
<p>Put those four results together and the architecture writes itself: the verifier must be a separate context, must be pointed at the world rather than at the author&rsquo;s reasoning, and must be given an objective that rewards refutation.</p>
<h2 id="what-is-vibe-verification">What Is &ldquo;Vibe Verification&rdquo;?</h2>
<p>&ldquo;Vibe verification&rdquo; is the named anti-pattern this whole skill family exists to interrupt. It is the moment an agent accumulates enough green signals to feel confident and stops looking for ways its conclusion could be false.</p>
<p>It is a real, observable behaviour, and it is easy to recognize once you have a name for it:</p>
<ul>
<li><strong>Vibe verification:</strong> &ldquo;Tests pass. Build is green. Looks good.&rdquo;</li>
<li><strong>Adversarial verification:</strong> &ldquo;What would let these tests pass while the original bug still exists?&rdquo;</li>
</ul>
<p>The second question is not rhetorical. It has concrete answers. The test might be asserting on a mocked collaborator that no longer matches the real one. The assertion might have been weakened — <code>expect(x).toBe(3)</code> quietly edited to <code>expect(x).toBeTruthy()</code> — during the same session that &ldquo;fixed&rdquo; the failure. The suite might be running against a stale build artifact. The test might be skipped, and the skip might be reported as a pass by a runner that only checks the exit code.</p>
<p>None of those require bad faith. They require an agent that stopped searching in the direction of falsity. That is what &ldquo;vibe verification&rdquo; describes, and it is why the fix is a change of search direction rather than a longer prompt saying &ldquo;be careful.&rdquo;</p>
<h2 id="how-does-prove-it-work-four-rules-one-verdict">How Does Prove-It Work? Four Rules, One Verdict</h2>
<p>Prove-It, published by Pablo-aps in August 2026, packages the refuter as an Agent Skill — Anthropic&rsquo;s <a href="https://docs.claude.com/en/docs/agents-and-tools/agent-skills/overview">documented Agent Skills format</a> for modular capabilities (instructions plus metadata plus optional scripts) that load on demand and extend Claude automatically when relevant. The skill implements a four-step loop.</p>
<p><strong>1. DEFINE — freeze a falsifiable claim and its acceptance criteria before inspecting any evidence.</strong> This ordering is load-bearing. If you read the evidence first and write the criteria second, verification degrades into moving the goalposts until the evidence fits. The claim must be specific enough to be wrong: not &ldquo;the endpoint works&rdquo; but &ldquo;a POST to <code>/api/export</code> with a valid token enqueues a job and the job completes with a downloadable artifact.&rdquo;</p>
<p><strong>2. BREAK — produce a ranked falsification plan.</strong> What are the most plausible ways this claim could be false given the evidence I am about to collect? What is the single cheapest decisive check — the one that, if it fails, ends the investigation? The skill asks for the safest decisive check first, not the most dramatic one.</p>
<p><strong>3. VERIFY — verify the outcome, not a proxy for the outcome.</strong> This is where most false assurance is manufactured, and the skill treats proxy-versus-outcome confusion as a first-class defect rather than an edge case.</p>
<p><strong>4. VERDICT — emit exactly one of four values.</strong> No prose hedging, no &ldquo;looks good to me,&rdquo; no confidence percentages.</p>
<p>The skill is explicitly out of scope for a lot of things people assume it does. It ships no runtime dependencies, no hooks, no background process, no telemetry, no MCP server, and no orchestration layer. It is a behavioural guardrail, not a test framework and not a security scanner. Its own prior conclusions are treated as claims to be re-verified, not as evidence — which is the detail that separates it from a session-level reminder to try harder.</p>
<h2 id="what-do-the-four-verdicts--proven-failed-not-proven-blocked--actually-mean">What Do the Four Verdicts — PROVEN, FAILED, NOT PROVEN, BLOCKED — Actually Mean?</h2>
<p>The verdict vocabulary is arguably the real product. Four values, deliberately calibrated, each with a narrow scope.</p>
<table>
  <thead>
      <tr>
          <th>Verdict</th>
          <th>Requires</th>
          <th>Does <em>not</em> claim</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>PROVEN</strong></td>
          <td>A decisive check ran and passed, within frozen acceptance criteria</td>
          <td>Formal correctness, future safety, or behaviour under untested conditions</td>
      </tr>
      <tr>
          <td><strong>FAILED</strong></td>
          <td>A decisive check ran and produced direct contradiction</td>
          <td>That the <em>approach</em> is wrong — only that the claim is</td>
      </tr>
      <tr>
          <td><strong>NOT PROVEN</strong></td>
          <td>Evidence is indirect, incomplete, stale, or narrower than the claim</td>
          <td>That the claim is false</td>
      </tr>
      <tr>
          <td><strong>BLOCKED</strong></td>
          <td>A specific named check cannot run, and the reason is identified</td>
          <td>That the claim is unverifiable in principle</td>
      </tr>
  </tbody>
</table>
<p><strong>NOT PROVEN is the most valuable of the four</strong>, and it is the verdict that collapses into a false &ldquo;done&rdquo; in every naive agent loop. Consider a claim that reads &ldquo;the migration completes cleanly on production data.&rdquo; A test suite against a 100-row fixture does not refute the claim and does not establish it. The honest verdict is NOT PROVEN. An agent with only two options — pass or fail — will inevitably report pass, because the tests are green.</p>
<p><strong>BLOCKED is narrower than it looks.</strong> It is reserved for a check you can name and cannot run: no credentials for the staging environment, a service the check depends on is down. &ldquo;I could not think of a check&rdquo; is NOT PROVEN, not BLOCKED. That distinction exists because BLOCKED creates a visible escalation, and padding it with vague uncertainty would destroy its signal value.</p>
<p><strong>FAILED requires direct contradiction.</strong> An agent that has stopped finding positive signals has not proven failure, and reporting FAILED on indirect evidence is the mirror-image error of reporting PROVEN on indirect evidence. Both are miscalibrated verdicts; both destroy trust in the verifier.</p>
<p>The ordering also prevents the failure mode that makes people abandon verification tools: infinite skepticism. A refuter that always answers NOT PROVEN is useless, which is why Prove-It warns explicitly against inventing unbounded hypothetical gaps after the frozen scope is covered.</p>
<h2 id="how-do-you-install-prove-it-in-30-seconds-claude-code-codex-cursor">How Do You Install Prove-It in 30 Seconds (Claude Code, Codex, Cursor)?</h2>
<p>The install is one command:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-bash" data-lang="bash"><span style="display:flex;"><span>npx skills add Pablo-aps/prove-it
</span></span></code></pre></div><p>The skill is a single file, so a manual install is just as easy — drop <code>SKILL.md</code> into the skills directory your host reads:</p>



<div class="goat svg-container ">
  
    <svg
      xmlns="http://www.w3.org/2000/svg"
      font-family="Menlo,Lucida Console,monospace"
      
        viewBox="0 0 400 57"
      >
      <g transform='translate(8,16)'>
<text text-anchor='middle' x='0' y='4' fill='currentColor' style='font-size:1em'>.</text>
<text text-anchor='middle' x='0' y='20' fill='currentColor' style='font-size:1em'>.</text>
<text text-anchor='middle' x='0' y='36' fill='currentColor' style='font-size:1em'>.</text>
<text text-anchor='middle' x='8' y='4' fill='currentColor' style='font-size:1em'>c</text>
<text text-anchor='middle' x='8' y='20' fill='currentColor' style='font-size:1em'>a</text>
<text text-anchor='middle' x='8' y='36' fill='currentColor' style='font-size:1em'>c</text>
<text text-anchor='middle' x='16' y='4' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='16' y='20' fill='currentColor' style='font-size:1em'>g</text>
<text text-anchor='middle' x='16' y='36' fill='currentColor' style='font-size:1em'>u</text>
<text text-anchor='middle' x='24' y='4' fill='currentColor' style='font-size:1em'>a</text>
<text text-anchor='middle' x='24' y='20' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='24' y='36' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='32' y='4' fill='currentColor' style='font-size:1em'>u</text>
<text text-anchor='middle' x='32' y='20' fill='currentColor' style='font-size:1em'>n</text>
<text text-anchor='middle' x='32' y='36' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='40' y='4' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='40' y='20' fill='currentColor' style='font-size:1em'>t</text>
<text text-anchor='middle' x='40' y='36' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='48' y='4' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='48' y='20' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='48' y='36' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='56' y='4' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='56' y='20' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='56' y='36' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='64' y='4' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='64' y='20' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='64' y='36' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='72' y='4' fill='currentColor' style='font-size:1em'>k</text>
<text text-anchor='middle' x='72' y='20' fill='currentColor' style='font-size:1em'>k</text>
<text text-anchor='middle' x='72' y='36' fill='currentColor' style='font-size:1em'>k</text>
<text text-anchor='middle' x='80' y='4' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='80' y='20' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='80' y='36' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='88' y='4' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='88' y='20' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='88' y='36' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='96' y='4' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='96' y='20' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='96' y='36' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='104' y='4' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='104' y='20' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='104' y='36' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='112' y='4' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='112' y='20' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='112' y='36' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='120' y='4' fill='currentColor' style='font-size:1em'>p</text>
<text text-anchor='middle' x='120' y='20' fill='currentColor' style='font-size:1em'>p</text>
<text text-anchor='middle' x='120' y='36' fill='currentColor' style='font-size:1em'>p</text>
<text text-anchor='middle' x='128' y='4' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='128' y='20' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='128' y='36' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='136' y='4' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='136' y='20' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='136' y='36' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='144' y='4' fill='currentColor' style='font-size:1em'>v</text>
<text text-anchor='middle' x='144' y='20' fill='currentColor' style='font-size:1em'>v</text>
<text text-anchor='middle' x='144' y='36' fill='currentColor' style='font-size:1em'>v</text>
<text text-anchor='middle' x='152' y='4' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='152' y='20' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='152' y='36' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='160' y='4' fill='currentColor' style='font-size:1em'>-</text>
<text text-anchor='middle' x='160' y='20' fill='currentColor' style='font-size:1em'>-</text>
<text text-anchor='middle' x='160' y='36' fill='currentColor' style='font-size:1em'>-</text>
<text text-anchor='middle' x='168' y='4' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='168' y='20' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='168' y='36' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='176' y='4' fill='currentColor' style='font-size:1em'>t</text>
<text text-anchor='middle' x='176' y='20' fill='currentColor' style='font-size:1em'>t</text>
<text text-anchor='middle' x='176' y='36' fill='currentColor' style='font-size:1em'>t</text>
<text text-anchor='middle' x='184' y='4' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='184' y='20' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='184' y='36' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='192' y='4' fill='currentColor' style='font-size:1em'>S</text>
<text text-anchor='middle' x='192' y='20' fill='currentColor' style='font-size:1em'>S</text>
<text text-anchor='middle' x='192' y='36' fill='currentColor' style='font-size:1em'>S</text>
<text text-anchor='middle' x='200' y='4' fill='currentColor' style='font-size:1em'>K</text>
<text text-anchor='middle' x='200' y='20' fill='currentColor' style='font-size:1em'>K</text>
<text text-anchor='middle' x='200' y='36' fill='currentColor' style='font-size:1em'>K</text>
<text text-anchor='middle' x='208' y='4' fill='currentColor' style='font-size:1em'>I</text>
<text text-anchor='middle' x='208' y='20' fill='currentColor' style='font-size:1em'>I</text>
<text text-anchor='middle' x='208' y='36' fill='currentColor' style='font-size:1em'>I</text>
<text text-anchor='middle' x='216' y='4' fill='currentColor' style='font-size:1em'>L</text>
<text text-anchor='middle' x='216' y='20' fill='currentColor' style='font-size:1em'>L</text>
<text text-anchor='middle' x='216' y='36' fill='currentColor' style='font-size:1em'>L</text>
<text text-anchor='middle' x='224' y='4' fill='currentColor' style='font-size:1em'>L</text>
<text text-anchor='middle' x='224' y='20' fill='currentColor' style='font-size:1em'>L</text>
<text text-anchor='middle' x='224' y='36' fill='currentColor' style='font-size:1em'>L</text>
<text text-anchor='middle' x='232' y='4' fill='currentColor' style='font-size:1em'>.</text>
<text text-anchor='middle' x='232' y='20' fill='currentColor' style='font-size:1em'>.</text>
<text text-anchor='middle' x='232' y='36' fill='currentColor' style='font-size:1em'>.</text>
<text text-anchor='middle' x='240' y='4' fill='currentColor' style='font-size:1em'>m</text>
<text text-anchor='middle' x='240' y='20' fill='currentColor' style='font-size:1em'>m</text>
<text text-anchor='middle' x='240' y='36' fill='currentColor' style='font-size:1em'>m</text>
<text text-anchor='middle' x='248' y='4' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='248' y='20' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='248' y='36' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='288' y='4' fill='currentColor' style='font-size:1em'>#</text>
<text text-anchor='middle' x='288' y='20' fill='currentColor' style='font-size:1em'>#</text>
<text text-anchor='middle' x='288' y='36' fill='currentColor' style='font-size:1em'>#</text>
<text text-anchor='middle' x='304' y='4' fill='currentColor' style='font-size:1em'>C</text>
<text text-anchor='middle' x='304' y='20' fill='currentColor' style='font-size:1em'>C</text>
<text text-anchor='middle' x='304' y='36' fill='currentColor' style='font-size:1em'>C</text>
<text text-anchor='middle' x='312' y='4' fill='currentColor' style='font-size:1em'>l</text>
<text text-anchor='middle' x='312' y='20' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='312' y='36' fill='currentColor' style='font-size:1em'>u</text>
<text text-anchor='middle' x='320' y='4' fill='currentColor' style='font-size:1em'>a</text>
<text text-anchor='middle' x='320' y='20' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='320' y='36' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='328' y='4' fill='currentColor' style='font-size:1em'>u</text>
<text text-anchor='middle' x='328' y='20' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='328' y='36' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='336' y='4' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='336' y='20' fill='currentColor' style='font-size:1em'>x</text>
<text text-anchor='middle' x='336' y='36' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='344' y='4' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='344' y='36' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='360' y='4' fill='currentColor' style='font-size:1em'>C</text>
<text text-anchor='middle' x='368' y='4' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='376' y='4' fill='currentColor' style='font-size:1em'>d</text>
<text text-anchor='middle' x='384' y='4' fill='currentColor' style='font-size:1em'>e</text>
</g>

    </svg>
  
</div>
<p>Invocation differs by host:</p>
<table>
  <thead>
      <tr>
          <th>Host</th>
          <th>How to invoke</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Claude Code</td>
          <td><code>/prove-it</code> slash command, or the trigger words below</td>
      </tr>
      <tr>
          <td>OpenAI Codex</td>
          <td><code>$prove-it</code></td>
      </tr>
      <tr>
          <td>Cursor</td>
          <td>Auto-activates on trigger words in the prompt</td>
      </tr>
  </tbody>
</table>
<p>The auto-activation trigger words matter more than they sound. In practice, the skill fires on <strong>prove</strong>, <strong>verify</strong>, <strong>validate</strong>, <strong>confirm</strong>, and <strong>double-check</strong> — which means the cheapest adoption path is not a new command in your workflow but a vocabulary change in the prompts you already write. When you would have typed &ldquo;make sure this works,&rdquo; type &ldquo;verify this and give me a verdict.&rdquo;</p>
<p>A realistic single invocation looks like this:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-text" data-lang="text"><span style="display:flex;"><span>/prove-it Fix the pagination bug in listUsers() — the per_page
</span></span><span style="display:flex;"><span>parameter is ignored when page &gt; 1.
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span>Claim: listUsers({page: 3, per_page: 10}) returns records 21-30.
</span></span><span style="display:flex;"><span>Acceptance: the returned array length is 10 and the first record&#39;s
</span></span><span style="display:flex;"><span>id matches the 21st record in the unfiltered table.
</span></span></code></pre></div><p>Then let it run. The value is not that it finds something every time — it is that the answer arrives as one of four words instead of a paragraph of reassurance.</p>
<h2 id="why-is-a-passing-signal-not-a-passing-outcome">Why Is a Passing Signal Not a Passing Outcome?</h2>
<p>The single most reusable table in this topic is the proxy-versus-outcome mapping. Every row is a signal that feels like proof and is not.</p>
<table>
  <thead>
      <tr>
          <th>Signal you observe</th>
          <th>What it actually establishes</th>
          <th>What still has to be true</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tests are green</td>
          <td>The assertions that ran passed</td>
          <td>The assertions cover the original bug</td>
      </tr>
      <tr>
          <td>Build succeeded</td>
          <td>Compilation finished</td>
          <td>The artifact being tested is this build</td>
      </tr>
      <tr>
          <td>Deploy reported success</td>
          <td>The control plane accepted the rollout</td>
          <td>Every replica runs the new version</td>
      </tr>
      <tr>
          <td>Healthcheck returns 200</td>
          <td>One endpoint answered</td>
          <td>Dependencies and workers are healthy</td>
      </tr>
      <tr>
          <td>HTTP 200 returned</td>
          <td>The request was accepted</td>
          <td>The async operation completed</td>
      </tr>
      <tr>
          <td>No ERROR lines in the log</td>
          <td>No error-level lines were emitted</td>
          <td>No relevant failure was logged at all</td>
      </tr>
      <tr>
          <td>One request succeeded</td>
          <td>That request succeeded</td>
          <td>There is no race under concurrency</td>
      </tr>
  </tbody>
</table>
<p>That last row is why verification skills keep insisting on outcomes. A race condition is not falsified by a single successful request any more than a flaky test is proven stable by one green run.</p>
<p>The 2026 maintainability data explains why this matters more now than it did in 2022. GitClear&rsquo;s <a href="https://www.gitclear.com/the_ai_code_quality_maintainability_gap">analysis of 623 million code changes</a> (2023-2026) found refactoring line moves down 70%, long-term legacy maintenance down 74%, cross-file function calls down 35% — while within-commit copy/paste rose 41%, duplicated code blocks rose 81% (from 40.3 to 73.0 duplicated lines per million changed lines), and error-masking constructs rose 47%. In their earlier 211-million-line study, 2024 was the first year measured where within-commit copy/pasted lines exceeded <em>moved</em> (refactored) lines, and refactoring fell from 21% of changed lines in 2022 to under 10% in 2024.</p>
<p>When duplication rises and refactoring falls, &ldquo;it passed&rdquo; matters less than &ldquo;it still holds.&rdquo; Verification has to check the outcome, because the structural signals that used to correlate with correctness are weakening.</p>
<h2 id="what-is-counterfeit-proof-the-green-signals-that-prove-nothing">What Is Counterfeit Proof? The Green Signals That Prove Nothing</h2>
<p>Counterfeit proof is the practical centerpiece of the skill, and it is a checklist rather than a theory. A construct is counterfeit when it hides, redefines, or removes the behaviour being verified.</p>
<table>
  <thead>
      <tr>
          <th>Construct</th>
          <th>Why it counterfeits proof</th>
          <th>Legitimate version</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><code>test.skip</code> / <code>xit</code></td>
          <td>Reports as a non-failure while testing nothing</td>
          <td>A tracked issue with a named owner</td>
      </tr>
      <tr>
          <td>Weakened assertion (<code>toBe(3)</code> → <code>toBeTruthy()</code>)</td>
          <td>Passes for the wrong reason</td>
          <td>Assert the specific value</td>
      </tr>
      <tr>
          <td>Ignored exit code</td>
          <td>Turns a failing command into a passing pipeline step</td>
          <td>Check the code, fail the step</td>
      </tr>
      <tr>
          <td>Empty <code>catch {}</code></td>
          <td>Converts an exception into silence</td>
          <td>Handle it, or let it propagate</td>
      </tr>
      <tr>
          <td><code>|| true</code> / blanket suppression</td>
          <td>Deletes the failure instead of the cause</td>
          <td>Suppress one specific, cited case</td>
      </tr>
      <tr>
          <td>Hardcoded result</td>
          <td>The test asserts the value it was given</td>
          <td>Compute the expected value independently</td>
      </tr>
      <tr>
          <td>Mock that no longer matches reality</td>
          <td>Verifies the mock, not the system</td>
          <td>Contract test against the real interface</td>
      </tr>
      <tr>
          <td>Unrelated mock substitution</td>
          <td>The wrong collaborator is stubbed</td>
          <td>Stub the actual dependency</td>
      </tr>
      <tr>
          <td>Timeout increased without a reproduced timing cause</td>
          <td>Masks a flake instead of explaining it</td>
          <td>Reproduce the timing failure first</td>
      </tr>
  </tbody>
</table>
<p>None of these are automatically wrong. Mocks are necessary. Suppressions are sometimes correct. Skips are sometimes honest. The point is that each one is <em>evidence against a claim</em> when it hides or redefines the behaviour you are claiming to have verified — and an agent that adds one during the same session it declares success has not verified anything.</p>
<p>A copy-pasteable pre-flight check, adapted from the skill&rsquo;s own doctrine:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-text" data-lang="text"><span style="display:flex;"><span>Before trusting any green signal, ask:
</span></span><span style="display:flex;"><span>[ ] Did I run the thing, or the tests around the thing?
</span></span><span style="display:flex;"><span>[ ] Did anything get skipped, and does the runner count skips as passes?
</span></span><span style="display:flex;"><span>[ ] Did I edit any assertion, mock, timeout, or suppression this session?
</span></span><span style="display:flex;"><span>[ ] Does the artifact under test come from the code I changed?
</span></span><span style="display:flex;"><span>[ ] Is this signal an outcome, or a proxy for an outcome?
</span></span><span style="display:flex;"><span>[ ] What would let this signal be true while the claim is false?
</span></span></code></pre></div><h2 id="what-do-false-done-claims-look-like-in-practice">What Do False &ldquo;Done&rdquo; Claims Look Like in Practice?</h2>
<p>The generic form of each failure is easier to recognize than the abstract rule, so here are the four that recur most often in practice.</p>
<p><strong>The weakened assertion.</strong> An agent fixes a bug, sees <code>expect(count).toBe(12)</code> fail, and changes it to <code>expect(count).toBeGreaterThan(0)</code>. The suite is green. The bug is intact. The falsifying check is not &ldquo;do tests pass&rdquo; but &ldquo;does any assertion in this session&rsquo;s diff constrain the behaviour I claimed to fix?&rdquo;</p>
<p><strong>The stale replica.</strong> A deploy reports success. The rollout controller accepted the change. Two of five replicas are still serving the previous image because the readiness probe passed on a cached layer. The proxy (deploy accepted) is true; the outcome (every replica runs the new version) is false. The decisive check targets the replicas, not the controller.</p>
<p><strong>The wrong-environment logs.</strong> A claim that a handler no longer throws is supported by a clean log file — from the pre-fix window, or from a different environment than the one under test, or from a service whose log level filters the relevant exception. The evidence is real and irrelevant. This is the canonical NOT PROVEN: indirect, stale, or narrower than the claim.</p>
<p><strong>The async export.</strong> An HTTP 200 from the export endpoint proves the request was accepted. It does not prove the job ran, that the artifact exists, or that the file contains the requested rows. The decisive check fetches the artifact and validates the row count against the source — which is exactly the kind of outcome check a refuter demands and a green status code discourages.</p>
<h2 id="how-does-prove-it-compare-to-prove_it-bug-hunt-adversarial-verify-and-production-audit">How Does Prove-It Compare to prove_it, bug-hunt, adversarial-verify, and production-audit?</h2>
<p>This niche is crowded, and one detail tells you how crowded: at least six unrelated repositories named <a href="https://github.com/Pablo-aps/prove-it"><code>prove-it</code></a> shipped between April and September 2026 — Pablo-aps, jkaraml, josharsh, jongouveia, ryanda9910, and dvhthomas. The name converged before the design did. That is evidence the need is real and the pattern is not yet standardized, which is why positioning matters more than ranking here.</p>
<table>
  <thead>
      <tr>
          <th>Tool</th>
          <th>Mechanism</th>
          <th>Enforcement point</th>
          <th>Best fit</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Pablo-aps/prove-it</strong> (Apache-2.0)</td>
          <td>Four-rule loop, one of four verdicts</td>
          <td>Agent&rsquo;s own claim, pre-&ldquo;done&rdquo;</td>
          <td>Portable falsification gate across Claude Code, Codex, Cursor</td>
      </tr>
      <tr>
          <td><strong>searlsco/prove_it</strong> (~198★, MIT)</td>
          <td>Lifecycle hooks that <strong>block</strong> stop/commit until configured tasks pass</td>
          <td>Harness level — cannot be skipped</td>
          <td>Claude Code shops that want mechanical enforcement</td>
      </tr>
      <tr>
          <td><strong>danpeg/bug-hunt</strong> (~146★, MIT)</td>
          <td>Three isolated agents: Hunter → Skeptic → Referee, competing incentives</td>
          <td>Independent review of a diff or project</td>
          <td>Finding bugs, as opposed to adjudicating a claim</td>
      </tr>
      <tr>
          <td><strong>fullo/claude-adversarial-skill</strong> (MIT)</td>
          <td>Chain-of-Verification plus tri-modal confidence scoring</td>
          <td>Protocol-level, multi-domain</td>
          <td>Teams wanting one heavyweight verification framework</td>
      </tr>
      <tr>
          <td><strong>Sahir619/fable-method</strong></td>
          <td>Treats completion reports as hostile testimony; re-executes claimed checks</td>
          <td>Verifying <em>another</em> agent&rsquo;s work</td>
          <td>Judging delegated work where re-execution is possible</td>
      </tr>
      <tr>
          <td><strong>apoorvjain25/production-audit</strong> (~19★, MIT)</td>
          <td>24 lenses, convergence stop when two consecutive sweeps find nothing</td>
          <td>Whole-product audit</td>
          <td>Pre-launch audits where 24 lenses fit the budget</td>
      </tr>
      <tr>
          <td><strong>henchmarketing-rgb/sub-zero-skill</strong></td>
          <td>Separate fresh-context verifier that checks the live world</td>
          <td>Post-deploy outcome check</td>
          <td>Claims that resolve to a URL, screenshot, or git ref</td>
      </tr>
      <tr>
          <td><strong>lucasfcosta/backpressured</strong> (~63★, MIT)</td>
          <td>Four-phase plan → implement → verify → ship loop with gates</td>
          <td>Long unattended runs</td>
          <td>Multi-hour autonomous runs needing process scaffolding</td>
      </tr>
  </tbody>
</table>
<p>The architectural distinction worth internalizing: <code>prove_it</code> enforces verification <em>mechanically</em> at the harness level and cannot be skipped, whereas Prove-It is a portable prompt-level discipline with zero runtime that works across three different hosts. Those solve different problems. If your team is all-in on Claude Code and you want a hook that physically prevents a commit, the hook-based skill is the stronger control. If you want the same discipline available in Codex and Cursor, or you want to verify a claim without installing infrastructure, the single-file skill is the cheaper primitive.</p>
<p><code>bug-hunt</code> and Prove-It also answer different questions. <code>bug-hunt</code> generates and filters findings with three isolated agents and load-bearing scoring incentives (Hunter +1/+5/+10 by severity; Skeptic earns for disproving but pays a 2× penalty for wrongly dismissing a real bug; Referee on symmetric +1/−1 ground-truth framing). Prove-It adjudicates one claim already made. One hunts; the other judges.</p>
<p>Cross-model verification deserves a note here. All-Claude adversarial panels share blind spots, so when the blast radius is high — authentication, cryptography, payments, migrations — adding a cross-vendor finder such as Codex to the panel buys genuine independence rather than more variance. <code>ng/adversarial-review</code> builds this in explicitly by treating agreement <em>across providers</em> as the strongest available signal.</p>
<h2 id="does-it-actually-work-reading-the-12-case-benchmark-honestly">Does It Actually Work? Reading the 12-Case Benchmark Honestly</h2>
<p>The repository ships a <a href="https://github.com/Pablo-aps/prove-it/blob/main/benchmark/README.md">reproducible 12-case benchmark</a> with three positive controls. The published directional run — Codex CLI 0.147.0, gpt-5.6-luna, low reasoning, one run per cell, 2026-08-18 — reports the following.</p>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>Baseline</th>
          <th>With the skill</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Correct verdict</td>
          <td>9/12 (75%)</td>
          <td>12/12 (100%)</td>
      </tr>
      <tr>
          <td>Positive-control accuracy</td>
          <td>2/3 (67%)</td>
          <td>3/3 (100%)</td>
      </tr>
      <tr>
          <td>Decisive-signal recall</td>
          <td>96%</td>
          <td>96%</td>
      </tr>
      <tr>
          <td>False-assurance rate</td>
          <td>0/12</td>
          <td>0/12</td>
      </tr>
      <tr>
          <td>Falsification attempt rate</td>
          <td>12/12</td>
          <td>12/12</td>
      </tr>
  </tbody>
</table>
<p>Read that honestly and the honest reading is: this is a <strong>transparent regression test, not an independent evaluation</strong>. The authors say so themselves, and the reason is straightforward — the cases were authored during the skill&rsquo;s development, by the people who wrote the skill. A 12-case suite with one run per cell cannot separate a real behavioural improvement from overfitting to twelve hand-picked scenarios.</p>
<p>Two further caveats from the repository&rsquo;s own metadata at research time: the project had 9 stars, 0 forks, 0 open issues, and one release, created and last pushed on 2026-08-18, listed on skills.sh. Its nearest neighbors by name and intent — <code>prove_it</code> at 198 stars and <code>bug-hunt</code> at 146 stars — mean this is an early-stage entrant in an already crowded niche. The metric that deserves the most attention is the one that <em>didn&rsquo;t</em> move: decisive-signal recall was flat at 96% in both arms, which suggests the skill is sharpening verdict discipline rather than making the underlying model better at finding evidence.</p>
<p>What the benchmark does establish, and what makes it worth citing, is the method. The results are published with a SHA-256-fingerprinted methodology and a positive-control set that guards against the obvious degenerate strategy — a skill that answers NOT PROVEN to everything would score 0/3 on positive controls and be visibly useless. That design choice is itself a good model for anyone building a verifier.</p>
<p>For the independent evidence, lean on the academic results instead: the 64.5% blind spot, the 23-93 percentage-point relabel lift, CRITIC&rsquo;s tool-interactive results, and the OpenAI critic study. Those were produced by people with no stake in this skill.</p>
<h2 id="what-does-adversarial-verification-cost-and-when-does-it-pay-for-itself">What Does Adversarial Verification Cost, and When Does It Pay for Itself?</h2>
<p>Adversarial verification is not free, and almost every write-up in this genre skips the cost. Three verifiers per finding triples the verification pass. Token cost scales linearly with verifier count, which means a naive &ldquo;adversarially verify everything&rdquo; policy is a budget leak.</p>
<p>The engineering answer is triage plus thresholds:</p>
<ul>
<li><strong>Route selectively.</strong> Run full adversarial review on high-confidence, high-blast-radius findings only. Mechanical checks — lint, typecheck, build, tests — are cheap and should run first and always.</li>
<li><strong>2-of-3 as the code-review floor.</strong> Independent verification is worth more than more variants of the same prompt. Running multiple copies of an identical verifier reduces variance but not bias, so diversify by <em>perspective</em> — correctness auditor, security reviewer restricted to trust boundaries, reproducibility auditor — rather than by count.</li>
<li><strong>Unanimous for security and compliance.</strong> Where a single refutation should block a change, a majority threshold is the wrong instrument.</li>
<li><strong>Tighten when escaped defects surface.</strong> A threshold is a tuning parameter, not a constant. If something reached production that passed verification, the threshold was too loose.</li>
</ul>
<p>The evidence that the budget is justified comes from the failure side. <a href="https://itnerd.blog/2026/09/28/producing-code-has-never-been-easier-but-ai-generated-bugs-and-rising-debugging-workloads-are-slowing-software-delivery">Undo research by Coleman Parkes</a> (July-August 2026, n=300 senior engineering leaders at $250m+ revenue organizations) found that 81% of organizations had a production incident or customer-visible outage in the previous six months attributed to AI coding tools, 93% had a root cause misdiagnosed because of an AI hallucination, and 91% had test escapes or serious defects reach production. In the same survey, 35% of AI-generated code reaches production before engineers fully comprehend it, and engineers spend 16.9 hours per week — 42% of the working week — debugging rather than writing.</p>
<p>Faros AI telemetry across roughly 10,000 developers and 1,255 teams points the same direction: high-AI-adoption teams closed 21% more tasks and produced 98% more pull requests, with PR size up 154%, review time up 91%, bugs up 9% per developer, and an incidents-to-pull-request ratio 242.7% higher. DORA&rsquo;s 2025 report put a number on the same effect: every 25% increase in AI adoption correlated with a 1.5% drop in delivery speed and a 7.2% decrease in stability. The bottleneck moved from writing code to reviewing it. A verification gate is a control on exactly that bottleneck, and it is cheap next to a misdiagnosed incident.</p>
<h2 id="how-do-you-build-your-own-refuter-isolation-default-refuted-fixed-output">How Do You Build Your Own Refuter? Isolation, Default-Refuted, Fixed Output</h2>
<p>If you are building verification into your own harness rather than adopting a skill, five rules generalize. They come from <a href="https://dsplce.co/agentic-engineering/core/adversarial-verification">the clearest editorial treatment of the mechanism available</a> and they are what separate a refuter from another opinion.</p>
<ol>
<li><strong>Pass the claim, not the conversation.</strong> The verifier receives the assertion and its cited evidence — never the authoring agent&rsquo;s reasoning chain. Inheriting that chain means inheriting its biases.</li>
<li><strong>Do not include the first model&rsquo;s rationale or a summary of it.</strong> A summary transmits framing even when it drops conclusions.</li>
<li><strong>Do not disclose where the finding came from.</strong> Authority framing is bias; a finding labeled &ldquo;from a senior reviewer&rdquo; is evaluated differently from an unattributed one.</li>
<li><strong>Default to refutation.</strong> Uncertainty is not a pass. If the evidence does not decide the claim, the verdict is NOT PROVEN.</li>
<li><strong>Fixed output schema.</strong> A contract with a fixed decision rule and a fixed output shape, not a longer essay. A working refuter returns something closer to <code>{&quot;refuted&quot;: true, &quot;evidence&quot;: &quot;&lt;code&gt;&quot;, &quot;locator&quot;: &quot;auth.py:1-10&quot;}</code> than four paragraphs of analysis.</li>
</ol>
<p>The control experiment behind rule five is worth knowing. In that write-up, a deliberately false finding about an empty-token auth bypass — framed as coming from a senior engineer — was correctly <em>rejected</em> by a fresh session, which replied with four paragraphs, a code block, two caveats, and an open question back to the human. The verdict was right and the process was unworkable: at 40 findings it &ldquo;isn&rsquo;t a verification step at all, it&rsquo;s just 40 more things to read.&rdquo; The fix was not more cynicism. It was a contract.</p>
<p>And one meta-rule that the skill family&rsquo;s own doctrine implies: <strong>adversarially verify the verifier.</strong> Isolation, default-refuted, read-only, and frozen acceptance criteria are the safety rails. A verifier that mutates production to obtain proof, or that relaxes the claim after a failed check, is worse than no verifier at all, because it manufactures false assurance with a credible label attached. Read-only by default is not a limitation; it is what makes the verdict trustworthy.</p>
<h2 id="how-big-is-the-2026-trust-gap-90-adoption-versus-24-trust">How Big Is the 2026 Trust Gap? 90% Adoption Versus 24% Trust</h2>
<p>The reason a niche this crowded still has room is the size of the gap between how much developers use AI and how much they believe it.</p>
<p>Google Cloud&rsquo;s <a href="https://cloud.google.com/resources/content/2025-dora-ai-assisted-software-development-report">DORA 2025 <em>State of AI-assisted Software Development</em></a> report, covering nearly 5,000 technology professionals, found 90% of professional developers now use AI at work — up 14% year over year — spending a median of two hours per day with AI tools, with 71% using AI for writing new code. Then the trust paradox: only 24% trust AI-generated output &ldquo;a lot&rdquo; or &ldquo;a great deal,&rdquo; while 30% trust it &ldquo;a little&rdquo; or &ldquo;not at all,&rdquo; and 49% trust it &ldquo;somewhat.&rdquo; More than 80% still report productivity gains. Autonomous agent adoption lags far behind assisted use: only 17% of developers use agent mode daily, while 61% never use it at all.</p>
<p>Stack Overflow&rsquo;s <a href="https://survey.stackoverflow.co/2025/ai">2025 developer survey</a> found the distrust is hardening: 46% of developers actively distrust AI accuracy against 33% who trust it, with only 3% reporting high trust, and distrust up from 31% in 2024. The same survey family shows 96% of developers do not fully trust that AI-generated code is functionally correct, yet only 48% always verify it before committing — and 66% name &ldquo;almost right, but not quite&rdquo; as their top frustration. Sonar&rsquo;s 2026 survey of 1,149 professional developers adds the operational detail: 38% say reviewing AI code takes more effort than reviewing a colleague&rsquo;s, and 61% say AI often produces code that looks correct but is unreliable.</p>
<p>Meanwhile the measured productivity effect is contested in exactly the direction skepticism suggests. METR&rsquo;s <a href="https://arxiv.org/abs/2507.09089">randomized controlled trial</a> had 16 experienced open-source developers complete 246 real tasks 19% <em>slower</em> when allowed to use early-2025 AI tools on their own repositories — while estimating afterwards that AI had made them 20% faster. That is roughly a 39-percentage-point gap between belief and measurement. A February 2026 METR re-run with late-2025 tools still placed returning developers around 18% slower as the central estimate, with a confidence interval spanning −38% to +9%; METR characterizes the evidence for speedup as &ldquo;very weak&rdquo; once selection effects are accounted for.</p>
<p>There is also the benchmark-versus-production gap to keep in view. Claude Opus 4.7 <a href="https://benchlm.ai/benchmarks/sweVerified">leads SWE-bench Verified at 87.6%</a>, but on SWE-bench Pro every top model drops 18-25 points (Opus 4.7 at 64.3%). Roughly 20 points of Verified performance looks like benchmark-specific optimization rather than general code reasoning. When your benchmark number and your production experience disagree, the production experience is data.</p>
<p>A verifier does not close that trust gap by making models better. It closes it by making the <em>claim</em> checkable — which is a process fix, not a prompt trick.</p>
<h2 id="what-does-a-verdict-not-claim-limits-and-scope-discipline">What Does a Verdict Not Claim? Limits and Scope Discipline</h2>
<p>Prove-It&rsquo;s most credible design choice is how little its best verdict asserts. PROVEN is scoped to the frozen acceptance criteria and nothing else. It is not a claim of formal correctness, not a claim about untested inputs, and not a claim about future behaviour. A tool that promised more would be lying.</p>
<p>Five limits worth stating plainly:</p>
<ul>
<li><strong>A verdict is only as good as the frozen claim.</strong> Vague acceptance criteria produce a confident verdict about nothing. Write the criteria before you look at the evidence, or you are verifying your ability to rationalize.</li>
<li><strong>Positive controls are mandatory.</strong> A verifier that always refutes is as useless as one that always confirms, and it is harder to notice because it feels rigorous. Include cases that must come back PROVEN, and treat a failure there as a bug in the verifier.</li>
<li><strong>Verification does not replace review.</strong> It constrains a single proposition; it does not search for the bugs nobody thought to claim were absent.</li>
<li><strong>Scope creep is the silent failure.</strong> If the claim gets easier to prove after a check fails, the verification was theater. Frozen means frozen.</li>
<li><strong>Read-only is non-negotiable.</strong> Any verification step that changes the system under test has invalidated its own evidence.</li>
</ul>
<h2 id="where-should-you-put-the-verification-gate-in-your-agent-loop">Where Should You Put the Verification Gate in Your Agent Loop?</h2>
<p>The cheapest adoption path is not a new tool but a new gate location. Three places pay for themselves fastest.</p>
<p><strong>Before any &ldquo;done&rdquo; is spoken.</strong> This is the default and the highest-value slot. The agent writes the claim and acceptance criteria, runs the refuter, and reports the verdict instead of a summary. The cost is one extra pass; the return is that &ldquo;done&rdquo; becomes a word with a defined meaning.</p>
<p><strong>Before a commit or a merge.</strong> This is where hook-based enforcement outranks a prompt-level skill, because it cannot be skipped. If your team is standardized on one host, <code>prove_it</code>-style hooks are the stronger control. If you are multi-host, use the portable skill and make the verdict a required line in the PR description.</p>
<p><strong>After deploy, against the live world.</strong> The verifier fetches the URL, takes the screenshot, checks the git ref. Only the verifier gets to call a win, and it has never seen the work — only the result. This is the mode where &ldquo;no evidence, no win&rdquo; stops being a slogan: a 200 response with the pricing table actually on screen is evidence; &ldquo;the deploy succeeded&rdquo; is a proxy.</p>
<p>Across all three slots, the same two questions do the work. <em>Is this a signal or an outcome?</em> And <em>what would let this be true while my claim is false?</em></p>
<h2 id="faq-adversarial-verification-agents-invocation-and-verdicts">FAQ: Adversarial Verification Agents, Invocation, and Verdicts</h2>
<p><strong>Is Prove-It a test framework?</strong>
No. It ships no runtime dependencies, no test runner, no hooks, and no orchestration. It is a behavioural guardrail that changes how an agent verifies a claim. You still need tests — the skill&rsquo;s job is to stop you from treating a green test as proof of an unverified outcome.</p>
<p><strong>Which agents does it support?</strong>
Claude Code (<code>/prove-it</code>), OpenAI Codex (<code>$prove-it</code>), and Cursor (auto-activation on trigger words). Installation is <code>npx skills add Pablo-aps/prove-it</code>, or a manual copy of <code>SKILL.md</code> into <code>.claude/skills/prove-it/</code>, <code>.agents/skills/prove-it/</code>, or <code>.cursor/skills/prove-it/</code>.</p>
<p><strong>What is the difference between NOT PROVEN and BLOCKED?</strong>
NOT PROVEN means the evidence you have is indirect, incomplete, stale, or narrower than the claim — the claim remains open. BLOCKED is reserved for a specific named check that cannot run, with the reason identified, such as missing credentials for the environment the check requires. &ldquo;I couldn&rsquo;t think of a check&rdquo; is NOT PROVEN.</p>
<p><strong>Should I use Prove-It or bug-hunt?</strong>
They answer different questions. bug-hunt generates and adversarially filters findings across three isolated agents with competing scoring incentives — it hunts for bugs. Prove-It adjudicates a claim that has already been made, using four rules and four verdicts. If your problem is &ldquo;I don&rsquo;t know what&rsquo;s wrong,&rdquo; hunt. If your problem is &ldquo;the agent says it&rsquo;s fixed and I don&rsquo;t believe it,&rdquo; adjudicate.</p>
<p><strong>Is adversarial verification worth the extra cost?</strong>
Not for everything. Run mechanical checks — lint, typecheck, build, tests — always and cheaply, then route only high-confidence findings and high-blast-radius changes to full adversarial review. A 2-of-3 independent panel is a reasonable floor for code review, unanimous agreement for security and compliance. Given that 93% of organizations surveyed had a root cause misdiagnosed because of an AI hallucination within a six-month window, the question is usually not whether the gate pays for itself but where to place it.</p>
<p><strong>Does the 75% → 100% benchmark prove the skill works?</strong>
No, and the authors do not claim it does. It is a 12-case, single-run, self-authored regression test published with a fingerprinted methodology and three positive controls. Read it as evidence of a reproducible method and of verdict discipline, then lean on the independent academic results — the 64.5% self-correction blind spot and the 23-93 percentage-point relabel lift — for the underlying mechanism.</p>
<p><strong>Where should the verification gate sit in a normal agent workflow?</strong>
Three places pay off fastest: before any &ldquo;done&rdquo; is spoken (the default slot, one extra pass for a verdict instead of a summary), before a commit or merge (where hook-based enforcement outranks a prompt-level skill because it cannot be skipped), and after deploy against the live world (where only a verifier that has never seen the work gets to call a win).</p>
]]></content:encoded></item></channel></rss>