<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>How to Pick the Best Agent Patch on RockB</title><link>https://baeseokjae.github.io/tags/how-to-pick-the-best-agent-patch/</link><description>Recent content in How to Pick the Best Agent Patch on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 06:53:02 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/how-to-pick-the-best-agent-patch/index.xml" rel="self" type="application/rss+xml"/><item><title>Best-of-N Coding Agents With LLM-as-a-Verifier in Oh My Pi</title><link>https://baeseokjae.github.io/posts/skills-omp-pstack-multi-agent/</link><pubDate>Thu, 01 Oct 2026 06:53:02 +0000</pubDate><guid>https://baeseokjae.github.io/posts/skills-omp-pstack-multi-agent/</guid><description>Best-of-N coding agents run N candidates and let an LLM verifier pick the winner. How omp-best-of works, what it costs, and where it fails.</description><content:encoded><![CDATA[<p>Best-of-N coding agents run the same task N times in isolated workspaces, then use an LLM-as-a-verifier to score every full trajectory and select the winning patch. In Oh My Pi, the <code>omp-best-of</code> plugin does this with <code>/best-of --n 5 --apply &quot;Fix the failing authentication test&quot;</code>, defaulting to 3 candidates and selection-only mode.</p>
<p>That is the short answer. The rest of this guide is the part that decides whether the technique saves you money or burns it: what selection can and cannot fix, why the maintainers&rsquo; own benchmark numbers are much smaller than the headline claims circulating about them, and how to choose N without paying for a fourth opinion nobody needed.</p>
<h2 id="what-does-best-of-n-actually-solve--and-what-cant-it">What Does Best-of-N Actually Solve — and What Can&rsquo;t It?</h2>
<p>The case for sampling is arithmetic, not intuition. If a single attempt succeeds with probability <code>p</code>, then the chance that at least one of <code>k</code> independent attempts succeeds is:</p>



<div class="goat svg-container ">
  
    <svg
      xmlns="http://www.w3.org/2000/svg"
      font-family="Menlo,Lucida Console,monospace"
      
        viewBox="0 0 200 25"
      >
      <g transform='translate(8,16)'>
<text text-anchor='middle' x='0' y='4' fill='currentColor' style='font-size:1em'>c</text>
<text text-anchor='middle' x='8' y='4' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='16' y='4' fill='currentColor' style='font-size:1em'>v</text>
<text text-anchor='middle' x='24' y='4' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='32' y='4' fill='currentColor' style='font-size:1em'>r</text>
<text text-anchor='middle' x='40' y='4' fill='currentColor' style='font-size:1em'>a</text>
<text text-anchor='middle' x='48' y='4' fill='currentColor' style='font-size:1em'>g</text>
<text text-anchor='middle' x='56' y='4' fill='currentColor' style='font-size:1em'>e</text>
<text text-anchor='middle' x='72' y='4' fill='currentColor' style='font-size:1em'>=</text>
<text text-anchor='middle' x='88' y='4' fill='currentColor' style='font-size:1em'>1</text>
<text text-anchor='middle' x='104' y='4' fill='currentColor' style='font-size:1em'>-</text>
<text text-anchor='middle' x='120' y='4' fill='currentColor' style='font-size:1em'>(</text>
<text text-anchor='middle' x='128' y='4' fill='currentColor' style='font-size:1em'>1</text>
<text text-anchor='middle' x='144' y='4' fill='currentColor' style='font-size:1em'>-</text>
<text text-anchor='middle' x='160' y='4' fill='currentColor' style='font-size:1em'>p</text>
<text text-anchor='middle' x='168' y='4' fill='currentColor' style='font-size:1em'>)</text>
<text text-anchor='middle' x='176' y='4' fill='currentColor' style='font-size:1em'>^</text>
<text text-anchor='middle' x='184' y='4' fill='currentColor' style='font-size:1em'>k</text>
</g>

    </svg>
  
</div>
<p>At <code>p = 50%</code>, three attempts cover 87.5% of the outcomes and five attempts cover 96.9%. At <code>p = 25%</code>, three attempts only reach 57.8%. At <code>p = 10%</code>, you need twenty attempts to crawl to 87.8%. The math is unforgiving in exactly one direction: samples rescue tasks that are already roughly half-solvable, and do almost nothing for tasks the model cannot approach at all.</p>
<p>The strongest empirical demonstration remains the Large Language Monkeys result, which took a mid-tier open model from 15.9% on SWE-bench Lite with a single attempt to 56% with 250 attempts — above the 43% single-attempt state of the art at the time. That is horizontal scaling beating vertical scaling: more samples from a cheap model outrunning a bigger model used once.</p>
<p>But notice what that result does <em>not</em> say. It measures <strong>oracle@N</strong> — the fraction of tasks where <em>some</em> candidate in the pool was correct. It assumes a perfect selector that always picks the right one. Real deployments do not have that. Coverage is what N buys you; selection is a separate problem, and it is the harder one.</p>
<p>This distinction is where most best-of-N write-ups quietly cheat. A pool with 87.5% oracle@3 and a 50% accurate verifier does not deliver 87.5% of anything. This is why best-of-N with LLM-as-a-verifier is best understood as an architecture decision rather than a benchmark trick: you are not making the model smarter, you are buying a larger pool and then paying a second model to referee it. If the referee is bad, you have added cost and latency to random selection.</p>
<p>There is also a sharp edge on the other end. Best-of-N pays off in the <strong>middle band</strong>, where single-shot success is neither near zero nor near one. If your agent already solves a task 95% of the time, you are paying four extra attempts to be told what you already knew. If it solves it 5% of the time, you are buying lottery tickets. The regime that matters is the 30–70% band, and knowing which band you are in is a prerequisite, not an afterthought.</p>
<h2 id="how-does-llm-as-a-verifier-score-a-trajectory">How Does LLM-as-a-Verifier Score a Trajectory?</h2>
<p>The upstream framework here is <code>llm-verifier</code> (Kwok et al., arXiv 2607.05391, MIT licensed, roughly 3.2k stars on GitHub), installable with <code>pip install llm-verifier</code> and exposing <code>select()</code>, <code>compare()</code> and token accounting.</p>
<p>What separates it from an ordinary LLM judge is granularity and how the score is read. A conventional judge asks a model for a rating and takes the argmax token. LLM-as-a-Verifier instead expects a model to produce a <em>distribution</em> over score tokens and computes the expectation over that full distribution. That single change is what converts a noisy ordinal judgment into a continuous score, which in turn makes paired comparisons stable enough to rank.</p>
<p>The framework&rsquo;s published self-verification table on Terminal-Bench 2.1, using DeepSeek V4 Flash for both generation and verification, is the cleanest illustration of the gap between coverage and selection:</p>
<table>
  <thead>
      <tr>
          <th>Configuration</th>
          <th>Random selection</th>
          <th>Verifier-selected</th>
          <th>Oracle</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Best-of-3</td>
          <td>79.4%</td>
          <td>86.5%</td>
          <td>92.1%</td>
      </tr>
      <tr>
          <td>Best-of-5</td>
          <td>78.7%</td>
          <td>88.0%</td>
          <td>96.6%</td>
      </tr>
  </tbody>
</table>
<p>Read the columns in order. Verifier selection recovers a real chunk of the distance between random and oracle — about 7 points at N=3 — but it does not close it. At N=5 the oracle ceiling is 96.6% and the verifier delivers 88.0%. That 8.6-point gap is the price of an imperfect judge, and it is the number that should anchor your expectations rather than the oracle figure.</p>
<p>Three scaling levers follow from that design: finer scoring granularity, scaling the number of repeated evaluations, and decomposing one holistic verdict into several narrow criteria. The upstream project reports cross-benchmark results of 86.5% on Terminal-Bench V2 (GPT-5.5, best-of-5), 78.2% on SWE-Bench Verified (best-of-3), 87.4% on RoboRewardBench and 73.3% on MedAgentBench (best-of-5).</p>
<p>The framework also requires a backend that returns token logprobs — DeepSeek, vLLM or SGLang, or any OpenAI-compatible server that honors the constraint. That requirement is not incidental; it is the whole reason the <code>logprob</code> backend in <code>omp-best-of</code> has installation prerequisites that the <code>sampled</code> backend does not.</p>
<p>Version 0.2.0 added a prefix-cache optimization that cuts uncached input tokens by roughly 3.4x on trajectory-heavy benchmarks. That is the fix aimed squarely at the cost problem described later, because verifying a long agent trajectory means re-reading the entire thing.</p>
<h2 id="how-do-you-install-omp-best-of-in-oh-my-pi">How Do You Install omp-best-of in Oh My Pi?</h2>
<p><code>omp-best-of</code> (wolfiesch/omp-best-of) is an MIT-licensed Oh My Pi plugin — around 70 stars, 53 commits, last pushed 2026-09-02. It runs N headless OMP candidate sessions in isolated copy-on-write workspaces and ranks full trajectories rather than just final diffs.</p>
<p>Requirements are specific and worth reading before you install:</p>
<ul>
<li><strong>Oh My Pi 17+</strong> (can1357/oh-my-pi, roughly 33.9k stars, shipping daily)</li>
<li><strong>Bun 1.3+</strong></li>
<li><strong>Git</strong> — the isolation mechanism depends on it</li>
<li><strong><code>uv</code></strong> — only for the <code>logprob</code> backend, which runs a pinned <code>llm-verifier==0.2.0</code> sidecar</li>
<li>A <strong>verifier route configured in omp</strong> — the plugin does not invent one for you</li>
</ul>
<p>One operational constraint deserves emphasis because it is a hard gate rather than a warning: <strong>a clean working tree is enforced before any candidate starts.</strong> Uncommitted changes stop the run. That is the correct design — the plugin cannot reason about patches against a dirty baseline — but it means the tool is not something you reach for in the middle of a debugging session with three half-edited files open.</p>
<h2 id="what-does-your-first-run-look-like-from-best-of-to-an-applied-patch">What Does Your First Run Look Like, from /best-of to an Applied Patch?</h2>
<p>The command surface is deliberately small, which is a virtue worth crediting. In-session:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-bash" data-lang="bash"><span style="display:flex;"><span>/best-of --n <span style="color:#ae81ff">5</span> --apply Fix the failing authentication test
</span></span></code></pre></div><p>Standalone, writing a JSON summary to stdout:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-bash" data-lang="bash"><span style="display:flex;"><span>omp-best-of --n <span style="color:#ae81ff">3</span> --apply -- <span style="color:#e6db74">&#34;Fix the failing authentication test&#34;</span>
</span></span></code></pre></div><p>Defaults matter more than the flags here. <code>--n</code> defaults to <strong>3</strong>. The default verifier is <strong>deepseek/deepseek-v4-flash</strong>. And <code>--apply</code> is <strong>off</strong> by default, which means the tool selects a winner and tells you about it without touching your branch. That default is the right one: you want to read the verifier&rsquo;s reasoning on a few real tasks before you let it mutate a working tree.</p>
<p>Two behaviours are worth knowing before the first run:</p>
<ol>
<li>A candidate that <strong>exits non-zero</strong>, or whose patch <strong>cannot be captured</strong>, is excluded before ranking. It never competes.</li>
<li>Both backends score candidates against the same three criteria: <strong>Requirements</strong>, <strong>Correctness</strong>, and <strong>Verification</strong>. Holding the rubric constant across backends is what makes the two comparable at all — a detail competitors often blur.</li>
</ol>
<p>Practically, run selection-only for a week. Compare the verifier&rsquo;s pick against the pick you would have made. If they agree most of the time, <code>--apply</code> is safe. If they don&rsquo;t, you have learned something valuable at zero risk.</p>
<h2 id="which-verifier-backend-should-you-pick-logprob-or-sampled">Which Verifier Backend Should You Pick: logprob or sampled?</h2>
<p>This is the decision that determines your infrastructure bill, and the distinction is architectural rather than cosmetic.</p>
<table>
  <thead>
      <tr>
          <th></th>
          <th><code>logprob</code> backend</th>
          <th><code>sampled</code> backend</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Mechanism</td>
          <td>Continuous-score pivot tournament</td>
          <td>2 sandboxed candidate audits + seeded all-pairs judge</td>
      </tr>
      <tr>
          <td>Requires</td>
          <td>Endpoint returning score-token distributions; honors constrained prefill</td>
          <td>A subscription model route, e.g. <code>openai-codex/gpt-5.6-luna</code></td>
      </tr>
      <tr>
          <td>Infra</td>
          <td><code>uv</code> + pinned <code>llm-verifier==0.2.0</code> sidecar; vLLM/SGLang eligible</td>
          <td>No API credits or local GPU required</td>
      </tr>
      <tr>
          <td>Fidelity to the paper</td>
          <td>Upstream continuous-score method</td>
          <td>Conventional pairwise judge — <strong>not</strong> the paper&rsquo;s method</td>
      </tr>
  </tbody>
</table>
<p>Pick by what your credentials can actually prove. If you have a DeepSeek key or your own vLLM or SGLang endpoint, <code>logprob</code> gives you the method the paper describes. If you are on a subscription coding plan with no API credits, <code>sampled</code> is the only door open to you — and that is fine, as long as you do not call it something it isn&rsquo;t.</p>
<p>The maintainers make this explicit, and their wording is worth quoting verbatim because a lot of secondary coverage drops it: <em>&ldquo;the sampled backend is a separate conventional pairwise judge, not the paper&rsquo;s continuous-score method&rdquo;</em>, and <em>&ldquo;this plugin has not established equivalent reliability&rdquo;</em>. Anyone presenting <code>sampled</code> results as a reproduction of the published method is misreading the README.</p>
<h2 id="what-do-the-maintainers-own-benchmarks-really-show">What Do the Maintainers&rsquo; Own Benchmarks Really Show?</h2>
<p>This is the section most guides omit, and it is the most useful one.</p>
<p>The plugin&rsquo;s own <code>bench/RESULTS.md</code> and README report two results that still stand, against baseline comparisons the maintainers constructed themselves:</p>
<ul>
<li><strong><code>logprob 1</code> evaluation: 5/10 (50.0%) selections correct</strong> against a <em>62.5% random pass@1 baseline</em>. The verifier underperformed random selection on two reused tasks.</li>
<li><strong><code>sampled</code> Luna, 1 round: 13/15 (86.7%)</strong> against a 32.0% random baseline across five reused discriminating pools (ancestor commit <code>06eefab</code>).</li>
</ul>
<p>A 50% selection accuracy on the first configuration is not a rounding error. It is the honest, unflattering number, and the maintainers published it.</p>
<p>Two earlier headlines — <strong>83.3% and 72.2%</strong> — were <strong>withdrawn</strong> because two pools (content-type, http-range) had defective oracles and, once the oracles were corrected, no candidate passed. The apparent successes had been measuring a broken test.</p>
<p>Three findings from the same file matter more than either headline:</p>
<p><strong>Score separation collapses on saturated pools.</strong> Where the task was easy enough that candidates were mostly equivalent, the verifier&rsquo;s score spread fell to <strong>0.000–0.006</strong>. Nearly every comparison was a tie, which means the reported winner was effectively decided by a tie-break rule, not by verification. Re-ranking one pool three times with identical settings and the same seed produced picks #2, #4, #2 — that is not a stable selector, that is a coin with extra steps.</p>
<p><strong>Repeated evaluation is not free accuracy.</strong> On the single pool with genuine separation (0.061), the verifier ranked a <em>failing</em> candidate first at both 1 and 3 evaluations per criterion. Tripling the evaluation count tripled the cost and changed nothing.</p>
<p><strong>Saturation and discrimination pull in opposite directions.</strong> To measure selection accuracy you need pools with headroom <em>and</em> trajectories containing validation evidence. Visible tests supply the second and destroy the first. The maintainers&rsquo; conclusion — harden the hidden contract instead of deleting the visible tests — is the right lesson for anyone building a benchmark, and it applies well beyond this plugin.</p>
<p>The generalizable principle: <strong>before you trust any selector, measure the spread on your own pool.</strong> If your candidates are scoring 0.000–0.006 apart, your verifier is not selecting, it is guessing.</p>
<h2 id="what-does-best-of-n-cost-and-how-do-you-choose-n">What Does Best-of-N Cost, and How Do You Choose N?</h2>
<p>Verification is not a rounding error in the bill. In the plugin&rsquo;s early live runs, verification accounted for <strong>25–51% of total cost</strong>, because the verifier reads long trajectories and its output and reasoning tokens dominate. In one measured run (<code>run C retry-transient</code>), <strong>147,533 of 153,547 output tokens were reasoning</strong>. You are paying for a model to think carefully about work it did not do.</p>
<p>The sampled backend&rsquo;s invocation count is a closed-form cost model:</p>



<div class="goat svg-container ">
  
    <svg
      xmlns="http://www.w3.org/2000/svg"
      font-family="Menlo,Lucida Console,monospace"
      
        viewBox="0 0 256 25"
      >
      <g transform='translate(8,16)'>
<circle cx='168' cy='0' r='6' stroke='currentColor' fill='currentColor'></circle>
<text text-anchor='middle' x='0' y='4' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='8' y='4' fill='currentColor' style='font-size:1em'>n</text>
<text text-anchor='middle' x='16' y='4' fill='currentColor' style='font-size:1em'>v</text>
<text text-anchor='middle' x='24' y='4' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='32' y='4' fill='currentColor' style='font-size:1em'>c</text>
<text text-anchor='middle' x='40' y='4' fill='currentColor' style='font-size:1em'>a</text>
<text text-anchor='middle' x='48' y='4' fill='currentColor' style='font-size:1em'>t</text>
<text text-anchor='middle' x='56' y='4' fill='currentColor' style='font-size:1em'>i</text>
<text text-anchor='middle' x='64' y='4' fill='currentColor' style='font-size:1em'>o</text>
<text text-anchor='middle' x='72' y='4' fill='currentColor' style='font-size:1em'>n</text>
<text text-anchor='middle' x='80' y='4' fill='currentColor' style='font-size:1em'>s</text>
<text text-anchor='middle' x='96' y='4' fill='currentColor' style='font-size:1em'>=</text>
<text text-anchor='middle' x='112' y='4' fill='currentColor' style='font-size:1em'>2</text>
<text text-anchor='middle' x='120' y='4' fill='currentColor' style='font-size:1em'>N</text>
<text text-anchor='middle' x='136' y='4' fill='currentColor' style='font-size:1em'>+</text>
<text text-anchor='middle' x='152' y='4' fill='currentColor' style='font-size:1em'>E</text>
<text text-anchor='middle' x='184' y='4' fill='currentColor' style='font-size:1em'>N</text>
<text text-anchor='middle' x='192' y='4' fill='currentColor' style='font-size:1em'>(</text>
<text text-anchor='middle' x='200' y='4' fill='currentColor' style='font-size:1em'>N</text>
<text text-anchor='middle' x='208' y='4' fill='currentColor' style='font-size:1em'>-</text>
<text text-anchor='middle' x='216' y='4' fill='currentColor' style='font-size:1em'>1</text>
<text text-anchor='middle' x='224' y='4' fill='currentColor' style='font-size:1em'>)</text>
<text text-anchor='middle' x='232' y='4' fill='currentColor' style='font-size:1em'>/</text>
<text text-anchor='middle' x='240' y='4' fill='currentColor' style='font-size:1em'>2</text>
</g>

    </svg>
  
</div>
<p>for N candidates and E evaluations per criterion. Four candidates at one evaluation is 14 invocations; eight candidates is 44. The quadratic term is why raising N is not a linear purchase.</p>
<p>Measured wall-clock and money on small fixtures: roughly <strong>$0.0084–$0.0119 per candidate run</strong>, i.e. <strong>$0.04–$0.05 for a four-candidate pool</strong> and <strong>$0.05–$0.08 including verification</strong>. One four-candidate task took <strong>95s to 693s</strong> end to end.</p>
<p>For comparison at production scale, coding agents on SWE-bench Verified cost roughly <strong>$0.35–$0.75 per instance per attempt</strong>, making best-of-5 <strong>$1.75–$3.75</strong> and best-of-10 <strong>$3.50–$7.50</strong>. Against a loaded developer hour around $150, that is still cheap — but only if the selection is better than choosing at random.</p>
<p><strong>Latency, not money, is the real wall.</strong> Candidates run concurrently within a task, but tasks run sequentially, so wall clock is dominated by the slowest candidate plus the judge pass rather than the sum of all candidates. Latency is max-of-N, not sum-of-N, and shared prompt prefixes are usually billed once, so the marginal attempt costs less than it looks. The 95s–693s range per four-candidate task is the number that decides whether best-of-N belongs in CI or in a nightly job. Task-level parallelism is the obvious missing optimization.</p>
<p>How to choose N, practically:</p>
<ol>
<li><strong>Estimate <code>p</code> first.</strong> Run the task once, ten times, and count. If <code>p</code> is above 0.8, best-of-N is waste. If it is below 0.2, samples will not save you.</li>
<li><strong>Start at N=3</strong>, the default. It captures most of the coverage jump from a 50% base (87.5%) at a quarter of the verification bill of N=5.</li>
<li><strong>Go to N=5 only when you have measured that the spread is real</strong> — that candidates land in distinguishable score bands.</li>
<li><strong>Stop at N=5 unless the task is high-value.</strong> Returns diminish sharply past N=10 while the quadratic judge term keeps growing.</li>
<li><strong>Use adaptive N by estimated difficulty</strong>, and <strong>early-exit the moment a candidate passes your tests.</strong> OpenHands measured early stopping averaging only <strong>1.35 attempts</strong> while adding <strong>+17.7 points over random</strong> — the single best cost-control pattern available.</li>
</ol>
<h2 id="where-does-best-of-n-fail-and-when-should-you-not-use-it">Where Does Best-of-N Fail, and When Should You Not Use It?</h2>
<p>Four failure modes are documented well enough to plan around.</p>
<p><strong>The verifier saturates on convincing-but-wrong self-reports.</strong> The named example is a failing run that claimed <em>&ldquo;4102 entries, zero mismatches&rdquo;</em> when the correct answer was 698. A model writing a confident summary of its own success is not evidence, and a verifier reading that summary can be fooled. <strong>Pure code diffs are the clean signal; terminal self-reports are not.</strong> Treat any verifier that reads self-reported success as structurally compromised.</p>
<p><strong>Plausible code can outrank correct code.</strong> A patch that looks idiomatic and passes the visible tests can beat a correct patch that missed an unobserved contract edge case. This is the flip side of the saturation problem: visible validation evidence makes a candidate <em>look</em> verified to the judge.</p>
<p><strong>Benchmark-trained critics do not transfer.</strong> OpenHands measured critics trained only on benchmark traces at <strong>AUC 0.45–0.48 on production outcomes — worse than random</strong> — versus <strong>0.69</strong> when trained with code-survival supervision. A verifier that looked fine in evaluation can be actively harmful in production. The same source shows a well-trained critic moving Best@8 to <strong>73.8% vs 57.9% random@8</strong>, so the ceiling is high but the floor is below chance.</p>
<p><strong>Planning-first best-of-N is a measured dead end.</strong> Drafting N plans, selecting one, then executing once changed <strong>0 of 5 outcomes</strong> on Terminal-Bench at <strong>3–10x the cost</strong>. Whatever problem best-of-N solves, it is not plan quality.</p>
<p>On safety, the framing that matters most is that <strong>isolation is not a sandbox</strong>. The copy-on-write workspaces prevent patch <em>collisions</em> — two candidates cannot stomp each other&rsquo;s changes. They do not constrain what a candidate can do. Candidates run headless in yolo approval mode with host filesystem and network access. Treat them exactly as you would any unsandboxed subprocess: no production credentials in the environment, no write access to anything that matters, and a container boundary if you can arrange one.</p>
<p>Do not use best-of-N when:</p>
<ul>
<li>The task is already reliably solved (you are buying tie-breaks).</li>
<li>The task is effectively unsolvable by the model (no amount of sampling helps at low <code>p</code>).</li>
<li>The bottleneck is <strong>planning</strong>, not execution.</li>
<li>Your verification budget is capped below a full trajectory read — a truncated verifier is a worse selector than a coin.</li>
</ul>
<h2 id="how-do-you-get-past-the-oraclen-ceiling-repair-and-alternatives">How Do You Get Past the oracle@N Ceiling? Repair and Alternatives</h2>
<p>The sharpest framing in this whole space is that <strong>selection can never beat oracle@N</strong>. Best-of-N selects from a pool; it cannot invent a candidate the pool never contained. This means the marginal value of another point of selection accuracy is bounded by the pool, while the marginal value of <em>improving the winning candidate</em> is not.</p>
<p><code>agent-ultramode</code> (maverick-tr/agent-ultramode) is the competing implementation that acts on this. It runs the task N times in isolated git worktrees for opencode, Claude Code and Grok, and uses the <strong>same model</strong> as the verifier — no cross-model dependency. Its v2 adds <strong>verifier-guided repair on the winner</strong>, kept only if it verifies better, reaching <strong>91.7% on a 24-task SWE-bench Lite slice against its own 87.5% oracle@5 ceiling</strong> by rescuing two tasks no attempt had solved. That is the ceiling being broken rather than approached.</p>
<p>The same project reports <strong>90.4% on the Terminal-Bench 2.1 coding subset (77 tasks)</strong> with a small non-vision flash model at best-of-5, against 89.5% / 89.1% / 88.4% pass@1 for GPT-5.6 Sol / Claude Opus 5 / Grok 4.6 — and frames it honestly as reaching that tier at a fraction of the cost, not as a like-for-like win. Its blended figure is 78.7% base to 87.6%, a <strong>+8.9 point lift</strong>.</p>
<p>A parallel line of work attacks the same problem by changing what gets compared. arXiv 2604.16529 observes that long-horizon coding agents &ldquo;violate the premise&rdquo; that outputs can be directly compared, ranked or refined, and that the real challenge is <em>representing prior experience</em> rather than generating more attempts. Converting each rollout into a structured summary that preserves hypotheses, progress and failure modes — then scaling it via Recursive Tournament Voting (parallel) or Parallel-Distill-Refine (sequential) — moved Claude-4.5-Opus from <strong>70.9% to 77.6%</strong> on SWE-Bench Verified and <strong>46.9% to 59.1%</strong> on Terminal-Bench v2.0. That compression step is what makes ranking long trajectories tractable at all.</p>
<p>Selector engineering, consolidated from the practitioner literature:</p>
<ul>
<li><strong>Randomize candidate order and evaluate both orders.</strong> Position bias is real; evaluating only one order bakes it in.</li>
<li><strong>Strip length cues.</strong> Longer is not better, but judges consistently think it is.</li>
<li><strong>Use a judge from a different model family</strong> than the generator. Self-preference is measurable.</li>
<li><strong>Prefer pairwise tournaments</strong> (N-1 comparisons) over absolute 1–10 scores. Absolute scales drift; comparisons are local and stable.</li>
<li><strong>Decompose the verdict into narrow aspect verifiers</strong> rather than one holistic judgment — this is what the three-criteria rubric is doing.</li>
<li><strong>Calibrate against a labeled gold set</strong> before trusting the selector on anything that matters.</li>
</ul>
<p>A useful multi-signal rubric weighting, if you are building your own: test pass rate 0.35, regression tests 0.35, diff size, linter and type checker for the remainder.</p>
<h2 id="what-is-a-repeatable-best-of-n-workflow">What Is a Repeatable Best-of-N Workflow?</h2>
<ol>
<li><strong>Confirm a clean working tree.</strong> The plugin enforces this; do not fight it.</li>
<li><strong>Estimate <code>p</code> on the task class</strong> with a handful of single runs. If it is outside the 0.2–0.8 band, stop here.</li>
<li><strong>Run selection-only</strong> (<code>--apply</code> off, the default) at N=3.</li>
<li><strong>Inspect the score spread, not just the winner.</strong> Spread under ~0.01 means the pool is saturated and the pick is a tie-break. Fix the task&rsquo;s discriminative power before adding candidates.</li>
<li><strong>Check the pick against your own judgment.</strong> If the verifier disagrees with you, read its reasoning before blaming it — sometimes the unobserved contract edge case is real.</li>
<li><strong>Promote to N=5 and <code>--apply</code></strong> only once selection looks better than your own coin flip.</li>
<li><strong>Add early exit</strong> as soon as a candidate passes your tests. This is the largest available saving.</li>
<li><strong>Read diffs, not self-reports.</strong> Never let terminal output claiming success count as verification evidence.</li>
<li><strong>Sandbox the candidates.</strong> Isolated workspaces are not a security boundary.</li>
<li><strong>Re-measure monthly.</strong> Both the models and the oracles drift; a selector that worked in September is not verified in December.</li>
</ol>
<h2 id="faq">FAQ</h2>
<p><strong>What are best-of-N coding agents in one sentence?</strong>
They run the same coding task N times in parallel isolated workspaces and use a verifier model to score each full trajectory and select the winning patch, trading compute for a higher success rate than a single attempt.</p>
<p><strong>Does best-of-N with LLM-as-a-verifier actually work?</strong>
Partially, and the honest numbers are modest. Upstream reports verifier selection at 86.5% (best-of-3) and 88.0% (best-of-5) against oracles of 92.1% and 96.6% on Terminal-Bench 2.1. The <code>omp-best-of</code> plugin&rsquo;s own standing result on its <code>logprob</code> backend was 5/10 correct selections against a 62.5% random baseline — the maintainers published that unflattering number themselves.</p>
<p><strong>What is the difference between the logprob and sampled backends?</strong>
<code>logprob</code> runs the upstream continuous-score pivot tournament and needs an endpoint that returns score-token logprobs (DeepSeek, vLLM, SGLang). <code>sampled</code> runs a conventional pairwise judge with 2 sandboxed audits plus a seeded all-pairs comparison, and works on subscription model routes with no API credits. The maintainers state explicitly that <code>sampled</code> is not the paper&rsquo;s method and has not established equivalent reliability.</p>
<p><strong>How much does best-of-N cost per task?</strong>
On small fixtures, roughly $0.0084–$0.0119 per candidate, about $0.04–$0.05 for a four-candidate pool and $0.05–$0.08 including verification, with 95s–693s wall clock. Verification alone was 25–51% of live-run cost because the verifier reasons over long trajectories. At production scale, SWE-bench Verified instances run $0.35–$0.75 per attempt, so best-of-5 is $1.75–$3.75.</p>
<p><strong>What is the oracle@N ceiling and why does it matter?</strong>
Oracle@N is the fraction of tasks where at least one candidate in the pool was correct — the theoretical maximum any selector could achieve. Because selection can never beat it, the useful question is not how to select better but how to improve the winner: verifier-guided repair reached 91.7% on a 24-task SWE-bench Lite slice against its own 87.5% oracle@5 ceiling.</p>
]]></content:encoded></item></channel></rss>