<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Agent Harness vs Model Benchmark Variance on RockB</title><link>https://baeseokjae.github.io/tags/agent-harness-vs-model-benchmark-variance/</link><description>Recent content in Agent Harness vs Model Benchmark Variance on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 29 Sep 2026 22:09:23 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/agent-harness-vs-model-benchmark-variance/index.xml" rel="self" type="application/rss+xml"/><item><title>Coding Agent Harness Design and Performance: How Much Does the Harness Actually Matter in 2026?</title><link>https://baeseokjae.github.io/posts/harness-design-coding-agent-performance-2026/</link><pubDate>Tue, 29 Sep 2026 22:09:23 +0000</pubDate><guid>https://baeseokjae.github.io/posts/harness-design-coding-agent-performance-2026/</guid><description>Harness design moves coding-agent Pass@1 by 27.4pp vs 29.4pp for model choice. Component-level evidence from 176 controlled settings.</description><content:encoded><![CDATA[<p>The harness matters roughly as much as the model. Under identical conditions, harness choice moves coding-agent Pass@1 by 27.4 percentage points and model choice by 29.4 points, while the same model under a minimal versus full adapter swings from 19.1% to 73.4%. The question is no longer whether harness design matters — it is which component buys what, for which model, at which context budget.</p>
<h2 id="what-a-coding-agent-harness-actually-is-and-what-it-is-not">What a coding agent harness actually is (and what it is not)</h2>
<p>A coding agent is a model plus a harness. That framing is now standard, but the term itself is overloaded. Martin Fowler&rsquo;s harness engineering article (Boeckeler, 2 April 2026) narrows it to two bounded contexts: the built-in harness that ships inside a coding agent (its context assembly, tool exposure, and execution loop) and the outer harness a team builds around that agent (instructions, linters, tests, review gates, CI rules).</p>
<p>This article is about the first one. Everything below concerns the runtime machinery that wraps a large language model during a coding run:</p>
<ul>
<li><strong>Planning</strong> — whether the agent produces an explicit plan before acting on the repository.</li>
<li><strong>Action space</strong> — whether the agent gets a predefined tool set (file read, file write, search, patch) or just bash.</li>
<li><strong>Context management</strong> — how the conversation is compacted, elided, summarized, and recovered as a long run approaches its window limit.</li>
</ul>
<p>What the harness is not: it is not the model weights, it is not the verifier, and it is not the benchmark. Those are the other two-thirds of the number you read on a leaderboard.</p>
<p>Vendor anatomy confirms the split. The GitHub Copilot in VS Code write-up (15 May 2026) describes the coding harness as four concrete responsibilities — context assembly, tool exposure, the agent loop, and the round-versus-turn distinction. A source-code study of eleven production systems (arXiv:2609.00006, 15 July 2026) analyzed roughly four million lines across Claude Code, Codex CLI, Gemini CLI, Mistral Vibe, OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode, and OpenClaw. It found no agent runtime that imports a general-purpose agentic framework and none that retrieves code with vector embeddings, then distilled 13 observations, 29 recurring design patterns, 18 design recommendations, and a 90-line minimum-viable-harness scaffold. Harnesses are built, not imported.</p>
<h2 id="the-2026-evidence-how-big-is-the-harness-effect-really">The 2026 evidence: how big is the harness effect, really?</h2>
<p>Four independent lines of evidence converged on &ldquo;as large as the model term,&rdquo; and they measured it in different ways.</p>
<table>
  <thead>
      <tr>
          <th>Evidence</th>
          <th>Setup</th>
          <th>Effect size</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>OpenClaw adapter swap (arXiv:2606.12344)</td>
          <td>Same GLM 5.1 backbone, minimal direct-diff adapter vs full adapter, Claw-SWE-Bench 350 tasks</td>
          <td>19.1% → 73.4% Pass@1 (+54.3pp)</td>
      </tr>
      <tr>
          <td>Locked-model split (arXiv:2606.12344)</td>
          <td>OpenClaw × nine models vs five claws × two models</td>
          <td>harness 27.4pp vs model 29.4pp</td>
      </tr>
      <tr>
          <td>Harness-only edits (arXiv:2605.23950)</td>
          <td>Terminal-Bench 2 pass@1, weights frozen</td>
          <td>69.7% → 77.0% (+7.3pp)</td>
      </tr>
      <tr>
          <td>Scaffold-only monitoring (arXiv:2605.23950)</td>
          <td>Third-party tracker, SWE-bench Verified</td>
          <td>up to +15pp (Kimi K2 Thinking), +11pp (GPT-5)</td>
      </tr>
  </tbody>
</table>
<p>The Claw-SWE-Bench paper is the cleanest of these because harness is a controlled experimental variable: 350 tasks across 8 languages and 43 repositories. Five claws running on GLM 5.1 span 60.9% to 73.4% — a 12.5-point spread with the model held fixed. On Qwen 3.6-flash the spread widens to 38.6%–66.0%. Weaker models are more sensitive to harness quality, not less.</p>
<p>The position paper &ldquo;Stop Comparing LLM Agents Without Disclosing the Harness&rdquo; (arXiv:2605.23950) supplies the ranking-instability evidence that makes this a methodological problem rather than a curiosity. Claude Opus 4.5 reaches 45.9% on SWE-bench Pro under the standardized SEAL scaffold, but 55.4% under Claude Code. HAL reports same-model cross-scaffold swings of nearly 48pp on SWE-bench Verified Mini, against a mere 4.9pp spread spanning six frontier models under a standardized scaffold. That inversion — where the harness term swamps the model term — is the empirical core of what that paper calls the Binding Constraint Thesis.</p>
<p>Third-party tracking agrees in the same direction. LangChain moved Terminal-Bench from 52.8% to 66.5% (+13.7pp) by changing only the system prompt, tool choice, and execution flow, with the model fixed. Anthropic reports that resource configuration alone swings scores about 6pp, and that differences below 3pp are noise. Practical consequence: any leaderboard delta under three points is not a finding.</p>
<p>There is also a synthesis worth citing directly. A controlled factorial by Zhang et al., compiled in the Deep Feed&rsquo;s June 2026 analysis, found harness-induced variance exceeding model-induced variance by 7.8× on a SWE-bench subset, with 6 of 9 model-pair rankings reversing under a different scaffold. Model-paper advances in the same period moved scores 2–4 points; harness swaps moved them 6–54.</p>
<h2 id="how-was-the-harness-effect-measured-inside-176-matched-settings-4-models-2-benchmarks">How was the harness effect measured? Inside 176 matched settings, 4 models, 2 benchmarks</h2>
<p>The anchor study for component-level design decisions is &ldquo;An Empirical Study of Harness Design for Coding Agents&rdquo; (arXiv:2609.20804v1, 17 September 2026). Where Claw-SWE-Bench treats the harness as one variable, this study decomposes it.</p>
<p>It holds the execution loop fixed and varies three components across 176 matched experimental settings — 22 settings per model-benchmark pair, spanning 20 context-management settings (five tiers from T0 to T4 at 32k, 64k, 96k, and 128k windows) plus one planning ablation and one bash-only ablation (Section 3.1).</p>
<table>
  <thead>
      <tr>
          <th>Dimension</th>
          <th>Configuration</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Models</td>
          <td>Nemotron-3 30B, 120B, 550B; Mistral-Medium-3.5-128B</td>
      </tr>
      <tr>
          <td>Serving</td>
          <td>Local SGLang, BF16, temperature 0</td>
      </tr>
      <tr>
          <td>Benchmarks</td>
          <td>SWE-Bench Verified (500 human-verified GitHub issues); Terminal-Bench 2.1 (89 command-line tasks)</td>
      </tr>
      <tr>
          <td>Step budget</td>
          <td>300 steps per task maximum</td>
      </tr>
      <tr>
          <td>Output budget</td>
          <td>16,384 output tokens per turn</td>
      </tr>
      <tr>
          <td>Tool result cap</td>
          <td>24,000 characters</td>
      </tr>
  </tbody>
</table>
<p>Temperature zero and a fixed loop make the comparisons matched rather than anecdotal. Note the study&rsquo;s own stated limits up front, because they constrain every number below: single runs per setting, and both the planning and action-space ablations were run only at the T4 tier with a 128k baseline.</p>
<p>The five context tiers are the backbone of the design:</p>
<table>
  <thead>
      <tr>
          <th>Tier</th>
          <th>Mechanism</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>T0</td>
          <td>Unmanaged full context — the control</td>
      </tr>
      <tr>
          <td>T1</td>
          <td>Elision only (truncation/dropping)</td>
      </tr>
      <tr>
          <td>T2</td>
          <td>Elision + <code>recall_event</code> (lossless recovery of elided content)</td>
      </tr>
      <tr>
          <td>T3</td>
          <td>Summarization alone</td>
      </tr>
      <tr>
          <td>T4</td>
          <td>Staged elision-then-summarization</td>
      </tr>
  </tbody>
</table>
<h2 id="finding-1-is-context-management-a-reasoning-upgrade-or-budget-insurance">Finding 1: Is context management a reasoning upgrade or budget insurance?</h2>
<p>This is the most actionable result in the paper, and it is routinely misread. The value of context management is almost entirely a function of how likely your runs are to run out of window.</p>
<p>The managed-minus-T0 success gap on SWE-Bench falls from 35.7pp at 32k to 15.9pp, 5.5pp, and 2.7pp at 64k, 96k, and 128k respectively. On Terminal-Bench the same gap goes 9.5pp, 7.5pp, 4.8pp, 2.8pp (Section 3.2, Tables 3 and 4).</p>
<table>
  <thead>
      <tr>
          <th>Window budget</th>
          <th>SWE-Bench gap (managed − T0)</th>
          <th>Terminal-Bench gap</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>32k</td>
          <td>35.7pp</td>
          <td>9.5pp</td>
      </tr>
      <tr>
          <td>64k</td>
          <td>15.9pp</td>
          <td>7.5pp</td>
      </tr>
      <tr>
          <td>96k</td>
          <td>5.5pp</td>
          <td>4.8pp</td>
      </tr>
      <tr>
          <td>128k</td>
          <td>2.7pp</td>
          <td>2.8pp</td>
      </tr>
  </tbody>
</table>
<p>The mechanism is overflow, and the numbers are unambiguous. The unmanaged T0 window-overflow rate falls from 78.7% to 8.7% on SWE-Bench and from 61.0% to 12.1% on Terminal-Bench as the budget grows. Every managed tier — T1 through T4 — overflows on exactly zero tasks at every budget (Section 3.2, Figure 3).</p>
<p>That is the whole story. Context management is not making the model smarter; it is preventing a catastrophic failure mode where the run blows its window.</p>
<p>The extreme cases show how much of a tight budget is really a harness problem. Nemotron-3 30B scores 0% on SWE-Bench at 32k with no context management (T0) versus 20.6% with elision only (T1); at 128k the same model goes 24.8% (T0) versus 25.0% (T1), and the harness difference nearly vanishes. Nemotron-3 550B on SWE-Bench at 32k goes from 0% at T0 to 58.4% at T3 — a 58.4-point swing from one harness component at a tight budget (Table 3).</p>
<p>The design decision this produces: <strong>buy window capacity or buy compaction machinery, but paying for both is the common mistake.</strong> If you are running at 128k, the marginal value of a sophisticated compaction stack measured against your own baseline is under three points. If you are at 32k, it is the difference between a working agent and a broken one.</p>
<h2 id="finding-2-should-you-elide-summarize-or-build-lossless-recall">Finding 2: Should you elide, summarize, or build lossless recall?</h2>
<p>Two sub-results here, and the second one is a warning about shipping features you never measure.</p>
<p>First, ordering matters. Staged elision-then-summarization (T4) matches the accuracy of T1 through T3 while posting the lowest cost in 7 of 8 model-benchmark panels and the lowest mean cost per task at every window budget (Section 3.2, Figures 4 and 5). The reason is arithmetic: early cheap elision removes material before it can be sent to a costly LLM summarization call. You pay for summarization only on what survived truncation.</p>
<p>Second, the lossless recall feature did not pay off. T2 (elision plus <code>recall_event</code>) beats T1 (elision alone) in 15 settings, loses in 14, and ties in 3, across 32 model-benchmark-window comparisons — an equal-weight mean difference of −0.36pp. More damning than the null result is the invocation data: 36 of 64 T2/T4 settings (56.3%) never called <code>recall_event</code> at all, the median invocation rate is zero, and mean calls per task fall from 0.540 at 32k to 0.069, 0.011, and 0.007 at 64k, 96k, and 128k (Section 3.2 and Table 13).</p>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>Value</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>T2 vs T1 accuracy, mean difference</td>
          <td>−0.36pp</td>
      </tr>
      <tr>
          <td>Settings where T2 beat T1</td>
          <td>15 of 32</td>
      </tr>
      <tr>
          <td>Settings that never invoked <code>recall_event</code></td>
          <td>36 of 64 (56.3%)</td>
      </tr>
      <tr>
          <td>Median <code>recall_event</code> invocations per task</td>
          <td>0</td>
      </tr>
      <tr>
          <td>Mean calls per task at 32k / 64k / 96k / 128k</td>
          <td>0.540 / 0.069 / 0.011 / 0.007</td>
      </tr>
  </tbody>
</table>
<p>The feature was built, documented, and almost never used. If you only looked at outcome metrics, you would never learn that. Instrument invocation rates alongside outcomes.</p>
<p>There is a competing design worth knowing about, because it argues the opposite of &ldquo;summarize more cleverly.&rdquo; CliffCompaction (arXiv:2609.26779, 22 September 2026) cuts cost by up to 50% under a bounded context while maintaining or improving Terminal-Bench performance, and adds over 10pp on Terminal-Bench for less than the cost of two full-context runs. Its rule is strict: only truncate or drop content, never rephrase, and never compact a compaction — prior compacted output is discarded rather than re-summarized.</p>
<p>A third position is to avoid compaction entirely. NVIDIA&rsquo;s NOOA reports 82.2% on SWE-bench Verified with GPT-5.5 using roughly half the tokens and LLM calls of comparison harnesses, and needs no context compaction because tool results pass by reference instead of being serialized into the window (NVIDIA Developer Blog, &ldquo;Six Agent Harness Capabilities for Higher Model Performance&rdquo;). Recall this interacts directly with Finding 1: pass-by-reference and a large window are substitutes for the same failure mode.</p>
<h2 id="finding-3-do-you-actually-need-a-planning-step">Finding 3: Do you actually need a planning step?</h2>
<p>Planning is not universally good or universally bad. Its sign depends on model capability, and the mechanism explains why.</p>
<p>For Nemotron-3 30B, planning adds 11.6pp on SWE-Bench and 4.5pp on Terminal-Bench, at higher cost. For Nemotron-3 550B and Mistral-Medium-3.5-128B, planning cuts SWE-Bench cost by roughly 30% and 32% while success drops 2.0pp and 0.4pp (Section 3.2, Figure 6).</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th>Planning effect on success</th>
          <th>Planning effect on cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Nemotron-3 30B</td>
          <td>+11.6pp (SWE-Bench), +4.5pp (Terminal-Bench)</td>
          <td>higher</td>
      </tr>
      <tr>
          <td>Nemotron-3 550B</td>
          <td>−2.0pp</td>
          <td>≈ −30%</td>
      </tr>
      <tr>
          <td>Mistral-Medium-3.5-128B</td>
          <td>−0.4pp</td>
          <td>≈ −32%</td>
      </tr>
  </tbody>
</table>
<p>The mechanism is in where runs stop. With planning off, 68.6% of Nemotron-3 30B SWE-Bench runs ended without an edit and 58.4% stalled in the Localize phase. With planning on, those collapse to 27.8% and 10.4%. For the strongest models, the turns planning removes are mostly post-edit verification — useful-looking work that was not changing the outcome (Table 6).</p>
<p>One more number reframes the planning debate. Fewer than 3% of runs ended without editing code regardless of planning. Planning is not rescuing otherwise-doomed attempts across the board; it is trimming repetition once the fix already exists, and preventing early abandonment for weak models that cannot find the problem without a scaffold.</p>
<p>The practical rule: <strong>turn planning on when your model abandons tasks before editing; turn it off when your model succeeds anyway and you want the ~30% cost back.</strong></p>
<h2 id="finding-4-should-you-ship-predefined-tools-or-bash-only">Finding 4: Should you ship predefined tools or bash-only?</h2>
<p>The action space is the component with the sharpest crossover, and the crossover point moves with both model capability and task type.</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th>SWE-Bench: tool set vs bash-only</th>
          <th>Terminal-Bench: tool set vs bash-only</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Nemotron-3 30B</td>
          <td>tool set +15.0pp</td>
          <td>tool set +10.1pp</td>
      </tr>
      <tr>
          <td>Nemotron-3 550B</td>
          <td>bash-only +3.6pp, −53% cost</td>
          <td>bash-only +5.6pp, −30% cost</td>
      </tr>
      <tr>
          <td>Mistral-Medium-3.5-128B</td>
          <td>tool set +23.2pp</td>
          <td>bash-only +6.7pp</td>
      </tr>
  </tbody>
</table>
<p>The same Mistral model needs opposite action spaces on two benchmarks. That is the whole point: there is no fixed answer, because the correct action space is a function of the model&rsquo;s shell competence and of whether the task is shell-centric.</p>
<p>The failure mechanism for bash-only on weak models is instructive. 66% of Nemotron-3 30B bash-only Terminal-Bench trajectories terminate after out-of-interface tool emissions — the model emits learned tool-call patterns that the bash registry cannot resolve, and the run dies. Average trajectory length collapses from 71 turns to 15. On models that can drive a shell, the same restriction produces the opposite effect: bash-only trajectories issue 32% fewer calls on SWE-Bench and 24% fewer on Terminal-Bench, consistent with denser composite shell commands (Section 3.2).</p>
<p>The minimalism baseline makes the ceiling visible from the other direction. mini-SWE-agent is roughly 100 lines of Python with bash as its only action space. It posted about 65% on SWE-bench Verified in mid-2025 and its own docs report 74%+ with newer frontier models. In the Lita comparison it trails OpenHands by 3.2pp on Claude Sonnet 4 (64.8 vs 68.0) but by only 0.2pp on Claude Opus 4 (67.6 vs 67.8). Scaffolding&rsquo;s marginal value shrinks as model capability rises — which is exactly the crossover the ablation study measures.</p>
<h2 id="why-does-each-component-work-the-trajectory-level-mechanism">Why does each component work? The trajectory-level mechanism</h2>
<p>Scores tell you that a component helps. Trajectory shape tells you why, which is what you need when debugging your own harness. The anchor study&rsquo;s Section 4 gives each component a distinct signature:</p>
<table>
  <thead>
      <tr>
          <th>Component</th>
          <th>Trajectory signature</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Context management</td>
          <td>Extends execution trajectories — median length 20–30 turns unmanaged vs roughly 50–180 managed at 32k — without changing phase ordering or behavior</td>
      </tr>
      <tr>
          <td>Planning</td>
          <td>Changes where trajectories stop (early-abandonment rate and Localize stalls)</td>
      </tr>
      <tr>
          <td>Action space</td>
          <td>Changes the granularity at which code is written (call density and composite command use)</td>
      </tr>
  </tbody>
</table>
<p>Context management does not make the agent reason differently. It lets the same reasoning run longer before the window kills it — consistent with Finding 1&rsquo;s overflow numbers. Planning does not change what the agent does; it changes whether the agent gets to the point of doing it. Action space does not change whether the agent acts; it changes how many turns each action costs.</p>
<p>This is also why &ldquo;the harness effect&rdquo; is better understood as sabotage capacity. A weak harness raises per-step difficulty repeatedly across a long loop, so the ceiling is set by how badly the harness obstructs a capable model. As that easy sabotage gets engineered out, expect the harness term to shrink — so build measurement capability rather than a permanent &ldquo;best harness&rdquo; claim.</p>
<h2 id="how-do-you-match-harness-design-to-model-task-type-and-budget">How do you match harness design to model, task type, and budget?</h2>
<p>Everything above collapses into a decision procedure. Use it as a starting configuration, then validate on your own workloads.</p>
<table>
  <thead>
      <tr>
          <th>Situation</th>
          <th>Recommended harness configuration</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tight window (≤32k), any model</td>
          <td>Managed context, T4 staged elision-then-summarization. Overflow risk dominates; this is the highest-ROI change available</td>
      </tr>
      <tr>
          <td>Large window (≥96k), managed baseline exists</td>
          <td>Skip sophisticated compaction. The measured gap is 2.7–2.8pp; spend the engineering elsewhere</td>
      </tr>
      <tr>
          <td>Weak model that abandons tasks early</td>
          <td>Planning ON (+11.6pp SWE-Bench for the 30B), predefined tool set (+15.0pp)</td>
      </tr>
      <tr>
          <td>Capable model that succeeds without help</td>
          <td>Planning OFF for ~30% cost back; bash-only for 30–53% cost reduction</td>
      </tr>
      <tr>
          <td>Shell-centric tasks (CLI, infra, build systems)</td>
          <td>Bash-only once the model can drive a shell reliably</td>
      </tr>
      <tr>
          <td>Repository-shaped code edits (SWE-Bench style)</td>
          <td>Predefined tool set unless the model is bash-competent</td>
      </tr>
      <tr>
          <td>You want to cut cost without losing accuracy</td>
          <td>Truncate/drop instead of summarizing (CliffCompaction), and prefer staged elision before summarization</td>
      </tr>
      <tr>
          <td>Your runs degrade as they get longer</td>
          <td>Inspect context policy and trajectory length before blaming the model</td>
      </tr>
  </tbody>
</table>
<p>Two configuration principles cut across the table. First, <strong>cost belongs on the same chart as accuracy</strong> — bash-only&rsquo;s 30–53% cost reduction for capable models only exists in the cost column, and an accuracy-only reading of the same data produces a different and worse decision. Second, <strong>every component has an optimal budget</strong>, and paying for the same insurance twice — a 128k window and a full compaction stack — is the single most common design error the 2026 evidence exposes.</p>
<p>A supporting layer sits outside these three components but changes the same numbers: the outer harness of linters, tests, and structural checks. Fowler&rsquo;s split into guides (feedforward controls) and sensors (feedback controls) is the right vocabulary, and the strongest sensors emit LLM-optimized signals — linters that include self-correction instructions, not just error codes. The concept of harnessability is the corollary: strongly typed languages, clear module boundaries, and opinionated frameworks raise an agent&rsquo;s chance of success before any runtime change is made. OpenAI&rsquo;s harness-engineering report describes layered architecture enforced by custom linters and structural tests, plus recurring &ldquo;garbage collection&rdquo; passes to remove drift.</p>
<h2 id="how-do-you-measure-your-own-harness-without-fooling-yourself">How do you measure your own harness without fooling yourself?</h2>
<p>The measurement problems in this literature are as instructive as the results, because they are the mistakes teams repeat internally.</p>
<p><strong>Disclose model, harness, context budget, and step budget together.</strong> A published number is model × harness × verifier, and two of those three are usually unreported. The position paper&rsquo;s core complaint (arXiv:2605.23950) is that leaderboards report each model under its own best-tuned scaffold — which is neither a locked-harness comparison nor a factorial decomposition. Every ranking it examines is unstable under that regime.</p>
<p><strong>Know which regime you are in.</strong></p>
<table>
  <thead>
      <tr>
          <th>Regime</th>
          <th>What varies</th>
          <th>What it tells you</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Locked harness</td>
          <td>Model only</td>
          <td>How models compare under one scaffold — not how they compare in general</td>
      </tr>
      <tr>
          <td>Factorial decomposition</td>
          <td>Harness × model, crossed</td>
          <td>How much variance each factor and interaction explains</td>
      </tr>
  </tbody>
</table>
<p>Zhang et al.&rsquo;s 7.8× harness-over-model variance result is a factorial finding and is not transferable from a locked-harness leaderboard.</p>
<p><strong>Remember the verifier is the third term.</strong> UTBoost found SWE-bench tests weak enough that wrong patches passed, with corrections affecting 40.9% of SWE-bench Lite entries and 24.4% of Verified entries (Yu et al., ACL 2025). Roughly two in five of the gradings you have been comparing may have been wrong. If your harness work is graded by a suite you have never audited, you are optimizing a noisy objective.</p>
<p><strong>Respect the noise floor.</strong> Anthropic&rsquo;s report that differences below about 3pp are noise is a useful default. The anchor study&rsquo;s own limits point the same way: single runs per setting, and planning and action-space ablations only at T4 with 128k. Run repeats before you ship a component based on a two-point delta.</p>
<p><strong>Instrument invocation, not just outcomes.</strong> The <code>recall_event</code> result is the canonical case: an entire feature shipping to production with a median invocation rate of zero.</p>
<p><strong>Use trajectory shape as a diagnostic.</strong> Length for context problems, stop-location for planning problems, call density and command size for action-space problems. A score tells you something changed; the trajectory tells you which component changed it.</p>
<h2 id="what-does-the-evidence-not-settle-limits-contradictions-and-open-questions">What does the evidence not settle (limits, contradictions, and open questions)?</h2>
<p>The honest position is that the 2026 literature is strong on magnitudes and weak on universals.</p>
<ul>
<li><strong>No universal best harness.</strong> The anchor study explicitly limits its claims to the components and implementations it tested, with single runs per setting and planning/action-space ablations only at the T4/128k baseline. Do not generalize its tier results to tiers you did not test.</li>
<li><strong>The effect size is probably inflated by immaturity.</strong> Harness-induced variance exceeds model-induced variance by 7.8× in one factorial, but that is a statement about current harnesses. Harness effects are large precisely because easy sabotage is still shipping.</li>
<li><strong>Compaction design is genuinely unresolved.</strong> LLM summarization (the anchor study&rsquo;s T3/T4), truncate-only CliffCompaction, and pass-by-reference NOOA represent three incompatible philosophies, each with supporting numbers. The disagreement is about whether losing information is worse than paying to compress it.</li>
<li><strong>Benchmark coverage is narrow relative to practice.</strong> SWE-Bench Verified and Terminal-Bench 2.1 are repository-patching and shell tasks. Frontend, data, and multi-repo agent work are barely represented.</li>
<li><strong>The cost axis is under-tabulated across studies.</strong> The anchor study is unusually good here; most harness comparisons still report only success rates, which is why bash-only&rsquo;s 30–53% savings stays invisible in casual reading.</li>
<li><strong>Vendor and third-party numbers come from different regimes.</strong> The 19.1% → 73.4% adapter swap, the 27.4pp vs 29.4pp split, LangChain&rsquo;s +13.7pp, and Kimi K2 Thinking&rsquo;s +15pp are not a single controlled ladder. Treat them as convergent signals about magnitude, not as a comparable scale.</li>
</ul>
<h2 id="faq">FAQ</h2>
<p><strong>Is the harness more important than the model?</strong></p>
<p>It depends on the comparison regime, and both answers are correct in their own regime. Under a locked harness, the model term dominates. In a factorial decomposition over immature harnesses, the harness term dominates — one controlled factorial reports harness-induced variance exceeding model-induced variance by 7.8× on a SWE-bench subset, with 6 of 9 model-pair rankings reversing under a different scaffold. Under fixed models, harness choice moved Pass@1 by 27.4pp and model choice by 29.4pp (arXiv:2606.12344). Public leaderboards do not report either regime cleanly, because they run each model under its own best-tuned scaffold.</p>
<p><strong>Should I keep my context window large or invest in compaction?</strong></p>
<p>The measured margin shrinks to 2.7–2.8pp by 128k on the reported benchmarks, so a large window genuinely substitutes for part of the machinery. But the two are not equivalent in failure terms: managed tiers had zero overflow failures at every budget while unmanaged 128k still lost 8.7–12.1% of tasks to window overflow. If your workload already fits comfortably, buy window and skip the stack. If you are anywhere near the limit, manage the context.</p>
<p><strong>Do I need a plan step?</strong></p>
<p>Only if the model abandons tasks too early. Planning added 11.6pp on SWE-Bench for Nemotron-3 30B by cutting runs that ended without an edit from 68.6% to 27.8% and Localize-phase stalls from 58.4% to 10.4%. For stronger models it mostly trims post-edit verification, cutting cost by roughly 30% (Nemotron-3 550B) and 32% (Mistral-Medium-3.5-128B) while success drops 2.0pp and 0.4pp. Test both; the sign depends on your model.</p>
<p><strong>Should I give the agent predefined tools or just bash?</strong></p>
<p>Predefined tools for models with weak shell control — up to 15.0pp on the smallest model, where 66% of bash-only Terminal-Bench trajectories died after out-of-interface tool emissions. Bash-only for bash-capable models, especially on command-line-centric tasks: 3.6–6.7pp better with 30–53% lower cost. Mistral-Medium-3.5-128B needs the full tool set on SWE-Bench (+23.2pp) and bash-only on Terminal-Bench (+6.7pp), so validate per task type rather than per model.</p>
<p><strong>Is lossless context recall worth building?</strong></p>
<p>In the controlled study, no. Elision plus <code>recall_event</code> beat elision alone in 15 settings, lost in 14, tied in 3, for an equal-weight mean difference of −0.36pp. 56.3% of eligible settings never invoked it, the median invocation rate was zero, and mean calls per task fell from 0.540 at 32k to 0.007 at 128k. That is a null result plus a usage collapse — but the deeper lesson is to instrument invocation rates before shipping features like this at all.</p>
<h2 id="where-should-you-go-next">Where should you go next?</h2>
<p>The component-level view is what makes harness work tractable. Once you know that context management is overflow insurance, planning is an abandonment scaffold, and action space is a capability-dependent cost lever, you can change one thing at a time and measure it. Start with the guide to <a href="/posts/ai-harness-engineering-guide-2026/">AI harness engineering</a> for the broader practice, then browse the <a href="/posts/awesome-ai-harness-curated/">curated harness collection</a> and the <a href="/posts/swe-bench-coding-benchmarks-guide-2026/">2026 coding benchmark guide</a> to see how these numbers are produced. If your immediate problem is long-run degradation, <a href="/posts/claude-code-context-management-2026/">Claude Code context management</a> and <a href="/posts/agentic-workflow-context-management-2026/">agentic workflow context management</a> cover the operational side. For harnesses that adapt themselves, see <a href="/posts/proteus-self-evolving-agent-harness-2026/">Proteus self-evolving harness</a>, <a href="/posts/trueforge-open-source-agent-harness/">TrueForge</a>, <a href="/posts/openharness-universal-agent-harness-2026/">OpenHarness</a>, and the <a href="/posts/deepseek-harness-handbook-2026/">DeepSeek harness handbook</a>.</p>
]]></content:encoded></item></channel></rss>