<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Coding Agent Response Clarity on RockB</title><link>https://baeseokjae.github.io/tags/coding-agent-response-clarity/</link><description>Recent content in Coding Agent Response Clarity on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 00:26:26 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/coding-agent-response-clarity/index.xml" rel="self" type="application/rss+xml"/><item><title>Nopus Coding Agent Review: Deterministic Prose Checks for Clearer Responses</title><link>https://baeseokjae.github.io/posts/nopus-deterministic-prose-checks/</link><pubDate>Thu, 01 Oct 2026 00:26:26 +0000</pubDate><guid>https://baeseokjae.github.io/posts/nopus-deterministic-prose-checks/</guid><description>Nopus is a deterministic prose checker for coding agents: it measures a finished response offline and asks for one clearer rewrite. No LLM judge.</description><content:encoded><![CDATA[<p>Nopus is a deterministic prose checker for coding agents: it measures a finished response with packaged English word data, and if several signals cross their thresholds it asks the same agent for exactly one clearer rewrite. There is no LLM judge, no network call, and no retry loop.</p>
<p>That is the whole product. <code>nopus</code> (npm <code>@syzom/nopus</code>, MIT, by GitHub user Vistyy) shipped on 2026-08-15 and reached its current feature set in five days across 14 commits. It supports exactly three hosts — Pi, Claude Code and Codex — and it does one thing: after an assistant finishes answering, it looks at the prose and decides, reproducibly, whether that prose was too hard to read. If so, the agent is told once to say it more plainly.</p>
<p>This review is about whether that narrow idea is worth installing in 2026, what its numbers actually mean, and where the many third-party write-ups about it have already started inventing features that do not exist.</p>
<h2 id="what-is-nopus-and-what-does-it-do-to-a-coding-agent">What Is Nopus, and What Does It Do to a Coding Agent?</h2>
<p>Nopus is a response-quality hook, not a code-quality tool. It does not lint the code your agent wrote, it does not check imports, and it does not review your diff. It reads the <em>prose around</em> the code — the explanation, the plan, the summary — and intervenes only when that prose is measurably complex.</p>
<p>Mechanically it is three phases. First, prose extraction: the response is stripped of code blocks, inline code, URLs, file paths and table rows, so identifiers and snippets never enter the measurement. Second, multi-metric analysis: the remaining words are scored against packaged lexical tables. Third, conditional rewrite injection: if the policy fires, a single rewrite instruction is injected into the same agent&rsquo;s conversation.</p>
<p>The repository describes the target as answers that &ldquo;disappear into abstract LLM babble&rdquo; — the specific complaint being long load-bearing paragraphs, abstract vocabulary and overloaded phrases from recent models. The design constraint is that the decision must be reproducible and offline, which is the opposite of asking a second model to grade your prose.</p>
<h2 id="what-does-deterministic-prose-check-actually-mean">What Does &ldquo;Deterministic Prose Check&rdquo; Actually Mean?</h2>
<p>It means the accept/rewrite decision is an arithmetic comparison over lookup tables, not a model judgment. Nothing is sampled, nothing is temperature-dependent, and no tokens are spent deciding whether to rewrite. The same prose at the same sensitivity always produces the same verdict — you can verify that by running the policy twice, which you cannot do with an LLM critic.</p>
<p>The linguistic data is real and pinned. Word rarity comes from SUBTLEX-US conversational frequencies (Brysbaert &amp; New, DOI 10.3758/BRM.41.4.977, ISC licensed) and Norvig&rsquo;s Google Web Trillion Word Corpus counts. Abstractness comes from Brysbaert, Warriner and Kuperman concreteness ratings (DOI 10.3758/s13428-013-0403-5), where &ldquo;database&rdquo; scores high concreteness and &ldquo;paradigm&rdquo; scores low. Technical terms are down-weighted using The Carpentries Glosario (CC-BY-4.0). The style-cue list is checksum-pinned to claudisms.ai (SHA-256 <code>f4a09fa8...</code>).</p>
<p>The data is the heavyweight part of the package. <code>data/broad-web-word-counts.json</code> alone is 20.9 MB, conversational frequencies are 4.4 MB, concreteness ratings are 2.5 MB — roughly 28 MB of tables shipped inside an 11.9 MB plugin chunk, against only about 1,087 lines of source across 13 files plus ~1,005 lines of tests. You are installing a lexicon with a small engine attached.</p>
<h2 id="the-seven-measurements-and-six-rewrite-signals">The Seven Measurements and Six Rewrite Signals</h2>
<p>Nopus publishes its decision logic rather than hiding it. Seven quantities are measured:</p>
<table>
  <thead>
      <tr>
          <th>Measurement</th>
          <th>What it captures</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Uncommon wording</td>
          <td>Share of words rare in conversational English</td>
      </tr>
      <tr>
          <td>Very uncommon wording</td>
          <td>Share of words rare even against broad web counts</td>
      </tr>
      <tr>
          <td>Abstract vocabulary</td>
          <td>Share of words rated low on human concreteness norms</td>
      </tr>
      <tr>
          <td>Abstract sentences</td>
          <td>Sentences whose abstract-word ratio is high</td>
      </tr>
      <tr>
          <td>Noun/modifier stacks</td>
          <td>Compressed strings of 3+ nouns and modifiers</td>
      </tr>
      <tr>
          <td>Phrase load</td>
          <td>Complex phrases per 100 words</td>
      </tr>
      <tr>
          <td>Formulaic style cues</td>
          <td>Distinct entries from the packaged claudism list</td>
      </tr>
  </tbody>
</table>
<p>Those feed six independent signal paths in <code>src/policy/decide-rewrite.ts</code>: sustained-abstractness, combined-complexity, stacked-phrasing, concentrated-complexity, pervasive-complexity, and style-cues. A rewrite fires when <strong>any one</strong> path is true.</p>
<p>The thresholds are public and numeric. At medium sensitivity, stacked-phrasing fires when the abstract ratio is at least 0.50, there are 3 or more noun stacks, and phrase load is at least 2.5 per 100 words. At high sensitivity the same path loosens to 0.48 / 3 / 2.0. That is unusually auditable for a tool this small: you can read exactly when your agent will be interrupted instead of trusting a black box.</p>
<p>Anti-over-triggering is built into the same file. Most paths require several measurements to cross together, a single rare word or dense phrase does not normally fire, and style cues need either two distinct cues or one cue supported by enough uncommon wording. The README is explicit that technical terms survive end to end: &ldquo;InteractiveSessionHost stays InteractiveSessionHost&rdquo;.</p>
<h2 id="how-sensitive-is-nopus-the-53--99--186-numbers">How Sensitive Is Nopus? The 5.3% / 9.9% / 18.6% Numbers</h2>
<p>Three sensitivity profiles are published with observed rewrite rates measured on the author&rsquo;s own corpus:</p>
<table>
  <thead>
      <tr>
          <th>Sensitivity</th>
          <th>Rewrites observed</th>
          <th>Rate</th>
          <th>Corpus</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Low</td>
          <td>284</td>
          <td>5.3%</td>
          <td>5,337 completed Pi responses</td>
      </tr>
      <tr>
          <td>Medium (default)</td>
          <td>531</td>
          <td>9.9%</td>
          <td>5,337 completed Pi responses</td>
      </tr>
      <tr>
          <td>High</td>
          <td>995</td>
          <td>18.6%</td>
          <td>5,337 completed Pi responses</td>
      </tr>
  </tbody>
</table>
<p>The corpus is 5,337 unique completed Pi assistant responses drawn from 515 session files, extracted 2026-08-14. The author labels these rates &ldquo;a rough comparison because results vary by agent and task&rdquo;, and that caveat matters: the corpus is deliberately private. Files are mode 0700, artefacts 0600, and nothing is committed. The repo ships only a public regression fixture of frozen scalar measurements exercising &ldquo;more than 5,000 policy inputs&rdquo; — no response text, so no peer reproduction of the headline rates is possible from the repository alone.</p>
<p>Read those numbers as a design calibration, not as a benchmark you can check. What they do tell you honestly is the shape of the trade-off: one response in ten gets rewritten at the default setting, and roughly one in five at high.</p>
<h2 id="what-does-the-agent-actually-receive">What Does the Agent Actually Receive?</h2>
<p>When the policy fires, the agent does not get a vague &ldquo;be clearer&rdquo; nudge. <code>constructRewriteRequest</code> injects the specific evidence into conversation history — the measurements that crossed, with examples drawn from the agent&rsquo;s own response — plus an instruction to rewrite it more plainly.</p>
<p>Two optional behaviours change the feel of that interaction. <strong>Extra-simple mode</strong> (<code>extraSimple</code>, or <code>/nopus extra-simple on</code>) pushes the rewrite further toward short sentences. <strong>Hide original response</strong> is on by default in Pi (<code>pi.hideOriginalResponse=true</code>), and it removes the rejected response from the terminal transcript while leaving it in the session and model history.</p>
<p>That second setting is the one to think about. nopus does not undo the answer; it hides the first draft from the human reading the terminal. The agent&rsquo;s reasoning about your task is unchanged — you are buying readability, not correctness, and you are creating a small deliberate divergence between what you see and what the model still carries.</p>
<p>The rewrite-model evaluation (2026-08-16, <code>openai-codex/gpt-5.6-luna</code>, medium thinking) is the most concrete evidence that the intervention works: of 12 historical branches tried, 4 originals were selected by the medium policy. The normal rewrite passed the medium policy in 2 cases, the extra-simple rewrite in 3. Extra-simple compressed three substantial examples from 279 words to 45, 287 to 106, and 435 to 156. The author still notes that human review of what the rewrite omitted remains required — which is exactly the right caveat for a tool that shortens answers.</p>
<h2 id="how-do-you-install-nopus-on-pi-claude-code-and-codex">How Do You Install Nopus on Pi, Claude Code, and Codex?</h2>
<p>Three hosts, three install paths, one Node requirement. Everything needs Node.js 22+ and <code>node</code> on PATH.</p>
<table>
  <thead>
      <tr>
          <th>Host</th>
          <th>Install</th>
          <th>Mechanism</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Pi</td>
          <td><code>pi install npm:@syzom/nopus</code></td>
          <td>Extension with lifecycle hooks; hides the rejected response by default</td>
      </tr>
      <tr>
          <td>Claude Code</td>
          <td><code>/plugin marketplace add Vistyy/nopus</code> then <code>/plugin install nopus@nopus</code></td>
          <td>Bounded Stop hook requesting one clearer response</td>
      </tr>
      <tr>
          <td>Codex</td>
          <td><code>codex plugin marketplace add Vistyy/nopus</code> then <code>codex plugin add nopus@nopus</code></td>
          <td>Plugin; a new session is required so the plugin and skills load</td>
      </tr>
  </tbody>
</table>
<p>The user-facing surface is two bundled skills, <code>nopus-configure</code> and <code>nopus-simplify</code>, plus a command set: <code>/nopus status</code>, <code>/nopus check</code>, <code>/nopus on</code>, <code>/nopus off</code>, <code>/nopus extra-simple on|off</code>, and <code>/nopus hide-original on|off</code>. <code>nopus-simplify</code> exists because you do not have to wait for the hook — you can ask it to rewrite the immediately preceding response on demand.</p>
<p>Configuration lives in <code>$XDG_CONFIG_HOME/nopus/config.json</code> (resolved through <code>NOPUS_CONFIG</code>, then XDG, then platform defaults). Defaults are <code>complexitySensitivity: medium</code>, <code>includeEvidence: true</code>, <code>extraSimple: false</code>, <code>pi.hideOriginalResponse: true</code>. Every field has an environment override: <code>NOPUS_COMPLEXITY_SENSITIVITY</code>, <code>NOPUS_INCLUDE_EVIDENCE</code>, <code>NOPUS_EXTRA_SIMPLE</code>, <code>NOPUS_PI_HIDE_ORIGINAL_RESPONSE</code>.</p>
<p>One trust note specific to Pi: extensions run in-process with your full permissions, which makes any prose-checking extension an arbitrary-code dependency as well as a style tool. Pi&rsquo;s own documentation tells users to review package source before installing. The npm package has zero runtime dependencies and is small enough to read, which helps.</p>
<h2 id="the-evidence-read-honestly-5337-responses-and-50-labels">The Evidence, Read Honestly: 5,337 Responses and 50 Labels</h2>
<p>The most interesting thing about nopus&rsquo;s evaluation is that it publishes a result that does not flatter the tool.</p>
<table>
  <thead>
      <tr>
          <th>Evidence</th>
          <th>Result</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Corpus</td>
          <td>5,337 unique completed Pi responses, 515 session files, collected 2026-08-14</td>
      </tr>
      <tr>
          <td>Policy rewrite rates</td>
          <td>5.3% low / 9.9% medium / 18.6% high</td>
      </tr>
      <tr>
          <td>Human labels</td>
          <td>50 labels on medium-sensitivity batches</td>
      </tr>
      <tr>
          <td>Agreements</td>
          <td>41</td>
      </tr>
      <tr>
          <td>Wrong rewrites (accepted responses sent back)</td>
          <td>5</td>
      </tr>
      <tr>
          <td>Missed rewrites (responses marked for rewrite, not caught)</td>
          <td>4</td>
      </tr>
  </tbody>
</table>
<p>Fifty labels is a small sample, and the author states plainly that the batches were sampled <em>at</em> the medium boundary and are &ldquo;not a representative population sample&rdquo;. That means the directional read is roughly four in five correct on a deliberately hard batch — and, equally, that the numbers bound nothing about how the tool behaves on ordinary traffic in either direction.</p>
<p>Two structural limits follow. There is no published precision, recall or F1 for the policy anywhere in the repository. And the corpus being private means the headline rates cannot be audited independently. The project&rsquo;s honesty about this is unusual and should be rewarded, but the claim must still be sized to the evidence: one maintainer&rsquo;s private corpus, fifty human labels, and a rewrite model evaluated on four flagged originals.</p>
<h2 id="four-limits-you-should-know-before-installing">Four Limits You Should Know Before Installing</h2>
<p><strong>English only.</strong> SUBTLEX, Norvig&rsquo;s counts and Brysbaert&rsquo;s ratings are English datasets with no multilingual equivalents packaged. Non-English or code-switching responses produce meaningless metrics. For teams working in other languages this is a hard blocker, not a tuning problem — though a 2026 wave of small variants (a Korean <code>prose-lint</code>, a German <code>schreibwaechter</code>) shows the gap being filled from the edges.</p>
<p><strong>It cannot save you tokens.</strong> Nopus evaluates completed responses, so the flagged babble has already been generated and billed. Any cost saving is human reading time and transcript clutter. Anyone shopping for a cost-reduction tool is looking at the wrong product.</p>
<p><strong>One rewrite, maximum.</strong> That is the best engineering decision in the package: a bounded Stop hook that requests exactly one clearer response cannot create the infinite politeness loop a critic-model hook can. The price is that a response which is bad twice stays partly bad — nopus will never chase it further.</p>
<p><strong>Maintenance risk.</strong> Two npm versions (1.0.0 on 2026-08-15, 1.1.0 on 2026-08-16), 14 commits, and the last commit on 2026-08-19T16:25:13Z — roughly six weeks of silence as of this writing. The only open issue is #1 &ldquo;Add OpenCode plugin support&rdquo; (opened 2026-09-01, zero comments), and the most requested integration remains unbuilt. Star counts continue to drift up (repo at 296 stars on 2026-10-01; third-party snapshots in the same period read 264/276/287), but an unmaintained extension that runs in-process with full permissions is a dependency to pin and read, not infrastructure.</p>
<p>Distribution context is worth stating plainly, because it explains the adoption picture. Host tools are enormous — <code>@openai/codex</code> runs 25,653,416 weekly npm downloads, <code>@anthropic-ai/claude-code</code> 14,495,496, and <code>@earendil-works/pi-coding-agent</code> 4,316,299 in the same window — while <code>@syzom/nopus</code> itself had 34 downloads in its last week and 204 in the last month (window 2026-08-31 to 2026-09-29). Discoverability came from agent-package catalogs (agentmods, pi.dev, claudepluginhub) rather than developer-tool press: Hacker News has no nopus story at all, with an Algolia URL-restricted query returning zero hits.</p>
<h2 id="nopus-vs-write-good-alex-vale-and-avoid-ai-writing">Nopus vs write-good, alex, Vale, and avoid-ai-writing</h2>
<p>The prose-linting family is old and much larger, but none of it was built to interrupt an agent.</p>
<table>
  <thead>
      <tr>
          <th>Tool</th>
          <th>Stars</th>
          <th>npm downloads/month</th>
          <th>Output</th>
          <th>Where it runs</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>write-good</td>
          <td>5,092</td>
          <td>266,386</td>
          <td>Suggestions list</td>
          <td>Human documents, CI gate</td>
      </tr>
      <tr>
          <td>alex</td>
          <td>5,102</td>
          <td>184,122</td>
          <td>Warnings list</td>
          <td>Human documents, CI gate</td>
      </tr>
      <tr>
          <td>Vale</td>
          <td>6,175</td>
          <td>— (standalone)</td>
          <td>House-style violations</td>
          <td>Human documents, CI gate</td>
      </tr>
      <tr>
          <td>proselint</td>
          <td>4,579</td>
          <td>314</td>
          <td>Warnings list</td>
          <td>Effectively dormant</td>
      </tr>
      <tr>
          <td>avoid-ai-writing</td>
          <td>4,803</td>
          <td>2,092 (detector pkg)</td>
          <td>AI-tell detection</td>
          <td>Content, not agent responses</td>
      </tr>
      <tr>
          <td><strong>nopus</strong></td>
          <td>296</td>
          <td>204</td>
          <td><strong>One rewrite instruction</strong></td>
          <td><strong>Agent Stop hook / response</strong></td>
      </tr>
  </tbody>
</table>
<p>The distinction is the output type. Existing linters measure prose and return a report you read; nopus measures prose and returns an instruction the <em>agent</em> acts on. That is a new wiring of an old idea, and it is the only reason the project is interesting despite being three orders of magnitude smaller than write-good in usage.</p>
<p>The closer modern rival is <code>avoid-ai-writing</code> (4,803 stars, detector package published 2026-08-02, last modified 2026-09-23), which targets AI-writing tells in published content and explicitly defers house style to Vale. It answers &ldquo;does this read like AI wrote it?&rdquo; Nopus answers &ldquo;was this response too hard to read — rewrite it&rdquo;. Different job, different failure modes.</p>
<p>For the general linter lineage and its CI wiring, see the <a href="/posts/agent-skills-supply-chain-security-guide-2026/">agent skills supply chain security guide</a>, which covers how these extension packages get discovered and trusted. If you are building the hooks yourself rather than installing one, the <a href="/posts/claude-code-hooks-guide-2026/">Claude Code hooks guide</a> documents the lifecycle events nopus runs on, and the <a href="/posts/codex-plugins-integrations-guide-2026/">Codex plugins guide</a> covers that host&rsquo;s plugin loading.</p>
<h2 id="do-third-party-nopus-reviews-get-the-facts-right">Do Third-Party &ldquo;Nopus Reviews&rdquo; Get the Facts Right?</h2>
<p>Mostly no, and this is worth documenting because it is now typical of agent-tooling search results.</p>
<p>One aggregator page claims nopus implements a paper &ldquo;Reducing Hallucinations in Coding Agents via Prose Constraints&rdquo; (2026, arXiv:2607.89123). That arXiv ID returns &ldquo;Article not found&rdquo; (404), and no such paper is referenced anywhere in the repo. The same page invents a browser runtime, a CDN bundle, a CLI wrapper, an <code>@vistyy/nopus</code> package name, &ldquo;12+ prose/coding rules&rdquo; that block &ldquo;I think&rdquo;/&ldquo;maybe&rdquo;, code-block fencing validation, 98% coding-agent specificity, 67% hallucination reduction, and three community adapters including a Slack bot. None of it exists: there are 7 measurements and 6 signals, three host integrations, no browser target, and no hallucination-rate claim anywhere in the source.</p>
<p>Other write-ups are merely stale — one shows 154 stars against an actual 296 — and at least one openly labels its own metrics &ldquo;heuristic approximations &hellip; not authoritative measurements&rdquo;.</p>
<p>The lesson generalizes beyond nopus: for agent tooling, go to the repository and the registry. Both are cheap to read here, and the README is unusually specific about thresholds, defaults and provenance.</p>
<h2 id="does-the-concept-have-real-academic-support">Does the Concept Have Real Academic Support?</h2>
<p>The idea behind nopus — that verbosity is a measurable defect rather than a style preference — does, even though nopus itself has no published paper.</p>
<p>&ldquo;Verbosity Bias in Preference Labeling by Large Language Models&rdquo; (arXiv:2310.10076, 2023) shows LLM judges prefer more verbose answers of similar quality, which is the reason a <em>deterministic</em> check is a legitimate design choice rather than an odd one. &ldquo;Verbosity != Veracity&rdquo; (arXiv:2411.07858, 2024) documents verbosity compensation as a learned behaviour. Most directly, &ldquo;Verbosity-Aware Rationale Reduction&rdquo; (arXiv:2412.21006, ACL 2025 Findings) removed redundant reasoning sentences for an average +7.71% task performance while cutting token generation by 19.87% — trimming verbosity can improve accuracy, not just readability.</p>
<p>The developer-demand side is equally well measured. Stack Overflow&rsquo;s 2025 survey found 46% of developers distrust AI accuracy against 33% who trust it, with 66% naming &ldquo;almost right, but not quite&rdquo; as their top frustration. Sonar&rsquo;s 2026 State of Code survey (1,100+ respondents) found 96% do not fully trust that AI-generated code is functionally correct, and only 48% always check AI code before committing. Agent adoption roughly doubled to 59% while trust fell to 29%. Deterministic, auditable output guards are exactly the category that trust gap creates demand for.</p>
<h2 id="verdict-who-should-install-nopus-today">Verdict: Who Should Install Nopus Today?</h2>
<p>Install it if you use Pi, Claude Code or Codex, you are genuinely annoyed by lecture-hall framing and abstract &ldquo;capability / governance&rdquo; prose in agent answers, and you are comfortable reading ~2,000 lines of TypeScript before running it in-process. Keep the default medium sensitivity for a week, watch the rewrite rate, and pin the version — the project has been quiet since 2026-08-19 and you do not want a silent upgrade changing how often your agent interrupts itself.</p>
<p>Wait if any of these describe you: you work in a language other than English; you want token or cost savings; you need published precision and recall before adopting a policy that edits your agent&rsquo;s output; or you run transcript-based audit trails, in which case Pi&rsquo;s hidden-original default needs an explicit decision rather than a default.</p>
<p>Skip it entirely if what you actually want is house style enforcement — that is Vale, and Vale has been maintained since 2016 with a fraction of the uncertainty. Nopus is a well-scoped, unusually honest, single-maintainer experiment: the first prose linter wired to rewrite an agent&rsquo;s own answer, with a real corpus, a small human-label set, and a public threshold table you can audit. That is worth 34 weekly downloads and a careful look, not a site-wide rollout.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Is nopus an LLM or does it call one?</strong>
No. It is a deterministic checker over packaged lexical tables — no model is invoked to decide whether to rewrite, and no network call is needed for the decision. That is why the same response and sensitivity always produce the same verdict, and why the tool cannot drift the way a prompt-tuned critic can. The only model involved is your agent itself, which performs the rewrite.</p>
<p><strong>Does nopus work outside Pi, Claude Code and Codex?</strong>
Only those three are supported. Pi installs through <code>pi install npm:@syzom/nopus</code>, Claude Code through its plugin marketplace, and Codex through <code>codex plugin marketplace add</code>. OpenCode support is issue #1 and is still unbuilt. Porting it means re-implementing the Stop-hook or extension lifecycle against another harness, plus supplying your own host integration for the one rewrite request.</p>
<p><strong>Does it change the answer or only the wording?</strong>
Only the wording of the response, and only the visible response. The rewritten answer replaces the reply you read, but the rejected version stays in the session and model history — and in Pi it is merely hidden from the terminal transcript by default, not deleted. The agent&rsquo;s understanding of your task is unchanged, so nopus buys readability, not correctness.</p>
<p><strong>Does nopus save tokens or reduce cost?</strong>
No, and this is the most common misconception about it. Nopus evaluates responses that have already been generated and billed, so the flagged verbosity has already been paid for. The savings are human reading time and a cleaner transcript. There is a second-order effect — fewer tokens in future context if you keep shorter responses in history — but that is not what the tool measures or promises.</p>
<p><strong>Can nopus get stuck rewriting forever, or fire on technical terms?</strong>
It cannot loop: the design requests exactly one automatic rewrite, so a response that fails the policy twice is simply left as it is. On false positives, the defence is explicit — code blocks, inline code, URLs, file paths and table rows are stripped before measurement, established computing terms are down-weighted, and most decision paths require several measurements to cross together. The realistic failure mode is not identifiers but legitimately abstract discussion (authorization models, governance, capability boundaries) being asked to simplify.</p>
]]></content:encoded></item></channel></rss>