<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Bugbountyrules Skill on RockB</title><link>https://baeseokjae.github.io/tags/bugbountyrules-skill/</link><description>Recent content in Bugbountyrules Skill on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 06:12:53 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/bugbountyrules-skill/index.xml" rel="self" type="application/rss+xml"/><item><title>BugbountyRules Review: Agent Behavior Rules Security for Methodical AI Hunters</title><link>https://baeseokjae.github.io/posts/bugbountyrules-agent-discipline/</link><pubDate>Thu, 01 Oct 2026 06:12:53 +0000</pubDate><guid>https://baeseokjae.github.io/posts/bugbountyrules-agent-discipline/</guid><description>BugbountyRules keeps 42 always-on agent behavior rules in context so a hunting agent stops shipping false positives, inflating severity, and quitting early.</description><content:encoded><![CDATA[<p>BugbountyRules is a Claude Code skill that keeps 42 always-active behavioral rules in context — scope, coverage, evidence, persistence — instead of payloads. It exists because an agent already knows what SQL injection is; it fails by shipping false positives, inflating severity, and quitting early. The rules govern method, not knowledge.</p>
<p>That inversion is the whole point of the project. Nearly every &ldquo;AI security skill&rdquo; on the market is a knowledge pack: a longer list of things to look for. BugbountyRules (<a href="https://github.com/sanjarbiy/bugbountyrules">github.com/sanjarbiy/bugbountyrules</a>) takes the opposite position, and its README states the thesis bluntly — the agent&rsquo;s problem is not that it does not know what an IDOR is, it is that it will not reload prior state after a context compaction, will not distinguish a confirmed finding from a plausible one, and will declare a surface clean the moment it runs out of ideas. Those are behaviours. Behaviours are what the 42 rules target.</p>
<p>Below: what the skill actually contains, the three failures its own measurements are built around, how the verification gate and coverage ledger work, how it compares to the closest alternatives, and why installing any third-party behavioral skill is now a supply-chain decision.</p>
<h2 id="what-is-bugbountyrules-and-how-is-it-not-a-payload-pack">What Is BugbountyRules, and How Is It Not a Payload Pack?</h2>
<p>It is a single always-loaded behavioural core plus a deeply bundled, on-demand knowledge base.</p>
<p>The always-on layer is <code>SKILL.md</code> — 1,489 lines, 42 rules, verified by parsing the repository file on 2026-10-01. Those rules define how an agent hunts: what it must do before a request, how it proves a finding, when it is allowed to stop, and what language it may and may not use in a report. None of them ship an exploit.</p>
<p>The on-demand layer is the arsenal: a bundled <code>portswigger-kb</code> with 32 vulnerability-class folders and 93 sub-folders (roughly 271 files in the tree), a <code>reference/</code> directory of 8 files, plus <code>writeup-library.md</code>, <code>waf-bypass-arsenal.md</code>, <code>vuln-taxonomy.md</code>, <code>vulnerability-rating-taxonomy.json</code>, and a <code>scripts/</code> directory. Only the behavioural core sits in context by default; the technique library loads when a specific class becomes relevant.</p>
<p>The repository is small and young. Per the GitHub API on 2026-10-01: MIT license, 35 stars, 6 forks, 0 open issues, created 2026-08-17, last push 2026-08-18, roughly 855 KB. That is a deliberate design constraint, not a shortcoming to paper over — the architecture only works because the always-loaded part is behaviour and the heavy part is deferred.</p>
<h2 id="why-is-behaviour-not-knowledge-the-bottleneck">Why Is Behaviour, Not Knowledge, the Bottleneck?</h2>
<p>Because the knowledge problem is largely solved and the discipline problem visibly is not.</p>
<p>Autonomous agents now produce results that would have been implausible three years ago. XBOW submitted 1,060+ autonomous vulnerability reports on HackerOne and became the first AI system at #1 on the US leaderboard; in one published comparison it matched a 20-year veteran pentester&rsquo;s 85% solve rate on a 104-challenge suite in 28 minutes against the human&rsquo;s roughly 40 hours (<a href="https://xbow.com/blog/we-ran-1060-autonomous-attacks">xbow.com</a>). Wiz&rsquo;s Red Agent surfaced 17,000+ unique findings across about 1,000 customer environments in its first month, with access-control failures accounting for 54% of discoveries and exposed secrets for 61% of critical/high findings (<a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-autonomous-red-team-agent-findings-2026">CSA research note</a>). A single AI agent reached top-3 on multiple HackerOne business leaderboards on a ~$5,000/month budget, filing 150 reports in its primary window with 19 (12.7%) accepted or triaged, 64.4% of severity-rated findings critical or high, and 0% N/A on critical findings (<a href="https://firecompass.com/blog-ai-penetration-testing-hackerone-top-3-press-release">FireCompass, July 2026</a>).</p>
<p>Demand is there too. HackerOne&rsquo;s 9th Hacker-Powered Security Report records valid AI vulnerability reports up 210% (prompt injection up 540%) and programs with AI in scope up 270%, against 580,000+ validated vulnerabilities, $81M paid out and about $3B in breach losses avoided (<a href="https://www.hackerone.com/report/hacker-powered-security">hackerone.com/report</a>).</p>
<p>And yet, in the same report, 58% of surveyed security researchers say AI misses business logic or chained exploits, and only 12% believe it could replace them. That gap is not a knowledge gap — it is a judgment gap. An agent that cannot tell a real chained ATO from a suggestive response body is not missing a payload, it is missing a rule.</p>
<h2 id="what-are-the-three-failures-that-cost-money">What Are the Three Failures That Cost Money?</h2>
<p>The skill names three, and each one comes with a measurement rather than an assertion.</p>
<table>
  <thead>
      <tr>
          <th>Failure mode</th>
          <th>What it looks like in practice</th>
          <th>Mechanism the skill applies</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>False positives</td>
          <td>Plausible-but-unproven findings flood the report</td>
          <td>Hunter → Skeptic → Referee gate</td>
      </tr>
      <tr>
          <td>Severity inflation</td>
          <td>Every finding becomes critical</td>
          <td>Per-metric anchoring; hedging language banned</td>
      </tr>
      <tr>
          <td>Quitting early</td>
          <td>&ldquo;Surface appears secure&rdquo; after a shallow pass</td>
          <td>Coverage ledger, depth ladders, stuck loop</td>
      </tr>
  </tbody>
</table>
<p>The false-positive number is the one worth sitting with. The repository states that a small model handed eight findings waved through 3 of the 4 textbook false positives it was given and inflated severity on 6 of 8. That is a 75% false-positive pass rate on planted, unambiguous traps — not adversarial edge cases, textbook ones. For a triager on the receiving end, an agent with that profile is not a contributor, it is a queue.</p>
<p>Severity inflation is the second-order version of the same problem. A report is not a claim about what could happen; it is a claim about what did happen, under stated preconditions, with a stated impact. Per-metric anchoring forces the agent to justify each scored dimension separately instead of assigning a single flattering number, and banning hedging language removes the escape hatch of writing &ldquo;could potentially lead to&rdquo; where a demonstration was required.</p>
<p>Quitting early is the cheapest failure to miss, because a clean report and an abandoned hunt look identical from the outside.</p>
<h2 id="how-do-the-hunt-reflex-and-operational-flow-work">How Do the Hunt Reflex and Operational Flow Work?</h2>
<p>Two runtime mechanisms turn 42 static rules into a repeatable sequence.</p>
<p>The <strong>Hunt Reflex</strong> is a pre-move checklist: before the agent issues a request, it checks the rule set that applies to that move. The purpose is to make discipline reflexive rather than retrospective — a rule read after a failed attempt only produces a better excuse.</p>
<p>The <strong>Operational Flow</strong> is the longer loop: load prior state → probe tool pack → read scope → passive recon → map surface → hunt → verify → report. Two steps in that sequence are load-bearing. &ldquo;Read scope&rdquo; comes before any probe; &ldquo;load prior state&rdquo; comes first of all, because an agent recovering from a context compaction with no restored state is the duplicate-generating failure the skill measured directly.</p>
<p>The flow also encodes a ratio that experienced hunters already follow: roughly 60–70% of time in reconnaissance and mapping, not payload firing. An agent that rushes to exploitation is not aggressive, it is uninformed — and uninformed attackers produce the highest-severity noise.</p>
<h2 id="what-does-progressive-disclosure-actually-buy-you">What Does Progressive Disclosure Actually Buy You?</h2>
<p>It makes governance nearly free at the token level.</p>
<p>A 1,489-line behavioural core is always in context. A 32-class knowledge base with 93 sub-folders is not. That split means the rules that prevent errors are present on every single action, while the technique descriptions that only matter for one class of target are loaded on demand. The consequence is that an agent can be governed without carrying a library — the constraint that makes a rule-heavy skill practical at all.</p>
<p>The distinction matters when you compare approaches. A 93-skill library where every skill loads separately has breadth, but no single always-on artefact that guarantees the agent&rsquo;s <em>behaviour</em> is consistent across the whole engagement. Progressive disclosure is the design answer to &ldquo;how do I enforce 42 rules without paying 42 rules of context on every turn.&rdquo;</p>
<h2 id="which-rule-clusters-carry-the-weight">Which Rule Clusters Carry the Weight?</h2>
<p>Five clusters do most of the work.</p>
<p><strong>Rule 0, Adaptive Thinking.</strong> The skill is explicitly anti-robotic. Tools listed in it are examples, not mandates; every target is different; the rule ships an anti-pattern list of robotic behaviour to avoid. This is a deliberate contrast with static &ldquo;look for these patterns&rdquo; security checklists. A rigid checklist is the failure mode, not the fix.</p>
<p><strong>Scope discipline, Rules 1 and 9.</strong> Rule 1 exists because one out-of-scope request can get you banned from a programme — scope enforcement is legal safety, not housekeeping. Rule 9 covers the two-way error: an agent must not write off in-scope assets merely because of where they are hosted.</p>
<p><strong>State and duplication, Rules 21, 28 and 29.</strong> Roughly 1 in 5 actions on real engagements were exact duplicates, and around 20% of actions were wasted because prior state was not reloaded after context compaction. Those are the measurements the coverage ledger rules are built on.</p>
<p><strong>Honest reporting, Rules 24 and 25.</strong> Rule 25 refuses to let the agent conclude &ldquo;secure&rdquo; while surface remains unexamined; Rule 24 explicitly permits an honest zero over a fabricated finding. Together they invert the incentive that autonomous pentest agents are usually optimised for — benchmark score rather than report quality.</p>
<p><strong>Multi-engine routing, Rules 3.12 and 3.14.</strong> Rule 3.12 offloads bulk reading to a local LLM as a context firewall, keeping the hunting context clean. Rule 3.14 routes findings and &ldquo;secure&rdquo; conclusions to a frontier peer agent for a kill-gate and a resurrection-gate — an external check on both over-claiming and premature closure.</p>
<h2 id="how-does-the-hunter--skeptic--referee-gate-stop-false-positives">How Does the Hunter → Skeptic → Referee Gate Stop False Positives?</h2>
<p>By making the agent argue against itself before it is allowed to write a report.</p>
<p>The gate puts one persona in the position of finding, a second in the position of attacking the finding, and a third in the position of adjudicating between them. It is adversarial self-review, and it is aimed squarely at the 3-of-4 false-positive pass rate the project measured in an ungated model.</p>
<p>The gate is not a novelty. It maps closely onto the control categories OWASP formalised for agentic systems — reasoning-integrity and rogue-agent controls — where the risk is not a malicious tool but an agent confidently asserting an unfounded conclusion. A verification gate that requires a demonstrable reproduction, plus a severity justified per metric, plus language that cannot hedge its way around a missing proof, is the practical implementation of that control in a hunting context.</p>
<h2 id="what-is-the-coverage-ledger-and-why-does-it-kill-duplicates">What Is the Coverage Ledger, and Why Does It Kill Duplicates?</h2>
<p>The ledger is a running record of which surfaces have been examined, to what depth, and with what result — so an agent that reconsiders a target has to consult what it already knows rather than re-probe from memory.</p>
<p>It attacks waste from two directions. The first is exact duplication: ~1 in 5 actions on real engagements were repeats. The second is context-loss duplication: about 20% of actions were wasted because prior state was not reloaded after compaction. Both are the same underlying defect — an agent that treats every moment as a fresh start will pay for the same knowledge twice.</p>
<p>Depth ladders and a stuck loop complete the mechanism. Instead of a binary &ldquo;tested / not tested&rdquo;, each surface carries a depth level, so &ldquo;I tried one payload&rdquo; and &ldquo;I worked the class properly&rdquo; are no longer the same entry. The stuck loop gives the agent a defined response to being out of ideas that is not silence and not a fabricated finding.</p>
<h2 id="why-is-scope-discipline-a-legal-question-not-a-tidiness-one">Why Is Scope Discipline a Legal Question, Not a Tidiness One?</h2>
<p>Because the downside is not a missed finding, it is exclusion from the programme.</p>
<p>An out-of-scope request is a real-world action against a system the operator was not authorised to touch, and bug bounty programmes treat it accordingly — the README&rsquo;s framing is that a single such request can get you banned. A rule that lives in the always-loaded core is the only kind that can fire <em>before</em> the request goes out, which is why scope enforcement belongs with behaviour rather than with technique.</p>
<p>Rule 9 closes the other half of the loop: an agent must not dismiss an asset that is genuinely in scope simply because of where it is hosted. Both directions of scope error cost the operator money — one in credibility, the other in coverage.</p>
<h2 id="what-do-the-local-llm-offload-and-the-frontier-peer-agent-add">What Do the Local-LLM Offload and the Frontier Peer Agent Add?</h2>
<p>They turn a single agent into a small pack with separate failure domains.</p>
<p>Rule 3.12&rsquo;s local-LLM offload is a context firewall: bulk reading — long response bodies, large source files, documentation dumps — is processed outside the primary hunting context, so the agent&rsquo;s working memory is spent on reasoning instead of transcription. Rule 3.14&rsquo;s frontier peer agent is the opposite trade: expensive, but it provides an independent reviewer for both findings and &ldquo;secure&rdquo; conclusions.</p>
<p>The pairing is what makes the kill-gate meaningful. A gate applied by the same context that produced the finding is a gate applied by a party with an interest in the outcome. An external reviewer with no stake in the claim and no shared context is a materially different check — and the resurrection-gate direction (challenging a <em>negative</em> conclusion) is the half that most agent pipelines omit entirely.</p>
<h2 id="how-does-bugbountyrules-compare-to-other-agent-security-skills">How Does BugbountyRules Compare to Other Agent Security Skills?</h2>
<p>Three architectures are competing for the same job.</p>
<table>
  <thead>
      <tr>
          <th></th>
          <th>BugbountyRules</th>
          <th>Murrtada&rsquo;s bb skills</th>
          <th>Sentry&rsquo;s security-review</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Structure</td>
          <td>One always-on behavioural core</td>
          <td>93 single-SKILL.md library</td>
          <td>One review skill + many references</td>
      </tr>
      <tr>
          <td>Emphasis</td>
          <td>Depth of behaviour</td>
          <td>Breadth of vulnerability classes</td>
          <td>Confidence-tiered review</td>
      </tr>
      <tr>
          <td>Verification</td>
          <td>Hunter → Skeptic → Referee</td>
          <td>7-question verify-or-kill gate</td>
          <td>HIGH/MEDIUM/LOW confidence</td>
      </tr>
      <tr>
          <td>Strongest claim</td>
          <td>Rules govern the agent globally</td>
          <td>Primitive chaining (A→B→C)</td>
          <td>Data-flow tracing discipline</td>
      </tr>
      <tr>
          <td>Weakness</td>
          <td>Single author, unproven at scale</td>
          <td>No single global behavioural guarantee</td>
          <td>Review-focused, not hunt-focused</td>
      </tr>
  </tbody>
</table>
<p><a href="https://github.com/murrtada/bug-bounty-agent-skills">github.com/murrtada/bug-bounty-agent-skills</a> is the closest ecosystem competitor: 93 offensive-security skills (<code>bb-methodology</code>, <code>bb-recon</code>, <code>bb-verify</code>, <code>bb-report</code> and a <code>hunt-*</code> family) rather than one core. Its verify-or-kill triage gate runs 7 questions on every finding before report-writing, and its <code>hunt-*</code> skills ship as primitives with explicit &ldquo;chain to X&rdquo; sections — so CSRF becomes account takeover and SSRF becomes cloud-metadata compromise. Per-class gate/verify/kill tables define what counts as proof, with rules like &ldquo;OOB callback or it didn&rsquo;t happen&rdquo; for blind SSRF. It covers modern classes including OWASP&rsquo;s ASI01–ASI10 agentic categories, CI/CD and GitHub Actions abuse, HTTP/2 single-packet races, and protocol-version shadow APIs. Verified 2026-10-01: NOASSERTION license, 1 star, created 2026-08-28, last push 2026-09-26, ~831 KB.</p>
<p>It is methodology-as-library. BugbountyRules is methodology-as-discipline. A hands-on comparison of five Claude Code security skills reached the same conclusion about what separates good from mediocre: the winning entry (Sentry&rsquo;s) defined a confidence system, carried false-positive awareness, and traced data flows rather than reciting patterns — &ldquo;the difference between a thin checklist and a methodology is the difference between noise and signal&rdquo; (<a href="https://timonweb.com/ai/i-checked-5-security-skills-for-claude-code-only-one-is-worth-installing/">timonweb.com</a>). That review also documents the install-count trap: an aggregator repo carrying 900+ skills inflated the top-ranked skill&rsquo;s installs, which is why distribution metrics are not quality signals here.</p>
<h2 id="what-is-the-ecosystem-risk-of-installing-a-behavioural-skill">What Is the Ecosystem Risk of Installing a Behavioural Skill?</h2>
<p>A behavioural skill is an executable instruction bundle, and it inherits the entire agent-skill supply chain problem.</p>
<p>The category is now formally named. OWASP&rsquo;s Top 10 for Agentic Applications (ASI01–ASI10) was published 2025-12-09 with 100+ contributors, and <strong>ASI04, Agentic Supply Chain Vulnerabilities</strong>, covers exactly what agents load at runtime: MCP servers, plugins, prompt templates, tool descriptors, and skills (<a href="https://blckalpaca.at/en/knowledge-base/ai-agents/ai-agent-security-owasp/owasp-agentic-asi-top-10-2026">OWASP ASI overview</a>). A companion Agentic Skills Top 10 (AST10) project promotes the skills layer to a first-class vulnerable component. The same source documents grounded incidents: postmark-mcp, the first malicious MCP server found in the wild (Koi Security, September 2025); EchoLeak, CVE-2025-32711, CVSS 9.3; and a Gemini memory-poisoning attack (Rehberger, February 2025). In Galileo AI&rsquo;s December 2025 simulation, one compromised agent poisoned 87% of downstream decisions within four hours.</p>
<p>The scale numbers are worse than most teams assume. A large-scale study scanning 42,447 agent skills found 26.1% exhibit at least one security vulnerability, and about 5.2% show signs of likely malicious intent. ClawHub hosted 49,592 community skills as of April 2026, implying 2,500+ potentially malicious packages. The ClawHavoc campaign introduced 1,184 malicious skills to ClawHub, and five of them still evaded the updated ClawScan and VirusTotal between February and May 2026 (<a href="https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-agent-skill-scanner-bypass-20260629-csa/">CSA research note</a>).</p>
<p>Scanner coverage is not the answer, and there is a structural reason. Trail of Bits bypassed ClawHub&rsquo;s malicious skill detector, Cisco&rsquo;s agent skill scanner, and all three scanners integrated into skills.sh; three of the four malicious skills were conceived and implemented in under an hour using standard tricks plus reading the scanner source. The bypasses included roughly 100,000 newlines to push a payload past ClawHub&rsquo;s truncation window, malicious <code>.pyc</code> bytecode sitting next to benign source, DOCX/ZIP archive indirection, and social-engineering the LLM-as-judge layer with compliance-policy-sounding prose (<a href="https://blog.trailofbits.com/2026/06/03/the-sorry-state-of-skill-distribution/">Trail of Bits</a>). Their conclusion is the ecosystem&rsquo;s framing: no amount of scanning or LLM analysis can reliably detect malicious content in agent skills — marketplaces are one layer, not a gate. One scanner-bypassing skill reportedly reached ~26,000 agents including corporate accounts, a vendor-reported figure that has not been independently verified (<a href="https://csoonline.com/article/4188840/how-a-malicious-ai-agent-skill-passed-security-checks-and-reached-26000-users.html">CSO Online</a>).</p>
<p>The structural limit is worth stating precisely: formal static analysis cannot reason about natural-language instructions in a <code>SKILL.md</code>, because the consuming LLM interprets them at runtime. Even SkillFortify, reporting 96.95% F1 on a 540-skill benchmark, does not close that gap. An always-loaded behavioural skill is precisely the artefact class this research says scanners cannot validate.</p>
<p>The practical mitigations follow from the OWASP design principles of least agency and strong observability: treat scanner passage as one signal among many, verify publisher identity out-of-band, and manually read the behavioural directives in <code>SKILL.md</code> before you let them into an agent&rsquo;s context on every turn.</p>
<h2 id="what-are-the-honest-limits">What Are the Honest Limits?</h2>
<p>BugbountyRules is a single-author project with 35 stars, 6 forks, zero open issues, and a repository history spanning two days (created 2026-08-17, last push 2026-08-18). It has not been independently benchmarked, and its numbers are its own.</p>
<p>Its most interesting verifiability claim is also its least tested. The repo ships <code>run-eval.sh</code>, which puts a rule&rsquo;s own text in front of an independent agent to check whether the rule actually instructs behaviour rather than merely reading well. That is a genuinely good idea — a rule set that can be evaluated for instructiveness is more than a prose artefact — but the results are not published against a third-party baseline, and no external party has reproduced them.</p>
<p>Two things the skill does not do: it does not replace the exploit knowledge in the bundled KB, and it does not make an agent correct. It makes an agent honest about what it knows. The economics explain why that is worth anything: one manual penetration test of a single application costs $2,400–$10,000, and complex engagements run to $40,000 or more (<a href="https://nhimg.org/articles/ai-agents-are-reshaping-penetration-testing-economics">nhimg.org</a>). An agent that converts that spend into triager noise produces negative value; an agent that produces twelve validated findings beats one that produces eighty plausible ones.</p>
<h2 id="who-should-install-it-and-how-do-you-start">Who Should Install It, and How Do You Start?</h2>
<p>It fits teams already running an agent against authorised targets — bug bounty programmes, internal red team scopes, CTF-style engagements — who have hit the false-positive wall rather than the knowledge wall.</p>
<p>A sensible adoption path:</p>
<ol>
<li>Read <code>SKILL.md</code> yourself, line by line. It is 1,489 lines of behavioural directives and it will be in your agent&rsquo;s context on every turn; treat it like code you are about to run.</li>
<li>Confirm the repository&rsquo;s provenance out-of-band, and pin the commit you reviewed rather than tracking a branch.</li>
<li>Verify that Rule 1&rsquo;s scope handling matches your actual programme rules — scope enforcement is the one failure with legal consequences.</li>
<li>Start with the behaviour rules and the bundled KB off. Run a scoped engagement and check the false-positive rate before adding the technique layer.</li>
<li>If you need class breadth more than behavioural depth, pair it with a library-style skill set such as Murrtada&rsquo;s rather than expecting one core to cover both.</li>
</ol>
<h2 id="faq">FAQ</h2>
<p><strong>What is BugbountyRules in one sentence?</strong>
A Claude Code skill holding 42 always-active behavioural rules — scope, coverage, evidence, persistence — that govern how an agent hunts rather than which payloads it fires.</p>
<p><strong>How many rules does it actually define?</strong>
Exactly 42, inside a 1,489-line <code>SKILL.md</code>, verified by parsing the repository file on 2026-10-01.</p>
<p><strong>What is the Hunter → Skeptic → Referee gate?</strong>
An adversarial self-review that runs before report-writing: one pass finds, one pass attacks the finding, one adjudicates. It exists because a small model given eight findings waved through 3 of the 4 textbook false positives and inflated severity on 6 of 8.</p>
<p><strong>What is the coverage ledger for?</strong>
Tracking which surfaces were examined and to what depth, so the agent does not repeat work. Roughly 1 in 5 actions on real engagements were exact duplicates, and about 20% were wasted because prior state was not reloaded after context compaction.</p>
<p><strong>Is installing a third-party agent skill safe if a scanner passed it?</strong>
No. Trail of Bits bypassed ClawHub, Cisco and all three skills.sh scanners, and a scan of 42,447 agent skills found 26.1% with at least one vulnerability. Read the behavioural directives yourself, verify the publisher out-of-band, and treat scanner passage as one signal, not a gate.</p>
]]></content:encoded></item></channel></rss>