<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AI Coding Agent Leadership on RockB</title><link>https://baeseokjae.github.io/tags/ai-coding-agent-leadership/</link><description>Recent content in AI Coding Agent Leadership on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 07:42:21 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/ai-coding-agent-leadership/index.xml" rel="self" type="application/rss+xml"/><item><title>Working With AI Feels More Like Leadership Than Coding: A Guide to Managing AI Coding Agents</title><link>https://baeseokjae.github.io/posts/working-with-ai-feels-more-like-leadership-than-coding/</link><pubDate>Thu, 01 Oct 2026 07:42:21 +0000</pubDate><guid>https://baeseokjae.github.io/posts/working-with-ai-feels-more-like-leadership-than-coding/</guid><description>Managing AI coding agents is real leadership: agent-wrangling skill correlates 0.81 with managing humans. Here is the delegation and review model.</description><content:encoded><![CDATA[<p>Working with AI coding agents feels like leadership because it is. The unit of work changed from writing code to decomposing it, delegating it, and verifying the result. That shift is measurable: a 2026 NBER working paper found that a leader&rsquo;s performance managing AI agents correlates at rho = 0.81 with their performance managing human teams — and still 0.69 after controlling for hard technical skills.</p>
<p>That single finding reframes everything else in this guide. If managing AI coding agents were a disguised coding skill, the correlation with human-team leadership would collapse once you controlled for task-specific ability. It does not. What the AI task is actually measuring is the soft, transferable part of leadership: how you set expectations, calibrate trust, and run a review loop.</p>
<p>So this is not a metaphor you reach for when the tooling gets confusing. It is an empirical description of the job — and, more usefully, a diagnosis with a treatment. The management canon you may have never read maps onto agents with almost no translation, and the failure modes you are about to hit have already been documented.</p>
<h2 id="you-got-promoted-and-nobody-told-you">You got promoted and nobody told you</h2>
<p>Most developers never applied for a management job. They applied to write software. Then between 2025 and 2026, the shape of the daily routine changed underneath them.</p>
<p>Anthropic&rsquo;s 2026 Agentic Coding Trends Report — which has been widely reported but whose primary PDF is not publicly hosted, so treat these figures as vendor-reported — describes the transition with session telemetry. Average Claude Code session length went from about 4 minutes in Q1 2025 to roughly 23 minutes in Q1 2026. The share of sessions involving multi-file edits rose from 34% to 78%. The average session now makes around 47 tool calls.</p>
<p>Read that as a job description rather than a product metric. You are no longer typing the lines. You are scoping a change, handing it to something that will touch a dozen files, and then deciding whether to accept what came back. That is a stand-up, a ticket hand-off, and a code review, compressed into one loop. It is management work with the meeting overhead stripped out.</p>
<p>The inversion shows up in how practitioners describe their own fleets. Boris Cherny runs five local and five to ten browser-based Claude Code sessions in parallel, a workflow Addy Osmani uses to make the point that AI coding at scale stops being a prompting problem and becomes a management problem. Among intensive OpenAI Codex users, 28.6% peaked at five or more concurrent agents, with p99 usage around 71 agent-hours per day — a number that only makes sense if a human is coordinating rather than executing.</p>
<p>If your day now consists of writing specifications, dispatching work, reading diffs, and deciding what to escalate, you are not &ldquo;using AI.&rdquo; You are running a team.</p>
<h2 id="the-evidence-that-this-really-is-leadership-not-a-metaphor">The evidence that this really is leadership, not a metaphor</h2>
<p>The strongest study here is NBER Working Paper 33662 by Ben Weidmann, Yixian Xu and David Deming, summarized in the NBER Research Digest. The design is what makes it convincing. Participants led both a team of human collaborators and a set of AI agents through structured work, and the researchers compared the two performance measures.</p>
<table>
  <thead>
      <tr>
          <th>Finding</th>
          <th>Result</th>
          <th>What it means for you</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Correlation, agent leadership vs human-team leadership</td>
          <td>rho = 0.81</td>
          <td>The two tasks share most of their skill content</td>
      </tr>
      <tr>
          <td>Correlation after controlling for hard skills</td>
          <td>rho = 0.69</td>
          <td>The overlap is leadership-specific, not technical ability</td>
      </tr>
      <tr>
          <td>Effect of leader quality on team outcome</td>
          <td>+1 SD leader ≈ +0.65 SD team performance</td>
          <td>Who leads matters more than which model you run</td>
      </tr>
      <tr>
          <td>Behavioural markers of strong leaders</td>
          <td>More questions, more turn-taking, more &ldquo;we&rdquo;/&ldquo;us&rdquo; language</td>
          <td>The winning behaviours are conversational, not technical</td>
      </tr>
      <tr>
          <td>Demographic predictors</td>
          <td>None significant</td>
          <td>Seniority and titles did not predict who led well</td>
      </tr>
      <tr>
          <td>Cost per assessment</td>
          <td>$23 (AI) vs $114 (human)</td>
          <td>Agent work makes leadership measurable at scale</td>
      </tr>
  </tbody>
</table>
<p>Two details deserve emphasis. First, the correlation survived controls for task-specific ability and fluid intelligence, which means the AI setting is capturing something real about how a person directs other workers. Second, in both settings the leader explained more than half the variation in team performance. The model was not the variable. The person directing the model was.</p>
<p>What does <em>not</em> transfer matters too. Agents do not negotiate roles, self-organize, or give each other psychological safety. All coordination remains centralized on you. Fortune&rsquo;s coverage of the same research quotes the analogy directly: with AI teams the human is closer to an orchestral conductor than a jazz-ensemble leader — every cue originates from one baton. You get the authority of a manager without the delegation of authority that a real team provides.</p>
<h2 id="the-trap-you-are-a-worse-judge-of-your-own-ai-workflow-than-you-think">The trap: you are a worse judge of your own AI workflow than you think</h2>
<p>Here is the part of the leadership literature that stings. Management instinct is miscalibrated by default, and the calibration error is largest precisely for experts.</p>
<p>METR&rsquo;s randomized controlled trial is the reference point. Sixteen experienced open-source developers worked 246 real issues on repositories averaging 22,000+ stars and over a million lines of code. When allowed to use AI, they were <strong>19% slower</strong>. Their forecast beforehand was a 24% speedup. Their post-hoc belief was that AI had sped them up by 20%. Reality and perception pointed in opposite directions, roughly 39 points apart.</p>
<p>METR&rsquo;s authors were careful about scope, and so should you be: the study does not show that AI fails to speed up most developers, and it explicitly notes that unfamiliar codebases, less experienced developers, and greenfield projects plausibly behave differently. Five contributing factors were identified — over-optimism, deep familiarity with the repository, large and complex codebases, low suggestion acceptance (under 44%), and implicit repository context the model could not see.</p>
<p>The leadership lesson is not &ldquo;AI is slow.&rdquo; It is that <strong>a confident manager with no instrumentation is how you get a fleet that looks productive and is not</strong>. METR&rsquo;s developers were not lazy or naive; they were experienced engineers who trusted their own feel for the work. Feel is exactly what stops being reliable when the work is distributed across agents.</p>
<p>The industry-level evidence rhymes. Google Cloud&rsquo;s DORA 2025 report (roughly 5,000 respondents) found 90% of technology professionals use AI at work and more than 80% report productivity gains — while 30% report little or no trust in AI-generated code. DORA&rsquo;s central framing is that AI is an amplifier: it shows a positive relationship with throughput and product performance alongside a <em>continuing negative relationship with delivery stability</em>. DORA 2024 had already estimated a 1.5% throughput reduction and a 7.2% instability increase for every 25% increase in AI adoption.</p>
<p>Faros AI&rsquo;s July 2025 telemetry (1,255 teams, 10,000+ developers) shows the same pattern in delivery terms: teams with high AI adoption completed 21% more tasks and merged 98% more pull requests, while PR review time rose 91%, average PR size rose 154%, and bugs per developer rose 9% — with no measurable organization-level DORA improvement.</p>
<p>Output went up. So did the cost of checking it. Both are true, and only one of them shows up in a demo.</p>
<h2 id="the-delegation-gap-60-usage-0-20-full-delegation">The delegation gap: 60% usage, 0-20% full delegation</h2>
<p>If you want a single number that defines the job, take this one from Anthropic&rsquo;s 2026 report: developers use AI in roughly 60% of their work, but can &ldquo;fully delegate&rdquo; only 0-20% of tasks without line-by-line review.</p>
<p>That 40-to-60-point gap is not a model-capability problem waiting for the next release. It is a trust-calibration problem, and trust calibration is a management discipline with forty years of prior art. Andy Grove&rsquo;s <em>High Output Management</em> (1983) introduced task-relevant maturity: the appropriate level of supervision depends on the subordinate&rsquo;s demonstrated competence with <em>this specific task class</em>, not on a global judgment about their talent.</p>
<p>Applied to agents, that produces a rule that is both more permissive and more useful than either extreme:</p>
<ul>
<li>Autonomy is granted per task class, not globally. &ldquo;This agent handles test scaffolding unsupervised&rdquo; is a valid, earned statement. &ldquo;I trust the agent&rdquo; is not.</li>
<li>Maturity is demonstrated by evidence, not by vibes. Track the failure rate per class. Raise autonomy only when the class has a track record.</li>
<li>New classes start at high supervision regardless of how good the model is. A better model does not shorten the trust ramp for work you have never delegated.</li>
</ul>
<p>Management 3.0&rsquo;s seven-level delegation spectrum — Tell, Sell, Consult, Agree, Advise, Inquire, Delegate — turns out to be literally implemented in your harness. Plan mode is Consult. Auto-accept is Advise. Full bypass is Delegate. You are not designing a new permission model; you are re-deriving a standard one, badly, from scratch, if you never read it.</p>
<h2 id="verification-is-the-binding-constraint-so-fleet-size-is-a-review-decision">Verification is the binding constraint, so fleet size is a review decision</h2>
<p>The instinct when you get faster at dispatching work is to dispatch more of it. That instinct runs straight into the actual bottleneck, which is not generation.</p>
<p>A February 2026 MIT working paper (Catalini et al., cited in analysis of management practice for agents) frames it plainly: the binding constraint on scaling agentic work is human verification bandwidth, not model intelligence. &ldquo;The code got cheap this year. The supervision didn&rsquo;t,&rdquo; as one practitioner write-up puts it — the unit of work changed from accepting a completion to handing over a ticket, and the ticket comes back as an obligation to review.</p>
<p>Look at what the review side is doing while output climbs.</p>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>Faros AI 2025 (1,255 teams)</th>
          <th>Faros AI 2026 (22,000 developers)</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Tasks completed</td>
          <td>+21%</td>
          <td>—</td>
      </tr>
      <tr>
          <td>PRs merged</td>
          <td>+98%</td>
          <td>—</td>
      </tr>
      <tr>
          <td>Time in PR review</td>
          <td>+91%</td>
          <td>+441% (median)</td>
      </tr>
      <tr>
          <td>Average PR size</td>
          <td>+154%</td>
          <td>—</td>
      </tr>
      <tr>
          <td>Bugs / incidents per PR</td>
          <td>+9%</td>
          <td>+242.7% per PR</td>
      </tr>
      <tr>
          <td>PRs merged with no review</td>
          <td>—</td>
          <td>+31% more</td>
      </tr>
      <tr>
          <td>PRs reviewed by an AI agent</td>
          <td>0%</td>
          <td>25%</td>
      </tr>
  </tbody>
</table>
<p>These are two different cohorts from the same vendor, not one time series, so do not read the numbers as a single trend line. Read them as two snapshots of the same structural problem: generation scaled, review did not, and the gap is being closed partly by <em>not reviewing</em>. Nearly a third more PRs now merge with no human review at all. A quarter are reviewed by another agent.</p>
<p>The seniority tax is real and usually unbudgeted. JetBrains&rsquo; ICSE 2026 telemetry of 800 developers (reported secondhand) found AI users performing roughly 100 delete/undo actions per month versus 7 for non-users — a 14x rework gap. That rework lands on whoever has to reason about the change afterwards, which is rarely the person who prompted it.</p>
<p>So the practical formula is not &ldquo;how many agents can I run.&rdquo; It is:</p>
<p><strong>Sustainable fleet size = your review capacity ÷ the review cost of one agent&rsquo;s output</strong></p>
<p>If the answer comes out below your ambition, the fleet you actually have is the smaller number. Everything above it is generating obligations, not value. Practitioner ceilings cluster at three to five concurrent agents for meaningful work, and five to seven CLI agents on a laptop before rate limits and merge conflicts eat the gains. Anthropic&rsquo;s own harness ships a 20-concurrent default with a &ldquo;fewer than 15 agents&rdquo; workflow guideline — a vendor capping its own feature because customers were drowning.</p>
<h2 id="you-are-not-inventing-management-the-canon-already-covers-this">You are not inventing management: the canon already covers this</h2>
<p>The most useful discovery available to a developer who has accidentally become a manager is that the field is not new. Mapping the canon onto agents is nearly mechanical.</p>
<table>
  <thead>
      <tr>
          <th>Management concept</th>
          <th>Source</th>
          <th>Agent equivalent</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Task-relevant maturity</td>
          <td>Grove, <em>High Output Management</em> (1983)</td>
          <td>Autonomy earned per task class, tracked by failure rate</td>
      </tr>
      <tr>
          <td>Delegation levels (Tell → Delegate)</td>
          <td>Management 3.0 delegation poker</td>
          <td>Plan mode → auto-accept → full bypass</td>
      </tr>
      <tr>
          <td>Decision rights</td>
          <td>Standard org design</td>
          <td><code>AGENTS.md</code> / <code>CLAUDE.md</code>: defaults plus escalation paths</td>
      </tr>
      <tr>
          <td>Director model</td>
          <td>Camille Fournier, <em>The Manager&rsquo;s Path</em></td>
          <td>Spend a third of your time on guardrails and tooling, not per-artifact oversight</td>
      </tr>
      <tr>
          <td>Standing meetings and onboarding docs</td>
          <td>Any management handbook</td>
          <td>Reusable context assets, evaluation harnesses, review checklists</td>
      </tr>
  </tbody>
</table>
<p>The decision-rights mapping is the one people underuse. A <code>CLAUDE.md</code> or <code>AGENTS.md</code> is not a prompt; it is a decision-rights document for a worker with no judgment of its own. That means it should contain the defaults and, crucially, the escalation triggers: build and test commands, hard constraints, do-not-touch paths, and explicit stop-and-ask conditions — &ldquo;stop before touching auth, migrations, payments, or CI.&rdquo; Keep it short. Bloated machine-generated context files reduce task success and raise inference cost, which is the documentation equivalent of writing a policy manual nobody reads.</p>
<p>Fournier&rsquo;s director model tells you where the leverage is. Reviewing each artifact one at a time does not scale, and never did for human teams either. Authoring the guardrail, the tooling, and the check that runs <em>before</em> you look is what compounds.</p>
<h2 id="the-supervision-loop-plan-monitor-wait-review-teach-manual-fix-update-assets">The supervision loop: Plan, Monitor, Wait, Review, Teach, Manual Fix, Update Assets</h2>
<p>If you want a defensible process rather than a pile of tips, there is now a framework paper. arXiv 2609.24234, &ldquo;The Work Behind Delegation: A Framework for Supervising AI Coding Agents&rdquo; (September 2026), derives its model from 19 experienced developers and reconfigures Sheridan&rsquo;s classic supervisory-control model into seven stages:</p>
<ol>
<li><strong>Plan</strong> — decide what the agent will do and what it must not do. This is the highest-leverage stage and the one people skip.</li>
<li><strong>Monitor</strong> — watch progress without hovering. Use checkpoints, not continuous attention.</li>
<li><strong>Wait</strong> — genuinely disengage while the agent works. Unproductive waiting is a cost, not diligence.</li>
<li><strong>Review</strong> — inspect the output against the plan, in a context that did not write the code.</li>
<li><strong>Teach</strong> — correct the agent when it goes wrong in a way that will recur.</li>
<li><strong>Manual Fix</strong> — take over directly when the task is faster to finish than to explain.</li>
<li><strong>Update Assets</strong> — turn recurring guidance into a reusable artifact: a rule in <code>AGENTS.md</code>, a lint check, an evaluation case, a test.</li>
</ol>
<p>Stage seven is the compounding lever, and it is why the loop is a loop. Every correction you make should either fix the output or fix the system. Corrections that only fix the output get paid for repeatedly.</p>
<p>The paper also documents how experienced developers manage supervisory load: they concentrate effort in planning, delegate supervisory work to <em>other agents</em> (a verifier, a reviewer), and convert recurring guidance into reusable assets. That is a manager delegating oversight — and it is the same move a good engineering manager makes when they turn a recurring review comment into a lint rule.</p>
<h2 id="when-a-fleet-is-the-wrong-answer">When a fleet is the wrong answer</h2>
<p>Parallelism has a real cost curve, and the evidence against naive fleets is stronger than the marketing.</p>
<p>Agents collaborating achieve roughly 50% lower success than solo agents, according to CooperBench (January 2026, 652 tasks, reported secondhand). Two-agent cooperation on frontier models succeeded only about 25% of the time. The failure split was informative: roughly 26% communication, 32% commitment, and 42% expectation failures — which is to say, this is a management problem, not an intelligence problem. It is the same taxonomy as a project with no clear owner.</p>
<p>Google Research evaluated 180 agent configurations (also reported secondhand) and found multi-agent coordination yielding up to +81% on parallelizable tasks while <em>degrading performance by 70%</em> on sequential tasks. The pattern is consistent and worth internalizing:</p>
<p><strong>Do not run a fleet on the critical path.</strong> If B depends on A, concurrency adds coordination cost and buys nothing. Fleet work is for independent, crisply specified, independently testable slices.</p>
<p>Beyond sequential work, avoid fleets for:</p>
<ul>
<li><strong>Shared files.</strong> Parallel agents touching adjacent code produce boundary failures, not tooling failures. Worktree-per-agent is the standard fix; Bun&rsquo;s 64-agent Zig-to-Rust port saw shared-tree agents destroy each other&rsquo;s work within minutes before the team moved to sharded worktrees and process isolation. Cursor&rsquo;s from-scratch agent version-control system recorded over 70,000 merge conflicts, with one file touched by 1,173 distinct agents at 7,771 conflicts — falling below 1,000 after the harness was rebuilt.</li>
<li><strong>Small tasks.</strong> Below roughly 30 minutes of human work, dispatch and review overhead usually exceeds the gain.</li>
<li><strong>Judgment work.</strong> Product intent, API design, architecture, and &ldquo;should we build this?&rdquo; are not delegation targets. Osmani calls this the over-delegation trap: the failure is not that the agent does it badly, it is that you stopped thinking about it.</li>
</ul>
<p>Also watch for agentic drift: parallel agents converging on conflicting implementations of the same concept while all tests still pass. Green tests are not coordination.</p>
<h2 id="the-first-30-days-as-an-agent-lead">The first 30 days as an agent lead</h2>
<p>If you want a concrete starting sequence rather than a philosophy, this is the one the evidence supports.</p>
<p><strong>Days 1-7: instrument before you optimize.</strong> Pick five metrics and collect a baseline. Time in review, PR size, share of PRs merged with no review, rework (undo/delete churn), and escaped bugs or incidents per PR. Faros&rsquo; 2026 numbers show exactly why review-side metrics matter: outputs rose, median review time rose 441%, and 31% more PRs merged with no review. Note that &ldquo;PRs merged&rdquo; is a vanity metric in this regime.</p>
<p><strong>Days 8-14: write your decision-rights document.</strong> Build and test commands, hard constraints, do-not-touch paths, escalation triggers. Keep it under a page. This is the artifact you will update most often, and updating it <em>is</em> the management work.</p>
<p><strong>Days 15-21: define your task classes.</strong> List the five to ten things you actually delegate. For each, write down the required evidence for acceptance. Start every class at high supervision, and promote a class only after it has a clean track record.</p>
<p><strong>Days 22-30: cap the fleet by review capacity and set kill criteria.</strong> Compute your honest review throughput, then set your concurrency to match — commonly three to five agents for meaningful work. Write down in advance the conditions under which you kill a running agent rather than letting it finish. Deciding that while the agent is mid-run is how review debt accumulates.</p>
<p>Then close the loop: every recurring correction goes into an asset — a rule, a check, a test — so that the supervision cost per task declines over time. That is the compounding return, and it is the same thing good managers have always done.</p>
<h2 id="faq-managing-ai-coding-agents">FAQ: managing AI coding agents</h2>
<p><strong>Why does working with AI coding agents feel like leadership instead of coding?</strong>
Because the unit of work changed from writing code to delegating and verifying it — roughly 60% AI usage against 0-20% full delegation in Anthropic&rsquo;s 2026 report — and the skills involved (decomposition, expectation-setting, trust calibration, review) are the same ones that predict success managing human teams, per NBER Working Paper 33662&rsquo;s rho = 0.81.</p>
<p><strong>Is managing AI agents really the same skill as managing people?</strong>
The overlap is large but not total. The raw correlation is rho = 0.81 and rho = 0.69 after controlling for hard skills. What does not transfer: agents do not negotiate roles, self-organize, or provide one another psychological safety, so coordination stays centralized on the human — closer to a conductor than a jazz-ensemble leader, as Fortune&rsquo;s reading of the Deming research puts it.</p>
<p><strong>How many AI coding agents should one developer manage?</strong>
Bound it by review capacity, not model capacity. Practitioner ceilings cluster at three to five concurrent agents for meaningful work, or five to seven CLI agents on a laptop before rate limits, merge conflicts, and review debt consume the gains. Anthropic&rsquo;s harness ships a 20-concurrent default with a &ldquo;fewer than 15 agents&rdquo; guideline. If nobody is reviewing, the fleet is generating obligations rather than value.</p>
<p><strong>How do I decide which tasks to fully delegate?</strong>
Use delegation levels rather than a binary: plan mode is Consult, auto-accept is Advise, full bypass is Delegate. Delegate mechanical, crisply specified, independently testable work. Keep ownership of product intent, architecture, shared interfaces, security-sensitive changes, and &ldquo;should we build this?&rdquo; decisions. Raise autonomy per task class only after that class earns it — Grove&rsquo;s task-relevant maturity, unchanged since 1983.</p>
<p><strong>What is the delegation gap?</strong>
It is the distance between how much you use AI and how much you can hand over without line-by-line review. Anthropic&rsquo;s 2026 report puts usage around 60% of work and full delegation at 0-20% of tasks. It is a trust and verification problem, not a model-capability problem: closing it requires automated evaluation, explicit escalation rules, and acceptance evidence — not just a better model.</p>
]]></content:encoded></item></channel></rss>