<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Zig Coding Agent on RockB</title><link>https://baeseokjae.github.io/tags/zig-coding-agent/</link><description>Recent content in Zig Coding Agent on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 05:15:34 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/zig-coding-agent/index.xml" rel="self" type="application/rss+xml"/><item><title>fx Coding Agent Review 2026: A Tiny, Open, Native Harness in Zig</title><link>https://baeseokjae.github.io/posts/fx-tiny-open-native-coding-agent/</link><pubDate>Thu, 01 Oct 2026 05:15:34 +0000</pubDate><guid>https://baeseokjae.github.io/posts/fx-tiny-open-native-coding-agent/</guid><description>fx coding agent review: Vercel Labs&amp;#39; Zig harness ships a ~6.2 MiB native binary, four interfaces over one core, and a ~7K-token request floor.</description><content:encoded><![CDATA[<p>The fx coding agent is Vercel Labs&rsquo; experimental coding-agent harness and CLI, written in Zig and open sourced under Apache-2.0 in August 2026. It ships as a single native binary of roughly 6.2 MiB on Apple Silicon — no Node.js, no Python, no <code>node_modules</code> — and drives four interfaces (shell CLI, one-shot JSON, an ACP server, and a WebAssembly SDK) from one agent core. As of 2026-10-01 it is at 3,237 GitHub stars, v0.0.12, and is still explicitly labeled experimental with breaking changes between releases.</p>
<h2 id="what-is-fx-and-what-do-tiny-open-native-actually-mean">What Is fx, and What Do &ldquo;Tiny, Open, Native&rdquo; Actually Mean?</h2>
<p>The three words in the tagline are doing marketing work, but each maps to a concrete engineering decision you can verify in the repository.</p>
<p><strong>Tiny</strong> refers to the runtime footprint, not the codebase. The macOS arm64 artifact is about 6.2–6.4 MiB, the Linux x86-64 build is 11.87 MB, and the README states a 7.8 MiB production ceiling enforced by release qualification on macOS arm64. Note the gap: the homepage advertises 6.39 MiB while the README quotes 7.8 MiB, because one is a representative Apple Silicon download and the other is a ceiling that every target must stay under. If you are sizing a container image, plan for 10–12 MB on Linux x86-64, not 6 MB.</p>
<p><strong>Open</strong> means the full Zig source under Apache-2.0 — not a source-available license with a commercial carve-out. That matters less for cost than for auditability: the system prompt, the tool definitions, and the permission logic are all readable, which is unusual among shipping coding agents.</p>
<p><strong>Native</strong> means machine code with no language runtime between you and the process, plus first-class WebAssembly targets. There is no interpreter startup, no package resolution step, and no dependency tree to reconcile on a build box.</p>
<p>For comparison, the same claim measured against an incumbent: <code>fx --help</code> completes in 2.7 ms on macOS arm64 where <code>claude --help</code> takes 112 ms, and the Claude Code installation it is being compared against consumes 197 MB. That is not a like-for-like comparison of capability, but it is a fair comparison of deployment cost, and deployment cost is the argument fx is actually making.</p>
<h2 id="fx-at-a-glance-version-license-binary-size-and-project-trajectory">fx at a Glance: Version, License, Binary Size, and Project Trajectory</h2>
<p>Project facts checked against the GitHub API and npm registry on 2026-10-01:</p>
<table>
  <thead>
      <tr>
          <th>Attribute</th>
          <th>Value (checked 2026-10-01)</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Repository created</td>
          <td>2026-08-11</td>
      </tr>
      <tr>
          <td>First public tags</td>
          <td>2026-08-17 to 2026-08-18</td>
      </tr>
      <tr>
          <td>First HN launch thread</td>
          <td>2026-08-18, 318 points</td>
      </tr>
      <tr>
          <td>Stars / forks / open issues</td>
          <td>3,237 / 364 / 250</td>
      </tr>
      <tr>
          <td>Language</td>
          <td>Zig (29.8 MB of source)</td>
      </tr>
      <tr>
          <td>License</td>
          <td>Apache-2.0</td>
      </tr>
      <tr>
          <td>Latest release</td>
          <td>v0.0.12, tagged 2026-09-30</td>
      </tr>
      <tr>
          <td>Release cadence</td>
          <td>8 releases in ~6 weeks (v0.0.5 → v0.0.12)</td>
      </tr>
      <tr>
          <td>Contributors</td>
          <td>26 (top: 2,449 commits; second: 82)</td>
      </tr>
      <tr>
          <td>npm package</td>
          <td><code>libfx</code>, 84 published versions, latest 0.0.12</td>
      </tr>
  </tbody>
</table>
<p>A project at 3,237 stars six weeks after its first tag is moving fast, but it is a fraction of the incumbents. The same day&rsquo;s check found opencode at 211,204 stars, claude-code at 148,751, codex (Rust) at 127,446, and pi at 110,818. fx has roughly 1.5% of opencode&rsquo;s community. That is not a reason to avoid it — early adoption is where the leverage is — but it does mean your questions will be answered by a small team and 250 open issues, not a forum with a decade of accumulated answers.</p>
<p>The velocity is also a warning label. Eight releases in six weeks means the provider authentication flow, permission behavior, and command surface all changed substantially within days of launch. The v0.0.12 notes are a good illustration of two things at once: <code>fx sessions</code> against a 14,600-session store dropped from 95 seconds to 0.17 seconds (a 560x improvement, and evidence that the earlier version was unusable at scale), and libfx error codes for unsupported model settings were renamed — a breaking change to a public SDK in a point release.</p>
<p>Binary size by target, from the v0.0.5 release assets:</p>
<table>
  <thead>
      <tr>
          <th>Target</th>
          <th>Size</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>macOS arm64</td>
          <td>6.43 MB (6.19 MiB measured)</td>
      </tr>
      <tr>
          <td>Linux arm64</td>
          <td>10.13 MB</td>
      </tr>
      <tr>
          <td>Linux x86-64</td>
          <td>11.87 MB</td>
      </tr>
      <tr>
          <td>macOS x86-64</td>
          <td>12.31 MB</td>
      </tr>
  </tbody>
</table>
<p>There are no Windows artifacts, which rules fx out for a meaningful share of enterprise developers regardless of everything else in this review.</p>
<h2 id="is-the-10-microsecond-cold-start-real-what-the-benchmarks-actually-measure">Is the &ldquo;10-Microsecond Cold Start&rdquo; Real? What the Benchmarks Actually Measure</h2>
<p>The 10-microsecond figure is the most quoted and least useful number in the launch material. It measures the time before fx accepts input — the point at which the process is alive and the terminal is ready for keystrokes. It does not measure the time before anything useful happens, and it is not the number the project enforces on itself.</p>
<p>What the project actually enforces is a 2 ms mean wall-clock budget. With <code>FX_BENCH=1</code>, fx parses arguments, dispatches the CLI path, and exits before TTY initialization; the Linux CI job runs six command paths (<code>fx</code>, <code>help</code>, <code>status --json</code>, <code>background --json</code>, <code>doctor --json</code>, <code>sessions --json</code>) 100 times each and fails the pull request if the mean exceeds 2 ms. That is the honest engineering claim: latency as a merge gate. A regression in startup cost cannot be merged, which is a stronger guarantee than any benchmark screenshot.</p>
<p>Then there is the number that a user actually experiences. Launch to first model request measures 31 ms on macOS arm64 — still an order of magnitude faster than the 112 ms it takes <code>claude --help</code> to print a help screen, but 3,100 times slower than 10 microseconds. Any article repeating &ldquo;10 microseconds&rdquo; without that context is repeating a framing, not a measurement.</p>
<h3 id="where-does-startup-latency-actually-matter">Where does startup latency actually matter?</h3>
<p>It matters when you start agents rather than run them. Hierarchical review fleets of 50–100 agents, one agent per pull request in CI, evaluation harnesses that spin up a fresh agent per test case, and disposable sandboxes that provision a coding agent for a single task — in all of those, process startup is paid hundreds of times per hour, and the difference between a 6 MB binary and a 250 MB installation is the difference between a container that fits and one that does not. One practitioner running review fleets at that scale framed it bluntly: 6 MB versus 250 MB is &ldquo;can do&rdquo; versus &ldquo;cannot.&rdquo;</p>
<p>If you are a human typing <code>fx</code> once and working for two hours, startup latency is noise. Buy fx for something else, or do not buy it.</p>
<h2 id="four-interfaces-over-one-core-cli-fx-ask-json-fx-acp-and-libfx">Four Interfaces Over One Core: CLI, fx ask &ndash;json, fx acp, and libfx</h2>
<p>This is the part of the product that has no direct equivalent in the incumbent harnesses, and it is the real reason to pay attention.</p>
<p><strong>1. Interactive shell.</strong> fx behaves closer to a Unix shell than an IDE embedded in a terminal. Sessions auto-name the terminal tab (session name, workspace fallback, model as context), so multi-window work stays legible; <code>fx sessions</code>, <code>fx session resume last</code>, and <code>fx session resume --id &lt;id&gt;</code> form a proper command group.</p>
<p><strong>2. One-shot JSON.</strong> <code>fx ask --json</code> runs a single request and returns machine-readable output. For CI and scripting this removes the entire category of TUI-scraping hacks that teams otherwise write to drive a coding agent from a pipeline.</p>
<p><strong>3. ACP server.</strong> <code>fx acp</code> speaks the Agent Client Protocol over stdio, letting an editor or host own the interface while fx owns sessions, prompts, tool calls, and permission requests. Vercel shipped <code>@ai-sdk/harness-fx</code> to connect fx into the AI SDK harness layer over ACP, which means the agent can be one component in a larger orchestrated system rather than the system itself.</p>
<p><strong>4. WebAssembly SDK.</strong> <code>libfx</code> exposes <code>createFxAgent()</code> on fx-core.wasm (headless) and <code>createFxTerminal()</code> on fx-term.wasm (interactive), both embeddable in JavaScript hosts, including a browser tab. The live demo at fx.sh/try runs the agent in the page. The browser build requires JavaScript Promise Integration (Chrome/Edge 137+).</p>
<p>The comparison to the nearest minimalist competitor is instructive: pi has print mode, JSON, RPC, and a TypeScript SDK, but it needs a JavaScript runtime and cannot run in a browser at all. If embedding an agent in a web page is your requirement, the field narrows to exactly one mature option.</p>
<p>The Wasm surface is not the native surface. It omits native process execution, OS sandboxing, native MCP, subagents, skills, auto-upgrade, arbitrary WASI filesystem access, and web search. Embeddability is bought with capability, and the buyer should know which capabilities they are giving up.</p>
<h2 id="how-much-does-the-fx-harness-cost-per-request">How Much Does the fx Harness Cost Per Request?</h2>
<p>Megabytes are a one-time cost. Tokens are a per-turn cost, and a harness&rsquo;s system prompt and tool definitions are charged on every single request. This is where &ldquo;tiny&rdquo; stops being a marketing word and becomes a budget line.</p>
<p>An independent measurement pointed the undocumented <code>FX_GATEWAY_BASE_URL</code> / <code>FX_GATEWAY_CHAT_URL</code> environment variables at a local mock gateway and captured the exact request a two-word prompt produced in a clean environment:</p>
<table>
  <thead>
      <tr>
          <th>Component</th>
          <th>Bytes</th>
          <th>Share</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>17 tool definitions</td>
          <td>17,376</td>
          <td>68.2%</td>
      </tr>
      <tr>
          <td>Base system prompt</td>
          <td>5,531</td>
          <td>21.7%</td>
      </tr>
      <tr>
          <td>Runtime context</td>
          <td>2,456</td>
          <td>9.6%</td>
      </tr>
      <tr>
          <td>User prompt (&ldquo;review this&rdquo;)</td>
          <td>59</td>
          <td>0.2%</td>
      </tr>
      <tr>
          <td>Total</td>
          <td>25,479</td>
          <td>100%</td>
      </tr>
  </tbody>
</table>
<p>That is roughly 6,400–7,300 tokens of harness overhead before your prompt is considered. The comparable counted figure for Claude Code on the same measurement basis is 29,061 tokens. The gap is roughly 4x, and it is not because fx&rsquo;s system prompt is cleverer — it is because fx ships 17 tool definitions instead of dozens, and because MCP server schemas are not inlined.</p>
<p>The detail worth stealing regardless of which agent you run: fx ships <code>capability_search</code> and <code>mcp_select_tool</code> so that an MCP server&rsquo;s schemas load only when a tool is actually selected. Lazy tool loading is the single largest lever on per-request overhead in any agent with more than a dozen integrations, and a 68% tool-definition share shows what happens when you get it wrong.</p>
<p>You can reproduce the measurement yourself, and for reviewers who never install fx, the recipe is worth knowing: set the gateway base URL to a local mock, send one two-word prompt, and count bytes. Vendor claims are cheap; byte counts are not.</p>
<h2 id="why-does-fx-read-claudeskills-the-skill-catalog-tax">Why Does fx Read ~/.claude/skills? The Skill-Catalog Tax</h2>
<p>Here is the hidden cost. The same two-word prompt, measured in a normal user&rsquo;s HOME directory rather than a clean one, balloons from 25,479 bytes to 42,727 bytes — a 68% increase, with 16,247 of those extra bytes arriving as a single system message that lists other agents&rsquo; skills.</p>
<p>fx scans 12 skill directories, including <code>~/.claude/skills</code>, <code>~/.codex/skills</code>, <code>~/.config/opencode/skills</code>, <code>~/.agents/skills</code>, and <code>~/.claw/skills</code>, and ships a catalog of skill names, descriptions, and paths on every turn. The catalog is capped by <code>skill_catalog_bytes</code> (default 16 KiB) and can be disabled entirely with <code>off</code>.</p>
<p>Two takeaways. First, if you already run Claude Code or Codex, fx inherits their skill directories by design — a deliberate interoperability choice, and also a per-turn tax you were not told about at install time. Second, if your per-request context budget is tight, <code>skill_catalog_bytes: off</code> is the first configuration change to make, before you touch the model.</p>
<h2 id="which-models-can-fx-use-and-where-does-billing-flow">Which Models Can fx Use, and Where Does Billing Flow?</h2>
<p>The default credential resolution order is: Vercel OIDC token → <code>AI_GATEWAY_API_KEY</code> → <code>fx login</code> OAuth → a saved API key. In practice that means every documented default path bills through Vercel AI Gateway, whose compiled default model is <code>moonshotai/kimi-k3</code>.</p>
<p>That is a real buying consideration, not a footnote. &ldquo;Model-agnostic&rdquo; here means the exits are documented and work, but the gravity is Vercel&rsquo;s gateway. The documented exits are:</p>
<ul>
<li><strong>Codex via ChatGPT subscription</strong> — <code>fx login codex</code>, OAuth, tokens stored locally at <code>~/.fx/chatgpt-auth.json</code> and never sent through the gateway.</li>
<li><strong>Grok via X Premium / SuperGrok</strong> — <code>fx login grok</code>, xAI OAuth, stored at <code>~/.fx/grok-auth.json</code>, likewise never routed through the gateway.</li>
<li><strong>Custom OpenAI-compatible connections</strong> — named connections to any Chat Completions endpoint, which is how you reach Ollama, vLLM, or OpenRouter.</li>
</ul>
<p>Two caveats on the local-model story. Local inference is configured as a custom connection, not as a one-flag mode, which is a heavier setup path than users coming from Ollama-first tools expect. And subscription OAuth is a personal-account mechanism: the <code>/fast</code> command on supported Codex models consumes ChatGPT credits at the higher Fast rate, which is fine for an individual and awkward for a team with a corporate billing policy.</p>
<p>On privacy, the picture is defensible: there is no product telemetry endpoint, and the auto-update check reads static release metadata every 30 minutes with no machine or installation identifier. It can be disabled with <code>FX_AUTO_UPGRADE=0</code>. Gateway-side, Vercel records model, token counts, latency, and cost but does not retain prompts after a request completes.</p>
<h2 id="do-fx-permissions-isolate-anything-permissions-vs-sandboxing">Do fx Permissions Isolate Anything? Permissions vs Sandboxing</h2>
<p>This is the section to read carefully, because conflating the two is how teams ship a vulnerability.</p>
<p>fx has three permission modes: <code>ask</code>, <code>auto</code> (default), and <code>full-access</code>. Auto mode runs routine work directly and escalates unresolved sensitive actions to a narrow safety-reviewer model — the same classifier shape, with the same caveats, as Claude Code&rsquo;s auto mode. Non-interactive runs (piped or redirected stdin) stay non-interactive and fail rather than waiting for an approval that can never arrive. That last behavior is correct for CI and surprising for anyone who assumed a prompt would block.</p>
<p>Permissions are a policy layer. They decide whether fx will attempt an action. They are not an isolation boundary. At the verified revision, the source says an absent sandbox setting means no sandbox, and changing permission mode does not select one. macOS has an optional OS sandbox; the source states no OS sandbox implementation was available on unsupported hosts, including Linux, in that release.</p>
<p>The practical consequence: on Linux, fx executes commands and touches files with the authority of the process it runs as. Secrets visible to that process are visible to the agent and to anything the agent invokes. If your threat model includes prompt injection from repository content, the mitigation is a container or a VM, not a permission mode.</p>
<p>Two more surfaces with the same trap. The Node addon <code>libfx</code> is a <code>.node</code> file — executable native code holding the host process&rsquo;s authority; N-API is not a sandbox, even though the headless core advertises no native tools and cannot launch commands or read workspace files unless the host grants that capability. And the skeptical framing from the security write-up deserves to be repeated: the launch materials do not establish that fx completes coding tasks faster, cheaper, more accurately, or more safely than competing agents, and no public fx-specific comparative benchmark existed at verification time. Smaller binary is a proven claim. Better agent is not.</p>
<h2 id="fx-vs-claude-code-vs-pi-vs-opencode-which-harness-fits">fx vs Claude Code vs pi vs opencode: Which Harness Fits?</h2>
<p>The useful taxonomy in this space is not &ldquo;which agent is best&rdquo; but &ldquo;how much does the harness decide for you.&rdquo; pi (earendil-works/pi, MIT, TypeScript) minimizes features: it refuses MCP, subagents, permission popups, plan mode, built-in todos, and background bash by design, shipping four default tools — read, write, edit, bash. fx minimizes footprint while keeping the features. Claude Code minimizes assembly time by shipping a finished platform.</p>
<table>
  <thead>
      <tr>
          <th>Harness</th>
          <th>Optimization</th>
          <th>Default posture</th>
          <th>Best fit</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Claude Code</td>
          <td>Assembly time</td>
          <td>Full platform, opinionated</td>
          <td>Teams that want a finished product now</td>
      </tr>
      <tr>
          <td>pi</td>
          <td>Features</td>
          <td>Four tools, refuses the rest</td>
          <td>Developers who want the smallest possible decision surface</td>
      </tr>
      <tr>
          <td>fx</td>
          <td>Footprint</td>
          <td>26 documented tools, MCP, subagents, skills — compressed runtime</td>
          <td>Embedding, CI, per-task agents, browser/editor hosts</td>
      </tr>
      <tr>
          <td>opencode</td>
          <td>Community breadth</td>
          <td>Large plugin ecosystem</td>
          <td>Teams optimizing for ecosystem size</td>
      </tr>
      <tr>
          <td>tmux + scripts</td>
          <td>Nothing</td>
          <td>No harness</td>
          <td>You are building the harness yourself</td>
      </tr>
  </tbody>
</table>
<p>On the spectrum from finished product to raw component, fx sits below pi despite having more features, because fx is explicitly built to be embedded in larger systems. That is the single most important sentence in this review: fx&rsquo;s feature list is not competing with Claude Code&rsquo;s; it is competing with <code>tmux</code> plus a shell script plus a provider API client.</p>
<p>The strongest critique of fx is a feature complaint rather than a footprint complaint: 26 tools with a dedicated tool for nearly every file operation is more surface area than a modern model needs, and the argument &ldquo;just use fewer tools&rdquo; has real merit. The counter-argument is that fx&rsquo;s lazy MCP loading means the extra built-ins cost 17 definitions in the base payload, not 40. Both positions are defensible; measure your own payload before choosing.</p>
<p>A note on honesty in the minimalism ledger: fx is small in runtime surface, not in codebase size. The <code>src/</code> tree holds roughly 693k lines of Zig across 560 files. A 6 MB binary is a compilation artifact, not evidence of a simple program.</p>
<h2 id="how-do-you-install-and-configure-fx">How Do You Install and Configure fx?</h2>
<p>The verified install path is a single script: <code>curl -fsSL https://fx.sh/setup.sh | bash</code>, which installs to <code>~/.local/bin</code>. Building from source requires Zig 0.16.0+ and <code>zig build -Doptimize=ReleaseSafe</code>.</p>
<p>Authentication has exactly three documented routes: <code>fx login</code> for Vercel AI Gateway OAuth, <code>fx setup</code> for a Gateway API key, or <code>fx login codex</code> / <code>fx login grok</code> for subscription-based access.</p>
<p>Beyond auth, the configuration surfaces that matter in practice:</p>
<ul>
<li><strong>Sessions</strong> — persistent, resumable, auto-named per terminal tab; deterministic replay is supported.</li>
<li><strong>AGENTS.md</strong> — scoped resolution, so project-level agent instructions apply without a global config file.</li>
<li><strong>MCP</strong> — servers are selected on demand through <code>capability_search</code> and <code>mcp_select_tool</code> rather than inlined into every request.</li>
<li><strong>Skills</strong> — installable capabilities plus the cross-agent skill directories discussed above; cap or disable the catalog via <code>skill_catalog_bytes</code>.</li>
<li><strong>Subagents</strong> — session-backed, for delegating narrow work without spawning a second process.</li>
<li><strong>Memory</strong> — persists to <code>~/.fx/memories.json</code> and is retrieved on demand rather than injected into every request, which is the right call for token economics.</li>
<li><strong>Compaction</strong> — in v0.0.12, a prompt-too-long error triggers compaction plus one retry instead of failing the turn outright.</li>
</ul>
<p>Two design details worth praising because they are counter to the grain: <code>semantic_search</code> is explicitly lexical rather than an embedding index, and fx does not pretend otherwise; large tool results are held out behind byte-range handles read via <code>read_tool_result</code> instead of being pasted wholesale into context. Both are context-discipline decisions that reduce silent token burn.</p>
<h2 id="who-should-use-fx-today--and-who-should-wait">Who Should Use fx Today — and Who Should Wait?</h2>
<p>Adopt fx now if you are building CI images, disposable evaluation containers, agent-per-task sandboxes, editor or ACP hosts, or anything that embeds an agent inside a larger program. The deployment argument is unambiguous: one executable removes <code>node_modules</code>, a virtualenv, a language package manager, and a separately provisioned runtime from the image, and upgrades and rollbacks become one file plus a checksum. When you spawn a hundred agents an hour, that is the whole story.</p>
<p>Wait, or plan around the gaps, if any of the following apply:</p>
<ul>
<li><strong>You need Windows.</strong> There are no Windows artifacts.</li>
<li><strong>You need isolation on Linux.</strong> There is no OS sandbox; run fx inside a container or VM.</li>
<li><strong>You need stability guarantees.</strong> It is v0.0.x with breaking changes, including renames in the public libfx SDK.</li>
<li><strong>You need the documented tool list to match the binary.</strong> The docs list about 26 tools while a captured v0.0.7 payload contained 17, with different names (the docs say <code>shell</code>, the payload said <code>terminal</code>; web search arrived as <code>perplexity_search</code>). Take the docs as intent and the payload as fact.</li>
<li><strong>You need local inference as a first-class mode.</strong> It works, through a documented OpenAI-compatible custom connection, not a single flag.</li>
<li><strong>You need proof it codes better.</strong> Nobody has published that proof, including Vercel.</li>
</ul>
<h2 id="verdict-an-embeddable-agent-runner-not-a-daily-driver">Verdict: An Embeddable Agent Runner, Not a Daily Driver</h2>
<p>fx is the most interesting thing that happened to coding-agent deployment in 2026, and it is not yet a replacement for your current assistant. Reviewers converged on the same conclusion from different angles: excellent as an embeddable harness, not yet a daily driver.</p>
<p>Judge it as what it is — an open implementation of a disposable, embeddable agent runner. The startup budget is enforced in CI. The request payload is a quarter of the incumbent&rsquo;s. The four interfaces over one core are a genuine category difference, and the WebAssembly surface has no equivalent in the minimalist competition. Against that: 250 open issues, no Windows, no Linux sandbox, v0.0.x churn, and a documented-versus-actual tool mismatch.</p>
<p>The right test is narrow and cheap. Point fx at a mock gateway, count the bytes your own configuration actually sends, and try <code>fx ask --json</code> in one pipeline that currently scrapes a TUI. If those two experiments pay for themselves, you have an answer. If your work is one developer in one terminal for two hours, the 10-microsecond number will never appear on your invoice.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Is the fx coding agent production-ready in 2026?</strong>
No, and Vercel does not claim otherwise. It is labeled experimental, it is on v0.0.x, and breaking changes ship in point releases — libfx error codes were renamed in v0.0.12. It is production-usable for embedded, CI, and per-task agent workloads where you control the version, and it is not a stable daily-driver replacement for a mature coding assistant.</p>
<p><strong>How large is the fx binary, and does size actually matter?</strong>
For fx, plan for about 6.2–6.4 MiB on macOS arm64 and 10.13–11.87 MB on Linux, against a 7.8 MiB README production ceiling and no Windows build. Size matters when you start agents rather than run them: one executable removes Node, Python, and package managers from CI images, and a 50–100 agent review fleet becomes feasible where a 250 MB installation per agent is not.</p>
<p><strong>How much context does fx send before my prompt?</strong>
Roughly 25.5 KB — about 6,400–7,300 tokens — for a two-word prompt in a clean environment, of which 68.2% is 17 tool definitions. In a populated HOME directory the same request grows to 42.7 KB because fx scans 12 skill directories, including <code>~/.claude/skills</code>, and ships a skill catalog each turn. Cap it with <code>skill_catalog_bytes</code> or disable it with <code>off</code>.</p>
<p><strong>Does fx sandbox the commands it runs?</strong>
No, not in the sense of isolation. fx has <code>ask</code>, <code>auto</code>, and <code>full-access</code> permission modes, which are policy rather than a boundary; the source states that an absent sandbox setting means no sandbox and that no OS sandbox implementation was available on Linux in the verified release. The Node addon has the host process&rsquo;s authority. Treat fx as unsandboxed and wrap it in a container or VM.</p>
<p><strong>Can fx run local models like Ollama, and does it work offline?</strong>
Yes to local models, through a named custom connection to any OpenAI-compatible Chat Completions endpoint (Ollama, vLLM, OpenRouter) rather than a dedicated local mode. Not fully offline: the default path bills through Vercel AI Gateway with <code>moonshotai/kimi-k3</code> as the compiled default model, and auto-update checks read release metadata every 30 minutes unless you set <code>FX_AUTO_UPGRADE=0</code> or use Codex/Grok subscription OAuth.</p>
]]></content:encoded></item></channel></rss>