<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Document Framework for Ai Agents on RockB</title><link>https://baeseokjae.github.io/tags/document-framework-for-ai-agents/</link><description>Recent content in Document Framework for Ai Agents on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 30 Sep 2026 04:15:07 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/document-framework-for-ai-agents/index.xml" rel="self" type="application/rss+xml"/><item><title>Open Doc: The Document Framework Built for AI Agents</title><link>https://baeseokjae.github.io/posts/open-doc-document-framework-for-agents/</link><pubDate>Wed, 30 Sep 2026 04:15:07 +0000</pubDate><guid>https://baeseokjae.github.io/posts/open-doc-document-framework-for-agents/</guid><description>open-doc is a React-first document framework for AI agents with real A4 page geometry and a layout checker — but 78 stars and no .docx export yet.</description><content:encoded><![CDATA[<p>open-doc is an MIT-licensed, React-first document framework (<code>@open-document/core</code>) in which a coding agent writes each report as React components while the framework owns the paper: A4/B4/A3 page geometry, self-filling contents, page numbers, DOM-measured auto-pagination, layout diagnostics, and headless PDF/HTML/PNG/SVG export. It is genuinely early — 78 GitHub stars, one human contributor and 487 npm downloads last month (2026-09-30).</p>
<p>That gap is the whole story. If you are evaluating it, you are not choosing between a finished product and a finished product — you are deciding whether a well-argued, correctly-scoped design is worth reading end to end and trialling in an existing React/TypeScript shop.</p>
<h2 id="what-open-doc-actually-is-and-the-one-sentence-that-explains-it">What open-doc actually is (and the one sentence that explains it)</h2>
<p>The project&rsquo;s own description string is the article&rsquo;s title: <em>&ldquo;The document framework built for agents.&rdquo;</em> Concretely:</p>
<ul>
<li><strong>Authored medium:</strong> TSX. One document is one directory, <code>docs/&lt;id&gt;/index.tsx</code>.</li>
<li><strong>Runtime:</strong> a React renderer that renders a stack of real pages in the browser, plus a studio viewer with a live text inspector.</li>
<li><strong>Agent surface:</strong> 23 MCP tools mounted at <code>http://localhost:5273/mcp</code>, plus a set of skills that ship inside the scaffolder.</li>
<li><strong>Verification surface:</strong> <code>open-doc check</code>, a CLI that renders every page at true size and reports what broke, with source locations, exiting non-zero.</li>
<li><strong>Output:</strong> PDF (browser print pipeline), self-contained HTML, per-page PNG and SVG.</li>
</ul>
<p>The one sentence that explains the design is the author&rsquo;s own analogy from the README: <strong>&ldquo;If open-slide is Google Slides for agents, open-doc is Google Docs.&rdquo;</strong> A slide deck is a single 1920×1080 canvas — overflow is invisible because there is nowhere to overflow to. A document is a stack of A4 sheets that has to survive a printer, so it needs a page-break algorithm and, critically, a way to <em>tell someone when the algorithm failed</em>. That difference is why open-doc ships a layout checker and open-slide does not.</p>
<h2 id="the-problem-it-names-correctly--agents-write-well-and-format-terribly">The problem it names correctly — agents write well and format terribly</h2>
<p>Most &ldquo;AI document generation&rdquo; tooling quietly assumes the hard part is prose. open-doc&rsquo;s opening argument is that the hard part is that the author cannot see the page. From the README, verbatim:</p>
<blockquote>
<p>&ldquo;An agent writing React has no idea whether the paragraph it just added pushed the last three lines off the sheet.&rdquo;</p></blockquote>
<p>That is a precise diagnosis, and it explains a failure mode anyone shipping agent-generated reports recognises: the model produces fluent, well-organised text, then the PDF arrives with a table split across two sheets, a heading stranded alone at the bottom of page 6, an image that silently failed to load, and three orphaned lines clipped past the trim edge. Nothing in the loop told the model. It wrote a plausible document and had no feedback channel to learn it was wrong.</p>
<p>The reframe is worth stating plainly because it is not obvious from the feature list: <strong>open-doc&rsquo;s real product is the feedback loop, not the renderer.</strong> The renderer is a browser print pipeline — the same one your browser&rsquo;s &ldquo;Print to PDF&rdquo; uses. The differentiator is that the agent can query the layout and get back &ldquo;page 4 runs 37px past the sheet, at <code>p:41:12</code>&rdquo; and then fix it.</p>
<h2 id="react-as-the-authored-medium-the-file-contract">React as the authored medium: the file contract</h2>
<p>The scaffold produces a workspace, and every document inside it is ordinary source:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-tsx" data-lang="tsx"><span style="display:flex;"><span><span style="color:#75715e">// docs/q3-report/index.tsx
</span></span></span><span style="display:flex;"><span><span style="color:#75715e"></span><span style="color:#66d9ef">export</span> <span style="color:#66d9ef">const</span> <span style="color:#a6e22e">meta</span>: <span style="color:#66d9ef">DocMeta</span> <span style="color:#f92672">=</span> { <span style="color:#75715e">/* title, theme, author, … */</span> };
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#66d9ef">export</span> <span style="color:#66d9ef">default</span> [<span style="color:#a6e22e">Cover</span>, <span style="color:#a6e22e">Contents</span>, <span style="color:#a6e22e">flow</span>(&lt;&gt;, { <span style="color:#a6e22e">footer</span>: <span style="color:#66d9ef">Footer</span> })] <span style="color:#a6e22e">satisfies</span> <span style="color:#a6e22e">DocEntry</span>[];
</span></span></code></pre></div><p>Two authoring modes coexist in the same file: an array of <strong>fixed pages</strong> (<code>DocPage</code> entries, used for a cover or a contents page) and <code>flow(&lt;&gt;)</code> sections, which the framework paginates itself by measuring the real DOM rather than by estimating text heights.</p>
<p>The project states hard rules to the agent, and they are unusually disciplined: one <code>index.tsx</code> plus an <code>assets/</code> directory, <strong>no sibling <code>.tsx</code> files</strong>, no new dependencies, only <code>react</code> + <code>@open-document/core</code> + standard web APIs, and do not touch <code>package.json</code> or <code>open-doc.config.ts</code>. That last constraint matters more than it looks — it means an agent cannot &ldquo;fix&rdquo; a layout problem by adding a library, which is the single most common way agent-authored code degrades.</p>
<p>Why React at all, rather than a document DSL? The README&rsquo;s argument is that <strong>React is a medium, not a format</strong> — models have deep, heavily-trained priors on JSX, component composition and props, while a bespoke markup language is something a model must learn at inference time from a reference document. The bet is that fluent React beats awkward markup for the same author. The cost is equally real and worth naming: you inherit a browser print pipeline and its <code>@page</code> quirks, which the project has already hit once (a landscape <code>@page</code> descriptor bug fixed in 0.4.0).</p>
<h2 id="real-page-geometry-a4-jis-b4-and-a3-and-why-the-list-is-short">Real page geometry: A4, JIS B4 and A3, and why the list is short</h2>
<p>open-doc deliberately supports exactly six sheets: <strong>A4, JIS B4 (257 × 364 mm) and A3, each portrait or landscape.</strong> Version 0.4.0 <em>removed</em> Letter, A5 and Legal.</p>
<p>That reads like an omission until you read the rationale: every supported sheet maps to paper a given print shop actually stocks, so a document always maps onto something a human can physically buy and bind. A4 is laid out at <strong>794 × 1123 px @ 96 dpi</strong> with a matching <code>@page</code> descriptor, so nothing is rescaled at print time — the px grid the framework lays out against <em>is</em> the print grid. Print fidelity here is a first-class constraint, not a rendering detail, and it is the thing that most distinguishes open-doc from a &ldquo;generate an HTML file and hope&rdquo; workflow.</p>
<h2 id="auto-pagination-that-knows-what-not-to-break">Auto-pagination that knows what not to break</h2>
<p>Programmatic pagination is where most document generators embarrass themselves, usually by splitting a table down the middle. open-doc&rsquo;s packer carries explicit rules:</p>
<ul>
<li><strong>Headings never end a page</strong> — a heading that would land at the foot of a page moves with its content.</li>
<li><strong>Captions stay with their figures.</strong></li>
<li><strong>Tables move whole</strong> rather than splitting across sheets.</li>
</ul>
<p>These are the same typographic rules a human typesetter applies, encoded as invariants the runtime enforces. The framework&rsquo;s page-break decisions are therefore deterministic with respect to content, not dependent on where an LLM guessed a <code>page-break</code> should go.</p>
<h2 id="the-long-form-furniture-footnotes-figures-cross-references-and-tables-from-csv">The long-form furniture: footnotes, figures, cross-references and tables from CSV</h2>
<p>The parts of document production that agents usually get wrong — numbering — are handled by the runtime:</p>
<ul>
<li><strong><code>&lt;Footnote&gt;</code></strong> prints on whichever page its marker landed on, and the framework subtracts its height from that page&rsquo;s budget <em>before</em> the packer breaks the page. This is the detail that separates real pagination from decorative footnotes: the footnote can change where the page break falls.</li>
<li><strong><code>&lt;Ref&gt;</code></strong> renders &ldquo;Figure 3&rdquo; and adds &ldquo;(p. 12)&rdquo; only when the target is on another sheet.</li>
<li><strong><code>&lt;ListOfFigures /&gt;</code> / <code>&lt;ListOfTables /&gt;</code></strong> build the lists from the document itself.</li>
<li><strong><code>&lt;DataTable&gt;</code></strong> reads <code>.csv</code> / <code>.tsv</code> files at build time into arrays of objects, so table numbers stay synchronised with the prose around them, and numeric columns right-align with <code>tabular-nums</code> without being asked.</li>
</ul>
<p>Numbering that maintains itself is a genuine correctness feature, not a convenience. Every manual figure number is a bug waiting for the next revision.</p>
<h2 id="the-mcp-server--23-tools-stateless-streamable-http-and-409s-instead-of-clobbered-edits">The MCP server — 23 tools, stateless Streamable HTTP, and 409s instead of clobbered edits</h2>
<p>Running <code>open-doc dev --mcp</code> mounts <strong>23 tools</strong> over stateless Streamable HTTP at <code>http://localhost:5273/mcp</code>, with no session handshake. Any MCP client can drive the document — including reading it, editing text, and rendering pages.</p>
<p>The design decision worth highlighting is that <strong>the MCP tools and the browser UI share one implementation</strong>. That is a correctness property, not an architectural elegance: it means a tool call and a human click cannot diverge in behaviour.</p>
<p>Concurrency is handled with optimistic concurrency control. Text writes take an <code>expected</code> value — the content the agent last read — and a write against stale content is <strong>refused with HTTP 409 instead of silently clobbering a concurrent human edit</strong>. For a collaborative authoring loop where a person is editing in the studio while an agent edits through MCP, this is the difference between a safe tool and a data-loss generator. Comments persist as <code>@doc-comment</code> markers that a later <code>/apply-comments</code> pass picks up.</p>
<h2 id="open-doc-check-giving-a-blind-agent-eyes"><code>open-doc check</code>: giving a blind agent eyes</h2>
<p>This is the strongest, most checkable part of the project, and it is where the article&rsquo;s centre of gravity belongs.</p>
<p><code>open-doc check</code> renders every page at true size and reports a concrete failure taxonomy:</p>
<table>
  <thead>
      <tr>
          <th>What the checker finds</th>
          <th>Why it matters in print</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Content clipped past the sheet edge</td>
          <td>Lines silently disappear in the PDF — the most common agent failure</td>
      </tr>
      <tr>
          <td>Blank sheets</td>
          <td>Usually a stray page break or an empty <code>DocPage</code></td>
      </tr>
      <tr>
          <td>Headings stranded at a page foot</td>
          <td>Reads as a formatting error to any human reviewer</td>
      </tr>
      <tr>
          <td>Type too small to print legibly</td>
          <td>Survives on screen, fails on paper</td>
      </tr>
      <tr>
          <td>Images that failed to load</td>
          <td>Renders as an empty box, invisible to a text-only agent</td>
      </tr>
  </tbody>
</table>
<p>Crucially, each finding carries a <strong>source <code>line:column</code> locator</strong> — the README&rsquo;s sample output reads like <code>p.4 content runs 37px past the sheet</code>, <code>p.7 image failed to load</code>, <code>p.6 heading ends the page</code>, each pointing back at the offending line. And the command <strong>exits non-zero</strong>, which means the same tool that advises an agent doubles as a CI gate. A pipeline can refuse to publish a document that overflows. That is the concrete answer to &ldquo;how does an agent avoid clipped content on A4 pages&rdquo; — not better prompting, but a check the agent can act on and a build that fails when it does not.</p>
<h2 id="headless-export-and-markdown-import">Headless export and Markdown import</h2>
<p><strong>Out:</strong> PDF (browser print pipeline at true page size, so what you see is what prints), self-contained HTML (zipped when the document has assets), per-page PNG and SVG. Headless export needs an optional peer dependency — <strong>playwright + chromium</strong> — rather than a bundled browser. There is <strong>no <code>.docx</code> export</strong> (see limitations below).</p>
<p><strong>In:</strong> <code>open-doc import notes.md --id q3-notes --contents</code> converts Markdown into a real document — a <code>flow()</code> body, a cover, a self-filling contents page, GFM tables, and local images copied into the document&rsquo;s <code>assets/</code> — and what you get is <em>ordinary authored TSX</em> you can then edit by hand or by agent. That bidirectional story (import, co-author, export) is what makes it usable as a pipeline step rather than a one-way generator.</p>
<p>Deployment is deliberately boring: <code>open-doc build</code> emits a plain static site that drops onto Vercel, Cloudflare Pages, Netlify or any static host.</p>
<h2 id="install-and-ship-a-first-document-step-by-step">Install and ship a first document, step by step</h2>
<p>The real path, in order:</p>
<div class="highlight"><pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;"><code class="language-bash" data-lang="bash"><span style="display:flex;"><span><span style="color:#75715e"># 1. Scaffold a documents workspace</span>
</span></span><span style="display:flex;"><span>npx @open-document/cli init my-docs
</span></span><span style="display:flex;"><span>cd my-docs
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># 2. Studio + viewer on :5273 (also serves the MCP endpoint with --mcp)</span>
</span></span><span style="display:flex;"><span>pnpm dev
</span></span><span style="display:flex;"><span>open-doc dev --mcp          <span style="color:#75715e"># mounts 23 tools at http://localhost:5273/mcp</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># 3. Author: run /create-doc with your coding agent</span>
</span></span><span style="display:flex;"><span><span style="color:#75715e">#    (asks four scoping questions and establishes source material before writing)</span>
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># 4. Verify the layout — the step people skip and regret</span>
</span></span><span style="display:flex;"><span>open-doc check
</span></span><span style="display:flex;"><span>
</span></span><span style="display:flex;"><span><span style="color:#75715e"># 5. Ship</span>
</span></span><span style="display:flex;"><span>open-doc export --pdf       <span style="color:#75715e"># or --format html|png|svg</span>
</span></span><span style="display:flex;"><span>open-doc build              <span style="color:#75715e"># static site for your host</span>
</span></span></code></pre></div><p>The <code>/create-doc</code> skill is worth calling out as a design choice. It asks four scoping questions, insists on establishing source material first, and — in the project&rsquo;s own words — <strong>&ldquo;will not invent your numbers.&rdquo;</strong> Prompt engineering treated as versioned, shipped source, distributed as a skill rather than as documentation. The other shipped skills are <code>/doc-authoring</code> (the technical reference), <code>/current-doc</code> (resolves &ldquo;this page&rdquo; by reading <code>node_modules/.open-doc/current.json</code>), <code>/apply-comments</code> and <code>/create-theme</code>.</p>
<h2 id="the-2026-mcp-context--why-a-stateless-protocol-is-what-made-this-design-viable">The 2026 MCP context — why a stateless protocol is what made this design viable</h2>
<p>open-doc&rsquo;s MCP endpoint assumes a protocol shape that only became standard a few weeks before the project appeared. The <strong>MCP <code>2026-07-28</code> revision</strong> changed exactly the things that make a tool server like this simple to run:</p>
<ul>
<li><strong>Sessions removed</strong> (SEP-2567) — protocol-level sessions and the <code>Mcp-Session-Id</code> header are gone, and list endpoints no longer vary per connection.</li>
<li><strong>Handshake removed</strong> (SEP-2575) — the <code>initialize</code> / <code>notifications/initialized</code> dance is gone; every request carries protocol version and client capabilities in <code>_meta</code>.</li>
<li><strong><code>server/discover</code> became a mandatory RPC</strong> so clients can negotiate a version up front.</li>
<li><strong>Multi round-trip requests</strong> (SEP-2322) — servers return <code>resultType: &quot;input_required&quot;</code> with <code>inputRequests</code> instead of pushing server-initiated requests.</li>
<li><strong>SSE stream resumability and <code>Last-Event-ID</code> removed</strong> — a broken stream loses the in-flight request, and clients MUST re-issue with a new request id. That is directly relevant when an agent&rsquo;s long <code>render_page</code> call dies mid-render.</li>
<li><strong>Smaller changes that bite local tool servers:</strong> required <code>Mcp-Method</code> / <code>Mcp-Name</code> headers on Streamable HTTP POSTs (SEP-2243), <code>ttlMs</code> + <code>cacheScope</code> on list results (SEP-2549), deterministic <code>tools/list</code> ordering for prompt-cache hits, and resource-not-found moving from <code>-32002</code> to <code>-32602</code>.</li>
</ul>
<p>A stateless protocol means an MCP document server does not have to hold per-client state, re-establish anything, or reconcile a reconnect — it is just a request handler over the same implementation the browser uses. open-doc offers a <code>legacy: 'stateless'</code> mode that serves 2025-era clients from the same endpoint, which is a sensible hedge for anyone whose client has not caught up. For the broader protocol landscape, see our <a href="/posts/ai-agent-protocols-mcp-a2a-acp-2026/">MCP, A2A and ACP comparison</a> and the <a href="/posts/api-vs-mcp-difference-guide-2026/">plain-language MCP vs API explainer</a>.</p>
<h2 id="how-it-compares">How it compares</h2>
<p>All figures below were collected on <strong>2026-09-30</strong> from the GitHub and npm APIs; they move, so treat them as a snapshot. &ldquo;Downloads&rdquo; is npm downloads for the last month.</p>
<table>
  <thead>
      <tr>
          <th>Project</th>
          <th>Stars</th>
          <th>Downloads/mo</th>
          <th>What it is</th>
          <th>Best for</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>open-doc</strong> (<code>@open-document/core</code>)</td>
          <td>78</td>
          <td>487</td>
          <td>Agent-native document workspace: page geometry, checker, MCP tools, themes</td>
          <td>A person and an agent co-authoring a printable report in a React shop</td>
      </tr>
      <tr>
          <td><a href="/posts/ai-documentation-generator-tools-2026/">react-pdf</a> (<code>@react-pdf/renderer</code>)</td>
          <td>16,813</td>
          <td>22,574,136</td>
          <td>React-to-PDF renderer with its own layout engine</td>
          <td>High-volume programmatic PDFs inside an application</td>
      </tr>
      <tr>
          <td>Typst</td>
          <td>56,342</td>
          <td>12,920</td>
          <td>Markup-based typesetting compiler with fast incremental builds</td>
          <td>Deterministic, compiler-driven typesetting</td>
      </tr>
      <tr>
          <td>Pandoc</td>
          <td>46,452</td>
          <td>6,780</td>
          <td>Universal markup converter</td>
          <td>Markdown to <em>anything</em>, including real <code>.docx</code></td>
      </tr>
      <tr>
          <td>Quarto</td>
          <td>6,032</td>
          <td>—</td>
          <td>Scientific/technical publishing system on Pandoc</td>
          <td>Books, reports, reproducible research</td>
      </tr>
      <tr>
          <td>WeasyPrint</td>
          <td>9,649</td>
          <td>1,419</td>
          <td>HTML/CSS to PDF with real paged-media support</td>
          <td>Server-side HTML-to-print pipelines</td>
      </tr>
      <tr>
          <td>open-slide (sibling project)</td>
          <td>8,587</td>
          <td>—</td>
          <td>Agent-native <em>deck</em> framework, 1920×1080 canvas</td>
          <td>Slides, not documents</td>
      </tr>
      <tr>
          <td>MarkItDown</td>
          <td>187,655</td>
          <td>—</td>
          <td>Files and office documents <em>into</em> Markdown</td>
          <td>Ingestion for agents</td>
      </tr>
  </tbody>
</table>
<p>Three comparisons deserve a sentence each because they are the ones readers actually search for.</p>
<p><strong>open-doc vs react-pdf</strong> is not a contest — it is a category difference. react-pdf is a renderer with its own layout engine and a PDF-specific component set (<code>Document</code>, <code>Page</code>, <code>View</code>, <code>Text</code>); it is what you reach for when a service must produce PDFs at scale inside an existing app. open-doc is a document <em>workspace</em> with a viewer, inspector, themes, skills, an MCP endpoint and a layout checker that tells the agent <em>which page overflowed and at which line:column</em>. react-pdf has roughly 46,000× the downloads, and it does not give an agent a feedback loop about clipped content.</p>
<p><strong>open-doc vs open-slide</strong> is the same author-side philosophy in a different medium, and the star gap is instructive: 8,587 stars for the deck framework against 78 for the document framework. Same skill pattern (<code>/create-*</code> with four scoping questions, <code>@…-comment</code> markers, <code>/apply-comments</code>), but a deck exports an editable PPTX where each page becomes native text boxes and shapes — done entirely in the browser, no headless browser needed. Documents have a paginated paper contract, which is precisely why open-doc needs a checker and open-slide does not.</p>
<p><strong>open-doc vs Pandoc/Quarto/Typst</strong> is the fidelity-versus-breadth trade. The markup toolchain wins decisively on output formats — <code>.docx</code>, LaTeX, EPUB, ODT — and loses on layout feedback: none of them can tell an agent &ldquo;the table on page 7 overflows by 40px.&rdquo; open-doc&rsquo;s counter-argument to Typst is that agents write prose well and markup badly, and that React is a medium the model already has deep priors for; Typst&rsquo;s counter-argument is determinism — a compiler has no browser print pipeline, no font race and no <code>@page</code> descriptor that Chromium might drop.</p>
<p>One more framing worth keeping: <strong>&ldquo;agent documents&rdquo; is two problems.</strong> Ingestion (getting existing files into a form the model can read) and production (emitting a printable artifact). MarkItDown — at 187,655 stars, an order of magnitude bigger than anything else here — solves the first. open-doc only solves the second, and in a real pipeline they compose: MarkItDown ingests, the agent reasons, open-doc produces.</p>
<h2 id="limitations-and-open-issues">Limitations and open issues</h2>
<p>Read this section before you put it in front of a team.</p>
<ul>
<li><strong>No <code>.docx</code> export.</strong> Issue <a href="https://github.com/simonliu-ai-product/open-doc/issues/35">#35</a>, filed 2026-09-21, requests an editable Word file precisely because reviewers want track changes and Word comments. PDF is currently the end of the line. If your organisation&rsquo;s sign-off process runs in Word, that is a review-workflow problem, not a rendering problem, and it disqualifies open-doc today.</li>
<li><strong>The MCP escape hatch is currently broken.</strong> Issue <a href="https://github.com/simonliu-ai-product/open-doc/issues/34">#34</a> (2026-09-08): <code>allowedHosts</code> from <code>open-doc.config.ts</code> is never passed to the MCP plugin, so <code>/mcp</code> returns <code>403 Invalid Host</code> for any non-loopback host — including a Docker Compose service name. The endpoint validates <code>Host</code> and <code>Origin</code> against loopback (blocking DNS rebinding, 403 on both rejections) and authenticates nothing, so anything beyond loopback needs an authenticating reverse proxy. Do not expose it.</li>
<li><strong>Single-column viewer.</strong> Issue <a href="https://github.com/simonliu-ai-product/open-doc/issues/36">#36</a>: no two-up or grid view, despite fit-page / fit-width / 100% zoom controls.</li>
<li><strong>Adoption is minimal, and the curve is falling.</strong> Created 2026-08-17, 78 stars, 7 forks, 9 open issues, one human contributor (LiuYuWei, 12 commits; the other committers are <code>github-actions[bot]</code> and <code>dependabot[bot]</code>). 23 lifetime commits, 2 in the trailing four weeks. Core downloads: 976 over 45 days — <strong>459 in launch week, 23 in the last full week</strong>, a ~95% decay. Six core versions shipped in the first 18 days; the latest release is core 0.6.0 / mcp 0.3.2 on 2026-09-03.</li>
<li><strong>The marketing page is three minor versions stale</strong> — it still advertises &ldquo;v0.3.0 is on npm&rdquo; while npm serves 0.6.0. A small but honest signal of how thin the operation is.</li>
<li><strong>Naming collisions.</strong> <code>RyanYahya/OpenDoc</code> (created 2026-09-12, 2 stars) carries a near-identical description and appeared twelve days after this one went public. Searching &ldquo;OpenDoc&rdquo; also surfaces older, unrelated projects. Check which OpenDoc you found.</li>
</ul>
<p>Note also what the download split implies: <code>@open-document/mcp</code> pulled <strong>445 downloads</strong> last month against core&rsquo;s 487. The MCP surface is not a subset of the audience — it <em>is</em> the audience. People are reaching for this as an agent tool, not as a rendering library.</p>
<h2 id="who-should-use-open-doc--and-who-should-not">Who should use open-doc — and who should not</h2>
<p><strong>Use it if</strong> you have a recurring, human-reviewed printable report — monthly board pack, quarterly investor update, technical whitepaper — authored by a coding agent inside an existing React/TypeScript shop, where a person steers in chat while the agent writes, and where &ldquo;does this fit on the page&rdquo; currently costs a manual review cycle. The <code>open-doc check</code> CI gate and the &ldquo;no proprietary format, it diffs in git&rdquo; property are the real wins: your documents become readable, reviewable source that shows up in a pull request like anything else, instead of a binary <code>.docx</code> nobody can diff.</p>
<p><strong>Do not use it if</strong> you need high-volume server-side PDF generation (react-pdf), markup-first scientific publishing (Pandoc, Quarto, Typst), anything that must end as <code>.docx</code> with track changes, or a hosted service with an SLA. And do not deploy the MCP endpoint beyond loopback while issue #34 is open.</p>
<p>The candid summary: this is a well-argued design you can read end to end in an afternoon, with a genuinely novel contribution — a layout checker that closes the loop for a blind author. It is not yet a load-bearing dependency.</p>
<h2 id="faq">FAQ</h2>
<h3 id="what-is-open-doc-in-one-paragraph">What is open-doc in one paragraph?</h3>
<p>An MIT-licensed React-first document framework (<code>@open-document/core</code>) where a coding agent writes each report as React components while the framework supplies real A4/JIS B4/A3 page geometry, a self-filling contents page, running page numbers, DOM-measured auto-pagination, layout diagnostics and PDF/HTML/PNG/SVG export. Scaffold it with <code>npx @open-document/cli init my-docs</code>, then <code>open-doc dev</code> for the studio on port 5273.</p>
<h3 id="can-it-run-headless-in-ci-and-how-does-that-work">Can it run headless in CI, and how does that work?</h3>
<p>Yes. <code>open-doc check</code> renders every page at true size, reports clipped content, blank sheets, stranded headings, too-small type and failed images with source <code>line:column</code> locations, and exits non-zero — so the same command that advises an agent can fail a build. Headless export additionally needs <code>playwright</code> with chromium, an optional peer dependency rather than a bundled browser. See our notes on <a href="/posts/agent-codemode-mcp-scripts-2026/">code-mode MCP tooling</a> for how agents typically consume tool surfaces like this.</p>
<h3 id="does-it-export-to-word-docx">Does it export to Word (.docx)?</h3>
<p>No. Output is PDF, HTML, per-page PNG and SVG. Issue #35 requests <code>.docx</code> specifically because organisations that review in Word need track changes and comments. Until that lands a PDF is the end of the line; use Pandoc if Word is a hard requirement.</p>
<h3 id="which-agents-can-drive-it-and-does-it-need-a-browser">Which agents can drive it, and does it need a browser?</h3>
<p>Anything that writes files or speaks MCP. Skills ship inside the scaffolder (<code>/create-doc</code>, <code>/doc-authoring</code>, <code>/current-doc</code>, <code>/apply-comments</code>, <code>/create-theme</code>) for agents that read skill directories, and <code>open-doc dev --mcp</code> mounts 23 tools at <code>http://localhost:5273/mcp</code> for any MCP client. A browser is required — deliberately, since PDF export uses the browser print pipeline at true page size so that what you see is what prints. Those skills follow the pattern covered in the <a href="/posts/agent-skills-marketplace-guide-2026-claude-codex-cursor-and-gemini-cli/">agent skills marketplace guide</a>.</p>
<h3 id="how-mature-is-it-and-is-the-mcp-endpoint-safe-to-expose">How mature is it, and is the MCP endpoint safe to expose?</h3>
<p>Very early: created 2026-08-17, 78 stars, 7 forks, one human contributor, roughly 487 npm downloads a month for core, and a decay from 459 downloads in launch week to 23 in the last full week. On safety, no — not as shipped. The endpoint validates <code>Host</code> and <code>Origin</code> against loopback (403 on both) to block DNS rebinding and authenticates nothing, so anything beyond loopback needs an authenticating reverse proxy, and issue #34 means the intended <code>allowedHosts</code> escape hatch returns 403 anyway.</p>
<h2 id="sources-and-further-reading">Sources and further reading</h2>
<ul>
<li>open-doc repository, README, MCP package README, core CHANGELOG and authoring skill: <code>github.com/simonliu-ai-product/open-doc</code></li>
<li>Product page (advertising v0.3.0): <code>costaffs.app/tools/open-doc/</code></li>
<li>Issues <a href="https://github.com/simonliu-ai-product/open-doc/issues/35">#35</a> (<code>.docx</code>), <a href="https://github.com/simonliu-ai-product/open-doc/issues/34">#34</a> (<code>allowedHosts</code> / 403), <a href="https://github.com/simonliu-ai-product/open-doc/issues/36">#36</a> (two-up view)</li>
<li>npm registry and downloads API: <code>@open-document/core</code>, <code>@open-document/mcp</code>, <code>@open-document/cli</code></li>
<li>MCP specification <code>2026-07-28</code> changelog (SEP-2567, SEP-2575, SEP-2322, SEP-2243, SEP-2549)</li>
<li>Comparators: <code>react-pdf</code>, <code>typst/typst</code>, <code>jgm/pandoc</code>, <code>quarto-dev/quarto-cli</code>, <code>Kozea/WeasyPrint</code>, <code>microsoft/markitdown</code>, <code>open-slide/open-slide</code></li>
<li>Related reading on this site: <a href="/posts/ai-agent-protocols-mcp-a2a-acp-2026/">MCP, A2A and ACP in 2026</a>, <a href="/posts/ai-documentation-generator-tools-2026/">AI documentation generator tools</a>, <a href="/posts/ai-code-documentation-tools-2026/">AI code documentation tools</a>, <a href="/posts/agents-memory-cross-agent-context-mcp/">cross-agent context over MCP</a></li>
</ul>
]]></content:encoded></item></channel></rss>