<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Cerebras OpenAI 750MW Deal on RockB</title><link>https://baeseokjae.github.io/tags/cerebras-openai-750mw-deal/</link><description>Recent content in Cerebras OpenAI 750MW Deal on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 03:11:27 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/cerebras-openai-750mw-deal/index.xml" rel="self" type="application/rss+xml"/><item><title>Cerebras CS-4 Review: The Fastest Inference Hardware Yet?</title><link>https://baeseokjae.github.io/posts/cerebras-cs-4-inference-hardware/</link><pubDate>Thu, 01 Oct 2026 03:11:27 +0000</pubDate><guid>https://baeseokjae.github.io/posts/cerebras-cs-4-inference-hardware/</guid><description>Cerebras CS-4 review: three doubled-clock WSE-3 Turbo wafers hit 4,400 tokens/sec per user — about 12.6x a GPU service, not the claimed 30x.</description><content:encoded><![CDATA[<p>On decode-heavy, latency-bound workloads, yes: the Cerebras CS-4 is the fastest inference hardware shipping in 2026. It is not a new chip. Three doubled-clock WSE-3 Turbo wafers in a redesigned rack deliver 4,400 tokens per second per user — about 12.6x the fastest GPU service, not the advertised 30x.</p>
<p>That one paragraph is the whole review in miniature, and the rest of this article exists because almost every number in it needs unpacking. Cerebras announced the CS-4 on 2026-08-18 at its Supernova event, shipped the first units in Q3 2026, and priced neither the box nor the rack power. The vendor headline says &ldquo;up to 30x faster than GPU-based solutions.&rdquo; The reproducible third-party measurement says roughly 12.6x. Both numbers are defensible depending on which workload you run and which GPU configuration you compare against — and neither is the number a buyer should use without doing arithmetic first.</p>
<h2 id="cerebras-cs-4-what-actually-shipped-in-q3-2026">Cerebras CS-4: What Actually Shipped in Q3 2026?</h2>
<p>The CS-4 is a rack-scale system built from <strong>three WSE-3 Turbo wafers</strong>, the first multi-wafer product Cerebras has shipped. Per the official launch blog and the press release of 2026-08-18, a full rack delivers:</p>
<ul>
<li><strong>750 PFLOPS</strong> sparse FP16 AI compute</li>
<li><strong>129.6 PB/s</strong> aggregate on-wafer memory bandwidth</li>
<li><strong>160.5 PB/s</strong> total compute fabric bandwidth</li>
<li><strong>7.2 Tbit/s</strong> system I/O</li>
<li>Wafer-to-wafer latency as low as <strong>2 microseconds</strong> via switchless Direct Wafer Links</li>
</ul>
<p>The system sits in a new modular <strong>Nexus rack</strong> with 50 percent fewer components, self-contained compute, power and I/O assemblies, and direct liquid cooling. Cerebras says deployment time drops from days to hours, and that the same Nexus platform is committed to carry the CS-5 and CS-6 later this decade. First CS-4 shipments began in Q3 2026.</p>
<p>What did not ship is equally important: there is no published purchase price, no published rack power rating, no cooling-water specification, no named CS-4 customer, and no reproducible MLPerf Inference result. Those are four unknowns a procurement team would normally expect in a launch quarter, and their absence is the first thing to notice about this release.</p>
<h2 id="is-the-cerebras-cs-4-a-new-chip-no--it-is-the-same-wafer-at-twice-the-clock">Is the Cerebras CS-4 a New Chip? No — It Is the Same Wafer at Twice the Clock</h2>
<p>This is the single most important fact in the launch, and it is the one the naming obscures. <strong>WSE-3 Turbo is the same die as the 2024 WSE-3.</strong> Same 4 trillion transistors. Same 900,000 AI cores. Same 44GB of on-wafer SRAM. Same TSMC 5nm process. Same 46,225 mm² of silicon. Cerebras did not tape out a new chip.</p>
<p>What it did was double the clock roughly from <strong>1.4GHz to 2.8GHz</strong>, and that required re-engineering the power and cooling path rather than the silicon. Two mechanisms make it work, both stated in the launch material:</p>
<ol>
<li><strong>Power conversion moved about 100x closer to the wafer</strong> — from roughly 50mm away to roughly 0.5mm. Shorter delivery distance means less resistive loss and less voltage droop, which is what makes the higher clock sustainable.</li>
<li><strong>Direct liquid cooling folded into a rear-mounted Wafer-Scale Backpack</strong>, which packages power conversion, cooling, high-speed I/O and control electronics into one module.</li>
</ol>
<p>Every performance gain in the spec sheet traces back to that clock ratio. Double the clock and per-wafer compute, memory bandwidth, fabric bandwidth, I/O and network all double — which is exactly what the numbers show. ServeTheHome&rsquo;s verdict is blunt and correct: WSE-3 Turbo is &ldquo;not an all-new design; virtually every critical element runs at twice the speed.&rdquo;</p>
<h3 id="cs-4-vs-cs-3-the-spec-table-with-the-multipliers-separated">CS-4 vs CS-3: The Spec Table, With the Multipliers Separated</h3>
<p>Cerebras&rsquo; own comparison table pits a <strong>one-wafer CS-3</strong> against a <strong>three-wafer CS-4</strong>. That means two independent multipliers are baked into every headline ratio: roughly <strong>2x per wafer</strong> from the clock bump, and <strong>3x from the wafer count</strong>. Read the table with both in mind.</p>
<table>
  <thead>
      <tr>
          <th>Specification</th>
          <th>CS-3 (1x WSE-3)</th>
          <th>CS-4 (3x WSE-3 Turbo)</th>
          <th>Ratio</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Sparse FP16 AI compute</td>
          <td>125 PFLOPS</td>
          <td>750 PFLOPS</td>
          <td>6x</td>
      </tr>
      <tr>
          <td>On-wafer memory bandwidth</td>
          <td>21.6 PB/s</td>
          <td>129.6 PB/s</td>
          <td>6x</td>
      </tr>
      <tr>
          <td>Compute fabric bandwidth</td>
          <td>26.8 PB/s</td>
          <td>160.5 PB/s</td>
          <td>6x</td>
      </tr>
      <tr>
          <td>System I/O</td>
          <td>1.2 Tbit/s</td>
          <td>7.2 Tbit/s</td>
          <td>6x</td>
      </tr>
      <tr>
          <td>Wafer-to-wafer latency</td>
          <td>5 µs</td>
          <td>2 µs</td>
          <td>2.5x better</td>
      </tr>
      <tr>
          <td>SRAM per wafer</td>
          <td>44GB</td>
          <td>44GB</td>
          <td>unchanged</td>
      </tr>
      <tr>
          <td>AI cores per wafer</td>
          <td>900,000</td>
          <td>900,000</td>
          <td>unchanged</td>
      </tr>
      <tr>
          <td>Transistors per wafer</td>
          <td>4 trillion</td>
          <td>4 trillion</td>
          <td>unchanged</td>
      </tr>
      <tr>
          <td>Process node</td>
          <td>TSMC 5nm</td>
          <td>TSMC 5nm</td>
          <td>unchanged</td>
      </tr>
  </tbody>
</table>
<p><em>(Sources: Cerebras press release 2026-08-18; HPCwire reproduction; ServeTheHome spec table; Devlery comparison table.)</em></p>
<p>Notice that 6x is not 30x. The remaining factor comes from the benchmark, not the hardware.</p>
<h2 id="where-does-the-30x-faster-claim-come-from--and-does-it-hold-up">Where Does the &ldquo;30x Faster&rdquo; Claim Come From — and Does It Hold Up?</h2>
<p>The claim is that the CS-4 delivers &ldquo;up to 30x faster tokens per second per user&rdquo; than GPU-based systems, illustrated by <strong>more than 4,400 output tokens per second per user on gpt-oss-120B</strong> with identical prompts. Cerebras prints its own caveat with the chart: the figure is sourced to Artificial Analysis and internal benchmarking, August 2026.</p>
<p>Three problems compress that 30x when you look for the underlying measurement:</p>
<ul>
<li><strong>The GPU is unnamed.</strong> No vendor, model, GPU count, serving runtime, precision, prompt length or concurrency level appears in the release. MLQ.ai&rsquo;s analysis makes this its central criticism: the release &ldquo;does not identify the competing GPU configuration and publishes no complete test protocol.&rdquo;</li>
<li><strong>The reproducible third-party number is about 12.6x.</strong> Artificial Analysis measured 4,400 tok/s per user on gpt-oss-120B against roughly <strong>350 tok/s</strong> on the fastest GPU-based inference service. That is a real, large lead — and it is 12.6x, not 30x.</li>
<li><strong>SemiAnalysis lands in between.</strong> Its independent August 2026 estimate puts real-world interactivity gains at <strong>20-40x</strong> for most frontier deployments, arguing that GPU inference rarely runs at theoretical peak in practice.</li>
</ul>
<p>The honest reading is that the 30x is a ceiling measured against an unstated and probably unoptimized GPU baseline, while 12.6x is a floor measured against the best GPU service available on the same open weights. Both are more than good enough for the CS-4 to be the fastest option in its class, which is why the inflated headline mostly hurts Cerebras&rsquo; credibility rather than its product.</p>
<table>
  <thead>
      <tr>
          <th>Claim</th>
          <th>Value</th>
          <th>Who measured it</th>
          <th>What it compares against</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Vendor headline</td>
          <td>up to 30x tok/s per user</td>
          <td>Cerebras, internal + Artificial Analysis data</td>
          <td>An unnamed GPU configuration</td>
      </tr>
      <tr>
          <td>Reproducible third-party</td>
          <td>~12.6x (4,400 vs ~350 tok/s)</td>
          <td>Artificial Analysis</td>
          <td>The fastest GPU inference service</td>
      </tr>
      <tr>
          <td>Analyst estimate</td>
          <td>20-40x interactivity</td>
          <td>SemiAnalysis</td>
          <td>Typical frontier deployments</td>
      </tr>
  </tbody>
</table>
<h2 id="is-750-pflops-the-number-an-llm-buyer-should-use-no--use-the-dense-figure">Is 750 PFLOPS the Number an LLM Buyer Should Use? No — Use the Dense Figure</h2>
<p>The 750 PFLOPS headline is a <strong>sparse FP16</strong> figure. Sparsity means the hardware can skip zero-valued weights when a model has them; most deployed LLM weights are dense, so structured sparsity gives you little or nothing on real inference traffic. The Register&rsquo;s dense analysis, relayed independently by TechTimes and The Clarity, puts WSE-3 Turbo at roughly <strong>25 PFLOPS per chip</strong> in dense FP16 — a tenfold gap from the 250 PFLOPS per-wafer headline.</p>
<p>That correction does not sink the product, because compute is not where the CS-4&rsquo;s advantage lives. It does mean that if you are sizing a cluster on the 250 PFLOPS number, you are sizing it on the wrong quantity. Use the dense figure for capacity planning and the memory bandwidth figure for decode planning.</p>
<h2 id="why-does-wafer-scale-sram-beat-hbm-for-decode">Why Does Wafer-Scale SRAM Beat HBM for Decode?</h2>
<p>Autoregressive token generation is <strong>memory-bandwidth-bound, not FLOP-bound.</strong> Every generated token requires streaming the active weights, so tokens per second scale with how fast you can move bytes out of memory, not with how many multiply-accumulates you can issue. That is why the CS-4&rsquo;s defining number is the one nobody advertises on a slide: <strong>43.2 PB/s of on-wafer memory bandwidth per wafer.</strong></p>
<p>On a GPU, weights live in HBM across a package boundary and reach the compute die through an interposer or an off-package hop. On a wafer-scale engine, the SRAM <strong>is</strong> the compute silicon — the same die, no serializer, no interposer. Cerebras claims roughly 2,000x the memory bandwidth of Nvidia&rsquo;s next-generation GPU on that specific metric, and while the multiple is vendor-framed, the architectural direction is not controversial.</p>
<p>The trade is capacity, not bandwidth. <strong>44GB of SRAM cannot hold a frontier model in FP16</strong>, and that ceiling has not moved since WSE-2 in 2021 — the same 44GB, three generations running. Everything about the three-wafer rack exists because of that ceiling: models get sharded across wafers, and wafer-to-wafer links at 2 microseconds have to beat a memory round-trip for the sharding to be worth it. Analysts expect the unannounced WSE-4 to double SRAM to roughly 88GB; that, not another clock bump, is what the next generation would have to deliver.</p>
<h2 id="tokens-per-second-per-user-is-not-throughput">Tokens Per Second Per User Is Not Throughput</h2>
<p>This is the distinction that decides whether the price premium pays for itself, and Cerebras does not put it on the chart. <strong>tok/s per user measures latency for one stream.</strong> It says a single person waits far less. It does not say that a rack serves more customers per hour — for that you need aggregate throughput under batch load, and Cerebras does not publish it. GPU vendors usually report aggregate batch throughput, so the two headline numbers are not directly comparable at all.</p>
<p>The practical consequence: for <strong>batch, offline or overnight work</strong>, tokens per dollar is the only metric that matters, and the CS-4 is not built to win that metric. Cerebras hardware exists to minimize the wall-clock time a human or an agent spends waiting, and it is priced accordingly.</p>
<h2 id="what-is-the-agentic-case-for-the-cerebras-cs-4">What Is the Agentic Case for the Cerebras CS-4?</h2>
<p>If there is one workload where the premium is unambiguously justified, it is the agent loop, because latency <strong>compounds across chained model calls</strong> rather than averaging out.</p>
<p>Do the arithmetic. A ReAct-style agent making <strong>12 tool calls</strong> at 300ms of model time per call burns <strong>3.6 seconds of pure waiting</strong> before any tool executes. At roughly 2,000 tok/s decode, the same loop finishes before a GPU-backed agent completes its second hop. Across a long autonomous session the difference is not &ldquo;faster responses&rdquo; — it is a different number of reasoning, verification and tool-use steps inside the same wall clock.</p>
<p>CTO Sean Lie&rsquo;s framing of the 30x claim is the launch&rsquo;s most defensible statement: it gives an agent an order of magnitude more reasoning, verification or tool use in the same amount of time. That is a real capability change, and it is why the fastest inference hardware matters even when its tokens cost more.</p>
<h2 id="how-does-disaggregated-inference-work-with-amd-helios-and-aws-trainium">How Does Disaggregated Inference Work With AMD Helios and AWS Trainium?</h2>
<p>The CS-4 does not do prefill at scale, and Cerebras is explicit about the architecture: a purpose-built <strong>prefill engine processes the prompt, transfers state, and the CS-4 performs ultra-low-latency decode.</strong> The named partners are AMD Helios and AWS Trainium.</p>
<ul>
<li><strong>AMD Helios + Cerebras</strong>, announced 2026-07-23, targets production in Q4 2026 and is reported at <strong>10x faster than GPUs alone and 5x more throughput than Cerebras decode-only.</strong></li>
<li><strong>AWS Trainium-backed Bedrock</strong> integration is targeted for Q1 2027.</li>
</ul>
<p>This is a coherent design — prefill is compute-bound and throughput-friendly, decode is bandwidth-bound and latency-sensitive, so you buy the right silicon for each phase. It is also the source of the CS-4&rsquo;s biggest operational risk.</p>
<h2 id="the-long-context-caveat-what-does-disaggregation-lock-in">The Long-Context Caveat: What Does Disaggregation Lock In?</h2>
<p>Two things.</p>
<p>First, <strong>the prefill-to-decode ratio is fixed at purchase.</strong> A split cluster is sized for one ratio of prompt processing to token generation. If your traffic mix shifts — longer documents, more retrieval, bigger agent contexts — you cannot rebalance in software. A homogeneous GPU cluster can be dynamically reallocated as workloads change; a disaggregated Cerebras cluster cannot, without new capex.</p>
<p>Second, <strong>long-context agentic turns push latency back onto the prefill side</strong>, which is exactly where the CS-4&rsquo;s advantage does not apply. Engineers report the possibility of <strong>90-second delays on very long contexts</strong> while the AMD or AWS side chews through the prompt. A 30x decode speedup does not help you if the prefill partner is the bottleneck on every turn.</p>
<h2 id="how-much-power-does-a-cerebras-cs-4-rack-draw">How Much Power Does a Cerebras CS-4 Rack Draw?</h2>
<p>Futurum Group&rsquo;s analyst note reports roughly <strong>120-140 kW per CS-4 rack</strong> — with three wafers aboard — against the estimated <strong>240-250 kW</strong> for AMD Helios and Nvidia Vera Rubin class racks. ServeTheHome independently infers about <strong>54 kW per WSE-3 Turbo wafer</strong> from Cerebras&rsquo; statement that the CS-4 delivers twice the power to the wafer, up from roughly 27 kW for WSE-3.</p>
<p>Both are estimates. Cerebras has not published a rack power rating, a cooling-loop flow rate, or a facility-water requirement, and voltage/frequency curves at a doubled clock are exactly where power estimates have historically drifted. Treat the ~54 kW per wafer as a planning figure, not a specification.</p>
<h2 id="what-does-the-cerebras-cs-4-cost-in-practice">What Does the Cerebras CS-4 Cost in Practice?</h2>
<p>Nobody knows what the CS-4 costs to buy, because the price is undisclosed. What is public is what Cerebras charges for inference on the same open weights its competitors serve — and the spread is wide.</p>
<table>
  <thead>
      <tr>
          <th>Provider (gpt-oss-120B)</th>
          <th>Input $/M tokens</th>
          <th>Output $/M tokens</th>
          <th>Measured speed</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Cerebras Cloud</td>
          <td>$0.35</td>
          <td>$0.75</td>
          <td>~1,641 tok/s</td>
      </tr>
      <tr>
          <td>Groq</td>
          <td>~$0.15</td>
          <td>~$0.60</td>
          <td>~500 tok/s</td>
      </tr>
      <tr>
          <td>DeepInfra / OpenRouter (commodity GPU)</td>
          <td>—</td>
          <td>~$0.20 combined</td>
          <td>Batch-tier throughput</td>
      </tr>
  </tbody>
</table>
<p><em>(Sources: Cerebras pricing page; PricePerToken snapshot 2026-08-21 via WealthEngine.)</em></p>
<p>Two things follow. The same open weights trade across a <strong>roughly 5x spread</strong> on price, so &ldquo;fast&rdquo; and &ldquo;cheap&rdquo; are genuinely different purchases. And on a tokens-per-dollar basis, third-party comparisons put the fastest tier at roughly <strong>10x the cost</strong> of commodity GPU tiers, while delivering about 4.6x the speed over Azure-class GPU endpoints. The premium pays when a human or an agent is waiting; it does not pay for batch or overnight work.</p>
<h2 id="the-business-behind-the-box-why-cloud-now-matters-more-than-hardware">The Business Behind the Box: Why Cloud Now Matters More Than Hardware</h2>
<p>Cerebras went public on Nasdaq under <strong>CBRS</strong> in May 2026 — a $5.55B listing at $185 per share, the largest US tech IPO since Snowflake — and the CS-4 is its first hardware shipped since the debut.</p>
<p>The Q2 FY2026 numbers, reported 2026-08-12, show where the business is actually going:</p>
<ul>
<li><strong>Core revenue $209.9M</strong> (+103% YoY)</li>
<li><strong>Core cloud and services revenue $127.7M</strong> (+287% YoY)</li>
<li><strong>Core hardware revenue $82.1M</strong> (+17% YoY)</li>
<li>Cloud moved from <strong>32 percent of GAAP revenue a year ago to 70 percent this quarter</strong></li>
</ul>
<p>Hardware is no longer the growth engine; it is the capacity input for the cloud business. Add the $25.4B backlog, a promised 10x manufacturing capacity expansion for 2026, 600+ MW of data center capacity live or contracted by end-2027, and the OpenAI deal — <strong>750 MW of Cerebras systems in a multi-year 2026-2028 agreement valued at more than $20B</strong>, the largest high-speed inference deployment announced to date.</p>
<p>The practical implication for a reader: you will almost certainly meet the CS-4 through an API — OpenAI&rsquo;s GPT-5.6 Sol Ultrafast routing (which required zero code changes for existing callers), Cerebras Cloud, or OpenRouter — rather than through a purchase order. A vendor willing to commit 750 MW to a single customer is a vendor optimizing for cloud margins, not for hardware list prices.</p>
<h2 id="how-does-the-cs-4-compare-with-nvidia-amd-groq-and-sambanova">How Does the CS-4 Compare With Nvidia, AMD, Groq and SambaNova?</h2>
<table>
  <thead>
      <tr>
          <th>System</th>
          <th>Architecture</th>
          <th>Where it wins</th>
          <th>Where the CS-4 wins</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Cerebras CS-4</td>
          <td>3x wafer-scale, on-die SRAM, no HBM</td>
          <td>Latency-bound decode, agent loops</td>
          <td>—</td>
      </tr>
      <tr>
          <td>Nvidia Vera Rubin / GB300</td>
          <td>HBM-based GPU + NVLink scale-up</td>
          <td>Throughput, model capacity, ecosystem</td>
          <td>Single-user decode latency</td>
      </tr>
      <tr>
          <td>AMD Helios</td>
          <td>HBM GPU rack, disaggregated prefill partner</td>
          <td>Batch throughput, dynamic rebalancing</td>
          <td>Paired (not head-to-head) on decode</td>
      </tr>
      <tr>
          <td>Groq (Nvidia-owned since 2026)</td>
          <td>Deterministic LPU, SRAM-based</td>
          <td>Price aggression, low-latency serving</td>
          <td>Peak single-stream speed</td>
      </tr>
      <tr>
          <td>AWS Trainium</td>
          <td>Custom accelerator, Bedrock-integrated</td>
          <td>Cost, managed integration</td>
          <td>Paired (not head-to-head) on decode</td>
      </tr>
  </tbody>
</table>
<p>Two structural notes. Nvidia&rsquo;s <strong>$20B acquisition of Groq assets</strong> — its largest deal on record — removed the most aggressive independent price competitor from the fast-inference market, leaving Cerebras the last large pure-play fast-inference silicon vendor. And because AMD and AWS now supply prefill for Cerebras rather than compete with it on decode, the competitive map has partly stopped being zero-sum.</p>
<h2 id="who-should-adopt-the-cerebras-cs-4-now--and-who-should-wait">Who Should Adopt the Cerebras CS-4 Now — and Who Should Wait?</h2>
<p><strong>Test it now if</strong> you run latency-bound, decode-heavy workloads: real-time agents, multi-tool ReAct loops, interactive coding or research assistants, voice or realtime interfaces where a human is staring at the screen. If your users measure success in perceived responsiveness, the 12.6x floor is worth real money.</p>
<p><strong>Wait if</strong> your workload is batch, offline, or throughput-dominated, where you optimize tokens per dollar and GPU tiers beat Cerebras by roughly 5x on identical open weights. Also wait if your traffic is dominated by very long contexts — you would be buying a decode advantage and paying a prefill penalty — or if you need dynamic rebalancing across a shifting prompt-to-generation mix, or if you need published power, cooling and price figures before you can budget the deployment.</p>
<h2 id="how-do-you-test-the-30x-claim-for-5">How Do You Test the 30x Claim for $5?</h2>
<p>Do not trust the headline and do not trust this review. The vendor&rsquo;s own $5 trial credit is enough to measure the claim on your own traffic:</p>
<ol>
<li><strong>Measure time-to-first-token (TTFT)</strong> on your real prompt distribution at your real context length.</li>
<li><strong>Measure sustained inter-token latency</strong> at the output length you actually produce, not at 200 tokens.</li>
<li><strong>Measure end-to-end agent-loop wall clock</strong> — the metric that decides whether an agent gets more work done per minute.</li>
<li><strong>Divide by cost per million tokens</strong> at your input/output mix.</li>
<li><strong>Compare against one commodity-GPU route of the same open weights</strong>, on the same prompts, in the same hour.</li>
</ol>
<p>If the resulting speedup on your workload is 12x, that is the number to plan with — and it is still, for interactive workloads, the largest available in 2026.</p>
<h2 id="verdict-the-fastest-inference-hardware--for-decode-not-for-everything">Verdict: The Fastest Inference Hardware — for Decode, Not for Everything</h2>
<p>The Cerebras CS-4 is the fastest inference hardware shipping today, on the workloads where speed is defined as single-user decode latency. It achieves that with an unchanged die clocked twice as high inside a genuinely new rack, at roughly half the rack power of the GPU systems it outruns. The three multipliers behind its 30x headline — a 2x clock bump, three wafers instead of one, and a self-run benchmark against an unspecified GPU — are all separately verifiable, and separating them gives a range of 12.6x to 30x depending on the baseline you accept.</p>
<p>What it is not is cheaper, more flexible, or more reproducible on paper. The 30x headline rests on a benchmark with no published protocol, dense FP16 compute is roughly a tenth of the sparse figure, disaggregation locks your prefill-to-decode ratio at purchase, and the price of the box is still undisclosed. Buy it because a person or an agent is waiting. Do not buy it because a slide said 30x.</p>
<h2 id="faq">FAQ</h2>
<h3 id="is-the-cerebras-cs-4-actually-faster-than-gpus">Is the Cerebras CS-4 actually faster than GPUs?</h3>
<p>Yes, on decode-dominated workloads. Artificial Analysis measured 4,400 tokens per second per user on gpt-oss-120B against about 350 tok/s on the fastest GPU-based inference service, which is roughly 12.6x. Cerebras&rsquo; own headline of &ldquo;up to 30x&rdquo; comes from an internal benchmark on the same model against an unnamed GPU configuration with no published test protocol. SemiAnalysis independently estimates 20-40x interactivity gains for most frontier deployments. The direction of the claim is not in dispute; the peak number is.</p>
<h3 id="does-the-cerebras-cs-4-use-a-new-chip">Does the Cerebras CS-4 use a new chip?</h3>
<p>No. The die is the same WSE-3 introduced in 2024 — same 4 trillion transistors, 900,000 AI cores, 44GB of SRAM, TSMC 5nm process and 46,225 mm² of silicon. The &ldquo;Turbo&rdquo; designation comes from roughly doubling the clock from about 1.4GHz to about 2.8GHz, which was only possible because Cerebras moved power conversion from about 50mm to about 0.5mm from the wafer and folded direct liquid cooling into a new Wafer-Scale Backpack. The CS-4 is best understood as a rack-generation product wearing a silicon-generation name.</p>
<h3 id="how-much-does-a-cerebras-cs-4-cost">How much does a Cerebras CS-4 cost?</h3>
<p>Cerebras has not disclosed the purchase price, the lease rate, or a cloud rate for dedicated CS-4 capacity, and no named customer has been announced. What is public is Cerebras Cloud pricing on gpt-oss-120B at $0.35 per million input tokens and $0.75 per million output tokens at about 1,641 tokens per second, versus roughly $0.20 per million combined for the same open weights on commodity GPU routes — a spread of about 5x on identical model weights.</p>
<h3 id="what-is-disaggregated-inference-and-why-does-the-cs-4-need-it">What is disaggregated inference and why does the CS-4 need it?</h3>
<p>The CS-4 performs decode only. A separate prefill engine processes the prompt, transfers state, and the CS-4 generates tokens at ultra-low latency — AMD Helios (production targeted Q4 2026) and AWS Trainium-powered Bedrock (targeted Q1 2027) are the named partners. This is efficient because prefill is compute-bound while decode is memory-bandwidth-bound, but it fixes your prefill-to-decode ratio at purchase and pushes the latency of very long prompts onto the partner hardware, where engineers report possible 90-second delays.</p>
<h3 id="is-wafer-scale-sram-better-than-hbm-for-llm-inference">Is wafer-scale SRAM better than HBM for LLM inference?</h3>
<p>For decode, yes — bandwidth matters more than capacity. Autoregressive generation is memory-bandwidth-bound, and the CS-4 keeps 43.2 PB/s of bandwidth per wafer on the same silicon as the compute, with no interposer or off-package hop to HBM. The catch is capacity: 44GB of SRAM cannot hold a frontier model in FP16, and that figure has not moved since WSE-2 in 2021, which is why models must be sharded across three wafers and why analysts expect the next generation to double SRAM to roughly 88GB before anything else changes.</p>
]]></content:encoded></item></channel></rss>