<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Mac Studio on RockB</title><link>https://baeseokjae.github.io/tags/mac-studio/</link><description>Recent content in Mac Studio on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Tue, 29 Sep 2026 01:18:26 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/mac-studio/index.xml" rel="self" type="application/rss+xml"/><item><title>Local LLM Rig Cost Payback: How Long Until It Pays for Itself?</title><link>https://baeseokjae.github.io/posts/local-llm-rig-cost-payback-2026/</link><pubDate>Tue, 29 Sep 2026 01:18:26 +0000</pubDate><guid>https://baeseokjae.github.io/posts/local-llm-rig-cost-payback-2026/</guid><description>Local LLM rig cost payback runs from ~4.8 months at 20h/day to never at 12h/day. Here is the five-number worksheet and the utilisation cliff behind it.</description><content:encoded><![CDATA[<p>A local LLM rig cost payback runs from about 4.8 months to infinity, and the machine is not what decides it. The same $3,500 RTX 5090 box pays back in ~4.8 months at 20 h/day against a $2/M API tier, stretches to ~39 months at 24/7 against a $0.50/M hosted open-weight model, and never pays back at all at 12 h/day — because a $133/month API bill undercuts a $137/month local rig. Payback is a property of your queue depth, not your GPU.</p>
<p>That is the whole answer in one paragraph. The rest of this post is the arithmetic, so you can plug your own four numbers in and get a defensible answer in about ten minutes.</p>
<h2 id="how-long-until-a-local-llm-rig-pays-for-itself-the-short-answer">How long until a local LLM rig pays for itself? (the short answer)</h2>
<p>Payback is amortised hardware plus electricity plus cooling plus ops time, divided by the gap between that and the API bill you avoided. If you produce under roughly 50M tokens a month on a cheap hosted model, the answer is almost always &ldquo;never.&rdquo; Between 50M and 500M tokens a month it depends on whether the data can leave your network and whether you needed the box for something else. Above 500M tokens a month sustained, local wins on cash.</p>
<p>Most people get a wrong answer because they compare a GPU running at 100% utilisation against an API invoice from a month when they barely used it — the local cost is charged on capability to produce, the API cost on tokens actually produced. Fix that framing and the numbers stop lying to you. And when the numbers still say &ldquo;never,&rdquo; you probably have to admit you are not asking about money at all.</p>
<h2 id="the-five-numbers-you-need-before-any-comparison">The five numbers you need before any comparison</h2>
<p>Skip the calculator shopping. You need five inputs, and only two of them require research.</p>
<table>
  <thead>
      <tr>
          <th>#</th>
          <th>Input</th>
          <th>Where it comes from</th>
          <th>Typical trap</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>1</td>
          <td>Throughput (tok/s)</td>
          <td>Benchmark for your exact quant</td>
          <td>Vendor prefill numbers quoted as decode</td>
      </tr>
      <tr>
          <td>2</td>
          <td>Duty cycle (hours/day)</td>
          <td>Honest log of your usage</td>
          <td>Assuming &ldquo;always on&rdquo; = &ldquo;always generating&rdquo;</td>
      </tr>
      <tr>
          <td>3</td>
          <td>Electricity rate (c/kWh)</td>
          <td>Your own utility bill</td>
          <td>Copying $0.12 from a 2023 blog post</td>
      </tr>
      <tr>
          <td>4</td>
          <td>Amortisation window (years)</td>
          <td>What you actually believe</td>
          <td>Using 3 years and then running the box for 6</td>
      </tr>
      <tr>
          <td>5</td>
          <td>The API tier you are truly replacing</td>
          <td>A like-for-like model, not the frontier</td>
          <td>Comparing a 12B local model to Claude Opus</td>
      </tr>
  </tbody>
</table>
<p>That last one deserves emphasis, because it is where most break-even math quietly becomes fiction. If your local model scores 42 on the Artificial Analysis Intelligence Index and the hosted model you compare it against scores 53, you have not replaced that model. You have replaced a cheaper one.</p>
<h2 id="step-1--turn-throughput-into-tokens-per-month">Step 1 — Turn throughput into tokens per month</h2>
<p>The formula is unglamorous: <code>tok/s x 3600 x hours_per_day x 30 x utilisation = tokens/month</code>.</p>
<p>Take a GeForce RTX 5090, which hits 205.5 decode tok/s versus 60.9 tok/s for an NVIDIA DGX Spark on the same benchmark — a 3.4x gap that most &ldquo;the hardware is a rounding error&rdquo; arguments ignore completely. Prefill tells the opposite story less dramatically: 8,518.6 tok/s for the 5090 against 2,054.0 tok/s for the Spark.</p>
<p>At 205 tok/s, 20 hours a day, your ceiling is about 443M tokens a month. Write that on paper now, because everything else divides into it.</p>
<p>Then subtract the fraction of that ceiling you will never reach. Real token demand is spiky, and a 20 h/day average with a 205 tok/s peak means your actual sustained utilisation is probably 25-40% of the peak. Use 30% if you are unsure and you will land closer to reality than any optimistic calculator will.</p>
<p><strong>The memory check you cannot skip.</strong> Before throughput, confirm the model fits — weights plus KV cache, not weights alone. Llama 3.3 70B Instruct at Q4_K_M needs 42.5GB of weights plus 10.7GB of KV cache at 32k context: 53.3GB total. The cache runs 327,680 bytes per token across 80 layers, so the context window you ask for is a hardware purchase decision, not a settings toggle. A 48GB card that &ldquo;runs 70B&rdquo; runs it at the context length you can afford, which on a 48GB card is short.</p>
<h2 id="step-2--price-your-electricity-honestly">Step 2 — Price your electricity honestly</h2>
<p>Most local-vs-cloud write-ups still assume $0.12-$0.15/kWh. The US average residential price was 18.34 c/kWh as of June 2026, up from 15.04 c/kWh in 2022 — a 22% rise in four years — with the EIA forecasting 3-5%/year through 2027. State extremes run from 13.11c in Nevada to 52.72c in Hawaii on the same June 2026 data.</p>
<p>So the common assumption understates power cost by roughly 20-50%. On a 600W rig running 20 h/day, moving from $0.15 to $0.183/kWh adds about $18/month. That is not fatal on its own. It becomes fatal when it stacks on an already marginal payback, or when you look at the Spark-vs-Mac gap below.</p>
<p><strong>The idle-power asymmetry.</strong> A DGX Spark draws roughly 120W idle, versus about 25W for a Mac Studio. A box you leave plugged in 24/7 but use four hours a day spends 20 hours a day at idle draw. That is a quiet, permanent tax that a lower upfront price can hide.</p>
<table>
  <thead>
      <tr>
          <th>Hardware</th>
          <th>Idle draw</th>
          <th>Load draw</th>
          <th>Notes</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>NVIDIA DGX Spark</td>
          <td>~120W</td>
          <td>600-800W</td>
          <td>Idle draw is 5x a Mac&rsquo;s</td>
      </tr>
      <tr>
          <td>Mac Studio (M-series Max)</td>
          <td>~25W</td>
          <td>150-200W</td>
          <td>Lower peak, far lower idle</td>
      </tr>
      <tr>
          <td>RTX 5090 desktop build</td>
          <td>60-100W</td>
          <td>450-600W</td>
          <td>Sleep states help if you use them</td>
      </tr>
      <tr>
          <td>Dual-H100 server</td>
          <td>300W+</td>
          <td>1,500W+</td>
          <td>Idle cost is a line item, not noise</td>
      </tr>
  </tbody>
</table>
<p>And the cooling bill: power and cooling add 15-25% on top of hardware TCO at US rates, more in Europe. If your spreadsheet has one &ldquo;electricity&rdquo; cell and nothing else, add 20% to it and move on.</p>
<h2 id="step-3--pick-an-amortisation-window-you-actually-believe">Step 3 — Pick an amortisation window you actually believe</h2>
<p>Three-year amortisation is the industry default and it is optimistic in a specific way: it assumes your hardware is worthless at month 37, and it assumes you will not run it for six years.</p>
<p>Here is a representative 2026 set, all on 3-year amortisation with 24/7 power at $0.15/kWh:</p>
<table>
  <thead>
      <tr>
          <th>Rig</th>
          <th>Upfront</th>
          <th>3-yr TCO</th>
          <th>Cost per year</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>RTX 5090 desktop build</td>
          <td>$3,500</td>
          <td>$5,120</td>
          <td>~$1,707</td>
      </tr>
      <tr>
          <td>Mac Studio M5 Max 128GB</td>
          <td>$3,499</td>
          <td>$4,429</td>
          <td>~$1,476</td>
      </tr>
      <tr>
          <td>NVIDIA DGX Spark</td>
          <td>$3,000</td>
          <td>$6,150</td>
          <td>~$2,050</td>
      </tr>
      <tr>
          <td>Dual-H100 server</td>
          <td>$60,000</td>
          <td>$78,900</td>
          <td>~$26,300</td>
      </tr>
  </tbody>
</table>
<p>Note the ordering flip: the Spark is the cheapest to buy and the second-most expensive to run, because 120W of idle draw plus 600-800W under inference eats its $500 purchase advantage inside the first year. This is why &ldquo;which box is cheapest&rdquo; and &ldquo;which box has the best payback&rdquo; are different questions.</p>
<p>The counter-argument is residual value, and it is real. Community reports from the Sunk Cost discussion thread describe used RTX 3090s bought for ~$500 after the merge, and commenters reporting their hardware value doubling or tripling since purchase. If your GPU appreciates, amortisation is a fiction in your favour. Do not plan around it — but do not plan around 100% depreciation either.</p>
<h2 id="step-4--pick-the-right-api-comparator-where-most-payback-math-goes-wrong">Step 4 — Pick the right API comparator (where most payback math goes wrong)</h2>
<p>This is the single largest source of wrong answers, so let me be blunt: comparing a local mid-tier model against frontier hosted pricing is a category error, not an aggressive assumption.</p>
<p>Live OpenRouter pricing on 2026-09-29 shows the spread clearly:</p>
<table>
  <thead>
      <tr>
          <th>Tier</th>
          <th>Example</th>
          <th>Output price</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Cheapest open-weight hosts</td>
          <td>mistral-nemo</td>
          <td>$0.030/M</td>
      </tr>
      <tr>
          <td></td>
          <td>llama-3.1-8b</td>
          <td>$0.080/M</td>
      </tr>
      <tr>
          <td></td>
          <td>gpt-oss-20b</td>
          <td>$0.090/M</td>
      </tr>
      <tr>
          <td></td>
          <td>gemma-3-4b</td>
          <td>$0.100/M</td>
      </tr>
      <tr>
          <td>Mid-tier hosted</td>
          <td>local-model-class equivalents</td>
          <td>~$0.50-1.25/M</td>
      </tr>
      <tr>
          <td>Frontier hosts</td>
          <td>claude-opus-5</td>
          <td>$25/M</td>
      </tr>
      <tr>
          <td></td>
          <td>claude-fable-5.1</td>
          <td>$50/M</td>
      </tr>
      <tr>
          <td></td>
          <td>gpt-6-astra</td>
          <td>$50/M</td>
      </tr>
      <tr>
          <td></td>
          <td>claude-opus-4.1</td>
          <td>$75/M</td>
      </tr>
  </tbody>
</table>
<p>The gap between $0.08/M and $75/M is nearly a thousandfold, and it is entirely a capability difference. Any payback number you compute against the frontier tier is a number about a model you did not deploy.</p>
<p>The honest comparator for a local rig in 2026 is the mid-tier hosted open-weight class at roughly $0.50-1.25/M output. The other defensible comparator is the specific invoice you are actually replacing — a GitHub Copilot seat, a ChatGPT Plus subscription, an existing API bill you can screenshot.</p>
<p>There is also a capability ceiling that hardware cannot buy past. The best open-weight model on Sunk Cost&rsquo;s leaderboard scores 42 on the Artificial Analysis Intelligence Index v4.3 — GLM-5.3-Flash, with 189GB of weights — against 53 for the best hosted model. That is an 11-point gap that no GPU closes. Any payback argument that assumes capability parity is quietly assuming you will be satisfied with less.</p>
<h2 id="step-5--add-the-costs-that-never-make-the-spreadsheet">Step 5 — Add the costs that never make the spreadsheet</h2>
<p>The items below are the difference between a calculator answer and a real one.</p>
<ul>
<li><strong>Ops labour.</strong> Model updates, quantization experiments, driver breakage, a broken CUDA matrix. Price your own hour at something nonzero and this is usually the largest hidden line.</li>
<li><strong>Idle power.</strong> Covered above, but it deserves a second mention because it is charged every hour of the year.</li>
<li><strong>Cooling.</strong> 15-25% on top of hardware TCO, and higher in warm climates or small rooms.</li>
<li><strong>Rate-limit engineering.</strong> If you still call an API for the hard 20% of requests, you maintain two stacks.</li>
<li><strong>Re-engineering cost when a vendor changes models.</strong> Prompt rewrites are a real cost on the API side — a genuine advantage for local, and most comparisons skip it because it does not have a number attached.</li>
<li><strong>Depreciation and residual value.</strong> Choose a number and write down why.</li>
</ul>
<h2 id="the-utilisation-cliff-same-rig-four-months-or-never">The utilisation cliff: same rig, four months or never</h2>
<p>Here is the table that answers the question in the title. Same $3,500 RTX 5090 rig, same 205 tok/s, electricity at the current US residential average of $0.183/kWh:</p>
<table>
  <thead>
      <tr>
          <th>Duty cycle</th>
          <th>Tokens/month (approx)</th>
          <th>Local cost/month</th>
          <th>API cost at $0.50/M</th>
          <th>API cost at $2.00/M</th>
          <th>Payback</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>4 h/day</td>
          <td>~89M</td>
          <td>~$75</td>
          <td>~$44</td>
          <td>~$177</td>
          <td>Only vs the expensive tier</td>
      </tr>
      <tr>
          <td>12 h/day</td>
          <td>~266M</td>
          <td>~$137</td>
          <td>~$133</td>
          <td>~$531</td>
          <td>Never vs $0.50/M</td>
      </tr>
      <tr>
          <td>20 h/day</td>
          <td>~443M</td>
          <td>~$163</td>
          <td>~$222</td>
          <td>~$886</td>
          <td>~4.8 months vs $2/M</td>
      </tr>
      <tr>
          <td>24/7</td>
          <td>~532M</td>
          <td>~$180</td>
          <td>~$266</td>
          <td>~$1,064</td>
          <td>~39 months vs $0.50/M</td>
      </tr>
  </tbody>
</table>
<p>Read the 12 h/day and 20 h/day rows together. Moving four hours a day flips the answer from &ldquo;never&rdquo; to &ldquo;under five months&rdquo; — not because the hardware changed, but because the monthly gap crossed zero. At 12 h/day the API bill is $133 and the local rig costs $137; the machine loses by $4/month and will lose forever.</p>
<p>Then read 20 h/day against 24/7. Going from 20 to 24 hours adds 89M tokens a month, which sounds like pure win, but it also pushes you down the API price ladder — because the cheaper hosted models are the ones you would actually use at that volume, and against $0.50/M the same rig needs ~39 months. The extra tokens are not extra savings if the comparator got cheaper.</p>
<p>That is the utilisation cliff: payback is not a property of the machine, it is a property of your queue depth.</p>
<h2 id="worked-example-a--3500-rtx-5090-rig-vs-a-hosted-open-weight-model">Worked example A — $3,500 RTX 5090 rig vs a hosted open-weight model</h2>
<p>Inputs: 205 tok/s decode, 20 h/day, $0.183/kWh, 450-600W under load, $3,500 upfront amortised over 3 years.</p>
<ul>
<li>Tokens/month: ~443M at the ceiling.</li>
<li>Electricity: roughly 500W average x 20 h x 30 days = 300 kWh/month = ~$55.</li>
<li>Hardware amortisation: ~$97/month.</li>
<li>All-in: ~$163/month, or about $0.37 per million tokens at full utilisation.</li>
</ul>
<p>Compare: a mid-tier hosted open-weight model at ~$0.50/M output is ~$222/month for the same 443M tokens. The rig is ahead by ~$59/month and pays back in about 4.8 months.</p>
<p>Now move the comparator. Against the cheap end of the market — $0.08/M output, which llama-3.1-8b and the $0.03-0.10 tier sit in — the same 443M tokens cost about $35/month. Payback becomes negative: it never happens.</p>
<p>Both of those numbers are correct. Which one applies to you depends on whether the model you would have called is worth calling.</p>
<h2 id="worked-example-b--used-rtx-3090-box-at-2000">Worked example B — used RTX 3090 box at $2,000</h2>
<p>This is the honest hobbyist configuration, and the community evidence is better than the spreadsheet evidence. One commenter runs Qwen 3.8 at 57 tok/s on a $2,000 used HP Omen with a 3090. Another replaced a $300+/month GitHub Copilot bill with a 32GB Radeon R9700 at $1,350 and reports a ~4-month payoff.</p>
<p>Run the numbers: 57 tok/s at 20 h/day is ~123M tokens/month at the ceiling. Electricity at ~$0.183/kWh and maybe 350W loaded lands near $38/month; amortised hardware over 3 years is ~$56/month. All-in ~$94/month, or roughly $0.76/M.</p>
<table>
  <thead>
      <tr>
          <th>Comparator</th>
          <th>Monthly API cost at 123M tokens</th>
          <th>Verdict</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>$0.50/M mid-tier</td>
          <td>~$62</td>
          <td>Local loses by ~$32/month</td>
      </tr>
      <tr>
          <td>$1.00/M</td>
          <td>~$123</td>
          <td>Local ahead ~$29/month, payback ~5.75 years</td>
      </tr>
      <tr>
          <td>A $300/month subscription it replaces</td>
          <td>$300</td>
          <td>Local ahead ~$206/month, payback under 6 months</td>
      </tr>
  </tbody>
</table>
<p>Notice what happened: the used-3090 rig&rsquo;s payback is essentially decided by what it replaces, not by its own efficiency. If you are replacing a flat-rate subscription with real usage, the numbers work fast. If you are replacing metered tokens at mid-tier prices, they are marginal.</p>
<p>The ~4-month community report is real and reproducible — but it is a subscription replacement, not a token-for-token comparison. Read claims like it with that in mind, including the optimistic framing.</p>
<h2 id="worked-example-c--mac-mini-m6-and-the-83-year-payback">Worked example C — Mac mini M6 and the 8.3-year payback</h2>
<p>Sunk Cost&rsquo;s own model, and the author&rsquo;s reply in the discussion thread, are the most deflating data points available. At 1M tokens/day, the quickest Sonnet-class local payback is Qwen3.8 27B on a Mac mini M6 with 32GB, at an estimated 6.9 tok/s — and the answer is <strong>8.3 years</strong>. Only at agent-scale, around 20M tokens/day, does it drop under a year. A user producing 50k tokens/day pays back in <strong>166 years</strong>.</p>
<p>The 6.9 tok/s figure is doing enormous work in that calculation, and it is explicitly labelled an estimate rather than a measurement. Two lessons:</p>
<ol>
<li><strong>Slow hardware makes payback unreachable, not long.</strong> Below roughly 100 tok/s you cannot produce enough tokens to matter, no matter how cheap the box was.</li>
<li><strong>Small-volume users should almost never buy.</strong> The floor on this calculation is not a few months; it is centuries. That is the correct order of magnitude for hobby token volumes, and it is why &ldquo;is a local rig cheaper than ChatGPT&rdquo; has a one-word answer at that volume: no.</li>
</ol>
<h2 id="does-local-ever-win-on-capability">Does local ever win on capability?</h2>
<p>No — and it is worth stating plainly rather than burying it in a table. The 11-point gap between the best open-weight model (42) and the best hosted model (53) on Artificial Analysis Intelligence Index v4.3 is a frontier gap, not a hardware gap. Buying a bigger GPU raises throughput; it does not raise that number.</p>
<p>What changes over time is the absolute capability of open weights, which has been improving steadily. The practical consequence: pick the local model you would genuinely be satisfied with, price the hosted service you would use instead at <em>that</em> capability level, and then compute payback. Every other version of the calculation is an argument for buying the machine you already wanted.</p>
<h2 id="when-local-wins-on-cash-alone-247-agent-and-batch-workloads">When local wins on cash alone: 24/7 agent and batch workloads</h2>
<p>There is exactly one regime where a local rig beats the API on cash for reasons that survive scrutiny: sustained, high-volume, latency-insensitive generation where the bottleneck is queue time rather than token price. Overnight batch jobs, embedding backfills, continuous subagent generation, eval sweeps.</p>
<p>Two caveats that the &ldquo;run it 24/7&rdquo; advice usually omits:</p>
<ul>
<li><strong>A flat-rate plan often serves more concurrency than one GPU.</strong> If your workload is API-shaped and bursty, a Max/Pro-tier subscription can absorb parallelism that a single 5090 cannot, because your constraint is VRAM and KV cache, not willingness to spend.</li>
<li><strong>Queue depth is the asset.</strong> Local wins when you have more work than you can get through, not when you have more money than you can spend. If your queue is already short, adding a GPU just makes idle hours more expensive.</li>
</ul>
<h2 id="decision-thresholds-hobby-prosumer-sustained">Decision thresholds: hobby, prosumer, sustained</h2>
<p>The brief asked for thresholds instead of vibes, so here they are. These assume a cheap-to-mid hosted comparator and a rig you would actually run hot.</p>
<table>
  <thead>
      <tr>
          <th>Your volume</th>
          <th>Recommended path</th>
          <th>Reasoning</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Under ~50M tokens/month</td>
          <td>Stay on the API</td>
          <td>Payback floor is years; cheap hosted tiers are $0.03-0.10/M</td>
      </tr>
      <tr>
          <td>50-500M tokens/month</td>
          <td>Buy only if you needed the hardware anyway, or the data cannot leave</td>
          <td>Marginal on cash; decisive on control and offline capability</td>
      </tr>
      <tr>
          <td>500M+/month sustained</td>
          <td>Local wins on cash</td>
          <td>The gap is large enough to survive electricity and ops surprises</td>
      </tr>
      <tr>
          <td>Privacy/regulated data at any volume</td>
          <td>Local, and stop pretending it is a cost decision</td>
          <td>Compliance cost is not on the API side of the ledger</td>
      </tr>
      <tr>
          <td>Replacing a $300/month subscription</td>
          <td>Usually pays back in months</td>
          <td>Compare subscriptions, not per-token rates</td>
      </tr>
  </tbody>
</table>
<p>For reference on where the market currently sits: one analysis found 7B-class models at 30% utilisation break even in 4-9 months and 70B-class on a DGX Spark in 3-6 months against frontier APIs, while stretching to 2-4 years under 10% utilisation. Another puts an RTX 4090 at ~$104/month all-in — beating GPT-4o at ~8M tokens/month but needing 204M+ tokens/month to beat GPT-4o-mini. Both land on the same rule: the cheaper your comparator, the harder local has to work.</p>
<h2 id="the-api-prices-keep-falling-trap">The API-prices-keep-falling trap</h2>
<p>Any payback calculation that assumes a flat API price predicts a payback that repricing can erase. This is the one place where the best available competitor tool is genuinely better than most write-ups: Sunk Cost lets you toggle &ldquo;assume API prices keep falling&rdquo; at a rate you set.</p>
<p>Model it explicitly. Take your payback in months and ask what the hosted price needs to do for the payback to exceed the amortisation window:</p>
<ul>
<li>If your comparator falls 20%/year and your payback is 24 months, you will be behind by month 18.</li>
<li>If your comparator falls 30%/year and your payback is 4.8 months, you are probably still fine — the fall does not have time to catch up.</li>
<li>If your comparator is $0.03-0.10/M today, there is almost nothing left to fall, which is good for local&rsquo;s relative position but bad for the case that local was ever cheaper at that tier.</li>
</ul>
<p>That last point is the cleanest way to state the 2026 situation: inference prices falling compresses the local advantage from above, but the cheap open-weight tier has already hit the floor. There is no version of this where a $3,500 rig beats a $0.03/M endpoint on cash.</p>
<h2 id="the-sunk-cost-test-five-questions-before-you-buy">The sunk cost test: five questions before you buy</h2>
<p>Run this before you order anything. If you cannot answer all five in a sentence each, the machine is a purchase and not an investment.</p>
<ol>
<li><strong>How many tokens per month will I actually produce, and how do I know?</strong> Cite your last 90 days of real usage, not a benchmark&rsquo;s ceiling.</li>
<li><strong>Which hosted service, at which price per million tokens, am I genuinely replacing?</strong> A like-for-like capability, not the frontier tier.</li>
<li><strong>What is my amortisation window, and what happens if the API price halves?</strong> If halving the API price kills the payback, say so out loud before you buy.</li>
<li><strong>What do the idle hours do?</strong> If the answer is &ldquo;nothing, but it&rsquo;s on anyway,&rdquo; the idle draw is a permanent subsidy to the argument.</li>
<li><strong>If this box could not be resold and produced no savings, would I still want it?</strong> An honest yes is a fine reason to buy — self-hosting is a hobby with real skills attached, and offline capability is genuinely worth something. The failure mode is not buying the box. It is buying it and calling a hobby a hedge.</li>
</ol>
<p>Once the box is on your desk, the trap closes in a specific, predictable way: the marginal token becomes nearly free, so the API comparison stops being a decision and becomes a justification. The sunk cost is real, but the comparison you start making afterwards is not. Make the decision while the money is still in your account.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Is running a local LLM cheaper than ChatGPT?</strong>
For almost everyone, no. A local rig&rsquo;s payback floor at hobby volumes is measured in years, and one model puts a 50k tokens/day user at a 166-year payback. The exceptions are people replacing a $300/month subscription with heavy real usage (community reports of ~4-month payoffs exist), and people running 500M+ tokens a month on 24/7 batch workloads.</p>
<p><strong>How is local LLM payback calculated?</strong>
<code>(hardware amortisation + electricity + cooling + ops) / (API cost avoided - local monthly cost)</code>. Take your sustained tok/s, multiply by hours per day, then by 30. Price electricity at your own rate, not a blog&rsquo;s $0.12. Amortise over a window you believe. Compare against a like-for-like hosted model, not the frontier tier.</p>
<p><strong>What is the difference between 12 h/day and 24/7 utilisation?</strong>
About 266M versus 532M tokens a month on a 205 tok/s rig — but the bigger effect is which hosted model becomes your comparator. At 12 h/day against $0.50/M output, a $3,500 rig never pays back ($133 API vs $137 local). At 20 h/day you are ~$59/month ahead against the same tier and pay back in ~4.8 months against $2/M. Small duty-cycle changes flip the sign.</p>
<p><strong>Does an RTX 5090 pay for itself against cloud APIs?</strong>
Against the frontier tier it appears to pay back in a month or two, which is a category error. Against a mid-tier hosted open-weight model at ~$0.50/M, a $3,500 rig at 20 h/day pays back in roughly 4.8 months. Against the cheapest hosted tier at $0.03-0.10/M, it never does.</p>
<p><strong>Is a Mac Studio or a DGX Spark better for break-even?</strong>
Neither, on cash terms. At 70B Q4 the Spark costs ~$0.60 per 1M output tokens in power against <del>$0.25 for the Mac Studio M5 Max at $0.15/kWh, but total amortised cost per 1M tokens lands within 3% of each other (</del>$1.70 vs ~$1.75), because Apple&rsquo;s higher hardware amortisation cancels its power advantage. The Spark&rsquo;s 120W idle draw is the detail that surprises people.</p>
<p><strong>What is the biggest hidden cost of a local LLM rig?</strong>
Ops labour, then idle power. Updates, driver breakage, quantization experiments, and a CUDA matrix that breaks on a Tuesday. Power and cooling add 15-25% on top of hardware TCO at US rates. Rate-limit engineering costs show up only if you still call an API for the hard cases — which most people do.</p>
<p><strong>Do local LLMs ever match frontier capability?</strong>
Not today. The best open-weight model scores 42 on Artificial Analysis Intelligence Index v4.3 against 53 for the best hosted model — an 11-point gap that no hardware closes. More VRAM raises throughput, not intelligence. Pick your comparator at the capability level you will actually accept, or the whole payback calculation is fiction.</p>
<h2 id="where-this-leaves-you">Where this leaves you</h2>
<p>The payback question has an honest answer and it is not a single number. At hobby volume it is &ldquo;never,&rdquo; and a spreadsheet will tell you so within ten minutes if you price electricity correctly and compare against the right model. At 50-500M tokens a month it is marginal, and control or compliance usually decides it. Above 500M tokens a month sustained, local wins on cash and the arithmetic gets comfortable.</p>
<p>If your numbers land in the &ldquo;never&rdquo; column and you still want the box — buy it. Just keep the two ledgers separate. The money ledger says no; the control ledger says you get a machine that runs when the network is down, that no rate limit can throttle, and that lets you own a skill that transfers. That is a legitimate purchase. What is not legitimate is re-running the calculator after the fact to prove the money ledger agreed.</p>
]]></content:encoded></item></channel></rss>