<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>GLM-5.3 CyberGym on RockB</title><link>https://baeseokjae.github.io/tags/glm-5.3-cybergym/</link><description>Recent content in GLM-5.3 CyberGym on RockB</description><image><title>RockB</title><url>https://baeseokjae.github.io/images/og-default.png</url><link>https://baeseokjae.github.io/images/og-default.png</link></image><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 01 Oct 2026 05:03:36 +0000</lastBuildDate><atom:link href="https://baeseokjae.github.io/tags/glm-5.3-cybergym/index.xml" rel="self" type="application/rss+xml"/><item><title>GLM-5.3 Benchmarks: Artificial Analysis Scores, Full Comparison Table, and Pricing</title><link>https://baeseokjae.github.io/posts/glm-5-3-artificial-analysis-benchmarks/</link><pubDate>Thu, 01 Oct 2026 05:03:36 +0000</pubDate><guid>https://baeseokjae.github.io/posts/glm-5-3-artificial-analysis-benchmarks/</guid><description>GLM-5.3 scores 45 on Artificial Analysis Intelligence Index v4.3.2, down from 60 in August. That drop is a re-base, not a regression.</description><content:encoded><![CDATA[<p>GLM-5.3 benchmarks today: 45 on Artificial Analysis Intelligence Index v4.3.2, 28.3 on Terminal-Bench 3.0, 66.9 on DeepSWE v1.1, and 1769 Elo on GDPval-AA v2 — the highest of any open-weight model. The famous &ldquo;60&rdquo; from August is the same model on an older index version.</p>
<p>That is the entire controversy compressed into two sentences. Everything below is the evidence behind it: the vendor table Z.ai published, the independent Artificial Analysis run, the price you actually pay per unit of work, and the capability gaps that did not close.</p>
<h2 id="what-glm-53-is-and-isnt-post-training-on-the-glm-52-base">What GLM-5.3 Is (and Isn&rsquo;t)? Post-Training on the GLM-5.2 Base</h2>
<p>GLM-5.3 is a 753B-parameter mixture-of-experts model with 40B active parameters, a 1M-token context window, and up to 128K tokens of output. It ships as FP8 E4M3 weights, is served by 27 API providers, and is capped by a 1,048,576-token context ceiling that Z.ai markets as &ldquo;full-repo&rdquo; scale.</p>
<p>The most important structural fact is not in the launch marketing: <strong>GLM-5.3 uses the same base model as GLM-5.2</strong>. Z.ai&rsquo;s own model card states it plainly. Every benchmark gain you see in the head-to-head tables below came from post-training — reinforcement learning, data curation, and harness work — not from a new pretraining run. That matters when you read the size of the deltas, because it tells you Z.ai&rsquo;s post-training pipeline is doing the heavy lifting, and it means the family will not plateau until that pipeline does.</p>
<h3 id="does-glm-53-support-image-input">Does GLM-5.3 support image input?</h3>
<p>No. GLM-5.3 is text-in, text-out only. There is no vision path at all, which is the single clearest gap between it and GPT-5.6 Sol, Fable 5, or Grok 4.6. If your pipeline needs screenshots, diagrams, or PDF-as-image ingestion, you need a different model in the loop — GLM-5.3-Flash (separate weights, image input supported) or a closed frontier model.</p>
<h3 id="can-you-disable-thinking-on-glm-53">Can you disable thinking on GLM-5.3?</h3>
<p>No, and this is a breaking change. Requests that disable thinking fail on GLM-5.3. It is a reasoning model with a <code>reasoning_effort</code> control that accepts low, high, or max, and the default is <code>max</code>. Migrating from GLM-5.2 with <code>thinking: disabled</code> in your request body will produce failures, not degraded answers. Budget your token spend accordingly.</p>
<h2 id="glm-53-on-artificial-analysis-45-today-60-in-august--why-the-number-moved">GLM-5.3 on Artificial Analysis: 45 Today, 60 in August — Why the Number Moved</h2>
<p>This is the question that generates more confusion than any other about GLM-5.3 benchmarks. Both numbers are correct.</p>
<table>
  <thead>
      <tr>
          <th>Date</th>
          <th>Index version</th>
          <th>GLM-5.3 score</th>
          <th>What changed in the harness</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Aug 18, 2026</td>
          <td>Intelligence Index v4.1.1</td>
          <td>60</td>
          <td>Original independent run by Artificial Analysis</td>
      </tr>
      <tr>
          <td>Sep 7, 2026</td>
          <td>Intelligence Index v4.3</td>
          <td>45</td>
          <td>GPQA Diamond removed; Terminal-Bench 4.0 and AutomationBench-AA added</td>
      </tr>
      <tr>
          <td>Oct 2026 (current)</td>
          <td>Intelligence Index v4.3.2</td>
          <td>45</td>
          <td>Ranked #2 of 117 in class; class median 18</td>
      </tr>
  </tbody>
</table>
<p>The August score of 60 put GLM-5.3 level with Kimi K3 (60), three points behind Claude Opus 5 (63), 8th of 181 models in its class against a class median of 35. It cost $0.68 per index task versus $0.84 for Kimi K3 and $2.34 for Claude Opus 5, and the full index run cost $1,238.50 on Z.ai&rsquo;s API.</p>
<p>The September re-base swapped a graduate-level science evaluation out and two agentic evaluations in. GLM-5.3&rsquo;s relative position barely moved — it still sits near the top of its class — but the absolute number fell by 15 points because the battery now weights the agentic work where GLM-5.3 is merely good rather than excellent.</p>
<p><strong>The practical rule: leaderboard snapshots are not comparable across index versions.</strong> A model that &ldquo;dropped 15 points&rdquo; overnight without a weight change did not get worse. A vendor that quotes the highest number it ever received is not lying, but it is not being useful either.</p>
<h2 id="the-full-benchmark-table-glm-53-vs-glm-52-kimi-k3-deepseek-v4-pro-and-gpt-56-sol">The Full Benchmark Table: GLM-5.3 vs GLM-5.2, Kimi K3, DeepSeek-V4 Pro and GPT-5.6 Sol</h2>
<p>The table below is Z.ai&rsquo;s launch-day vendor-run table from August 14, 2026. Treat it as the vendor&rsquo;s claim, corroborated where Artificial Analysis ran the same evaluation independently.</p>
<table>
  <thead>
      <tr>
          <th>Benchmark</th>
          <th>GLM-5.3</th>
          <th>GLM-5.2</th>
          <th>GPT-5.6 Sol</th>
          <th>Fable 5</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Terminal-Bench 3.0</td>
          <td>28.3</td>
          <td>4.6</td>
          <td>34.6</td>
          <td>33.7</td>
      </tr>
      <tr>
          <td>DeepSWE v1.1</td>
          <td>66.9</td>
          <td>46.2</td>
          <td>72.7</td>
          <td>—</td>
      </tr>
      <tr>
          <td>SWE-Marathon</td>
          <td>42.5</td>
          <td>19.4</td>
          <td>—</td>
          <td>—</td>
      </tr>
      <tr>
          <td>Agents&rsquo; Last Exam</td>
          <td>28.5</td>
          <td>23.8</td>
          <td>—</td>
          <td>—</td>
      </tr>
      <tr>
          <td>AutomationBench v1.0.6</td>
          <td>48.2</td>
          <td>26.2</td>
          <td>—</td>
          <td>—</td>
      </tr>
      <tr>
          <td>CyberGym</td>
          <td>84.5</td>
          <td>77.2</td>
          <td>—</td>
          <td>—</td>
      </tr>
      <tr>
          <td>ExploitBench</td>
          <td>54.4</td>
          <td>24.4</td>
          <td>76.5</td>
          <td>78.0</td>
      </tr>
      <tr>
          <td>GDPval-AA v2 (Elo)</td>
          <td>1769</td>
          <td>1508</td>
          <td>1730</td>
          <td>—</td>
      </tr>
  </tbody>
</table>
<p>GDPval-AA v2 was run by Artificial Analysis rather than Z.ai, which makes it one of the more trustworthy rows: GLM-5.3&rsquo;s 1769 Elo beats GPT-5.6 Sol (1730), Kimi K3 (1682), DeepSeek-V4 Pro (1590), and Claude Opus 4.8 (1588). On real-world professional deliverables, the open-weight model is at the front of the field, not chasing it.</p>
<p>The rows where it loses are equally specific. GLM-5.3 trails GPT-5.6 Sol on Terminal-Bench 3.0 (28.3 vs 34.6) and DeepSWE v1.1 (66.9 vs 72.7), and it trails both Sol (76.5) and Fable 5 (78.0) badly on ExploitBench (54.4). The pattern is consistent: GLM-5.3 is competitive on general agentic work and clearly behind on adversarial and security-flavored coding.</p>
<h2 id="coding-and-agentic-benchmarks-where-the-post-training-gains-actually-landed">Coding and Agentic Benchmarks: Where the Post-Training Gains Actually Landed</h2>
<p>The headline number — Terminal-Bench 3.0 going from 4.6 to 28.3 — deserves an honest read. A 6x improvement looks extraordinary, and it is real, but it is measured off a near-floor baseline. GLM-5.2 was barely functional on that harness, so the percentage overstates the absolute capability you are purchasing. The 28.3 figure itself is the number to plan against.</p>
<p>Where the gains are genuinely large against a meaningful baseline:</p>
<ul>
<li><strong>DeepSWE v1.1: 46.2 to 66.9</strong> — a 20-point move on a software-engineering benchmark where GLM-5.2 was already respectable.</li>
<li><strong>SWE-Marathon: 19.4 to 42.5</strong> — more than double, on long-horizon software tasks.</li>
<li><strong>ExploitBench: 24.4 to 54.4</strong> — more than double, and the reason the safety discussion below exists.</li>
<li><strong>CyberGym: 77.2 to 84.5</strong> — a solid gain in an area that was already a strength.</li>
</ul>
<p>Independent harnesses tell a similar but not identical story. In akitaonrails&rsquo; OpenCode evaluation, GLM-5.3 scored 94 — two points behind Claude Fable 5 (96) — in 80 minutes at roughly $2.59 in API-equivalent cost, the highest score ever recorded on that harness. On Reinvently&rsquo;s Ed-o-meter, which runs 28 everyday tasks, GLM-5.3 became the first model to clear all five categories at a 100% pass rate with a 9.3 rubric score at $0.28 per lap, roughly one-fifth the cost of GPT-5.5. The cost of that quality is latency: a median 16.3 seconds to first token.</p>
<p>Independent numbers also expose the split the composite hides. On Artificial Analysis&rsquo;s own detailed evals, GLM-5.3 posts Terminal-Bench 4.0 at 42% versus 33% for Flash, and SciCode at 59% versus 52% — but it <em>trails</em> Flash on GDP.pdf (11% vs 15%) and ties Flash at 80% on AA-LCR v1.1. The 45-versus-42 composite makes the flagship look unambiguously better. It is not, on every task.</p>
<h2 id="cost-speed-and-token-efficiency-vs-the-closed-frontier">Cost, Speed and Token Efficiency vs the Closed Frontier</h2>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>GLM-5.3</th>
          <th>Notes</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Input price (per 1M tokens)</td>
          <td>$1.40</td>
          <td>Artificial Analysis listing</td>
      </tr>
      <tr>
          <td>Output price (per 1M tokens)</td>
          <td>$4.40</td>
          <td>81% cache discount available</td>
      </tr>
      <tr>
          <td>Blended (7:2:1)</td>
          <td>$0.90</td>
          <td>Weighted mix</td>
      </tr>
      <tr>
          <td>Avg cost per Intelligence Index task</td>
          <td>$2.01</td>
          <td>Independent run</td>
      </tr>
      <tr>
          <td>Output speed</td>
          <td>67.8 tok/s</td>
          <td>44.7 tok/s for GLM-5.3-Flash</td>
      </tr>
      <tr>
          <td>Time to first token</td>
          <td>3.32 s</td>
          <td>Slowest tier of frontier models</td>
      </tr>
      <tr>
          <td>Output tokens across index run</td>
          <td>210M</td>
          <td>vs 140M median — very verbose</td>
      </tr>
  </tbody>
</table>
<p>The cost-per-unit-of-work picture is more favorable than the raw prices suggest. Artificial Analysis measured $0.68 per index task for GLM-5.3 against $0.84 for Kimi K3 and $2.34 for Claude Opus 5 — a 3.4x advantage over Anthropic&rsquo;s model at comparable index score. An independent September 25 snapshot recorded 63 tok/s for GLM-5.3 against 43 tok/s for Flash.</p>
<p>The token-efficiency angle is where GLM-5.3 is most underrated. On Z.ai Code Bench at High reasoning effort, GLM-5.3 scores 31.4% using roughly 50K output tokens, while Claude Opus 4.8 scores 29.5% using roughly 120K tokens. GLM-5.3 reaches a higher score with less than half the spend. At Max effort it hits 34.5% at ~75K tokens (GLM-5.2: 23.4% at 96K). Fable 5 still leads at Max with 39.5%, but it is not twice as good for the price.</p>
<p>The counterweight is verbosity. 210M output tokens across the index run versus a 140M median means GLM-5.3 thinks at length, and at $4.40 per 1M output tokens that verbosity is billed.</p>
<h2 id="the-cyber-capability-story-and-the-delayed-open-weights-release">The Cyber Capability Story and the Delayed Open-Weights Release</h2>
<p>GLM-5.3 was released on August 14, 2026, but the weights were withheld at launch and only published on August 28, 2026 under the GLM-5.3 Licence — roughly a two-week safety hold. This is unusual for a Z.ai release and it was driven by the model&rsquo;s security profile.</p>
<p>Anthropic&rsquo;s research reports that GLM-5.3&rsquo;s safeguards can be bypassed 64% to 100% of the time using simple techniques in simulated tests. NIST&rsquo;s CAISI called it &ldquo;the most cyber-capable open-weight model released to date&rdquo; and placed it roughly four months behind the US frontier. Those two statements, combined with ExploitBench nearly doubling from 24.4 to 54.4, explain the hold.</p>
<p>Practical consequence for deployers: the license permits commercial use but carries restrictions, and the safety posture means you should not expose GLM-5.3 to untrusted code execution, security tooling, or autonomous network access without your own guardrails. The model is capable; the model&rsquo;s built-in refusals are not the thing stopping a determined attacker.</p>
<p>Hacker News response tracked the shift in perception: the launch post &ldquo;GLM-5.3: Frontier coding with emergent cyber capabilities&rdquo; reached 1,171 points and 584 comments, and &ldquo;GLM-5.3 is now open-weight&rdquo; reached 806 points two weeks later. On Hugging Face, zai-org/GLM-5.3 recorded 1,372,510 downloads and 2,024 likes as of October 1, 2026.</p>
<h2 id="glm-53-vs-glm-53-flash-which-one-should-you-actually-deploy">GLM-5.3 vs GLM-5.3-Flash: Which One Should You Actually Deploy?</h2>
<p>The 753B/40B flagship is not a realistic self-hosting target for most teams. GLM-5.3-Flash is a different model with different weights, and for many workloads it is the better buy.</p>
<table>
  <thead>
      <tr>
          <th></th>
          <th>GLM-5.3 (flagship)</th>
          <th>GLM-5.3-Flash</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Artificial Analysis Intelligence Index</td>
          <td>45</td>
          <td>42</td>
      </tr>
      <tr>
          <td>Active/total parameters</td>
          <td>40B / 753B</td>
          <td>18B / 320B</td>
      </tr>
      <tr>
          <td>License</td>
          <td>GLM-5.3 Licence (restrictions)</td>
          <td>MIT</td>
      </tr>
      <tr>
          <td>Context</td>
          <td>1M</td>
          <td>1M</td>
      </tr>
      <tr>
          <td>Image input</td>
          <td>No</td>
          <td>Yes</td>
      </tr>
      <tr>
          <td>Price (in/out per 1M)</td>
          <td>$1.40 / $4.40</td>
          <td>$0.15 / $0.50</td>
      </tr>
      <tr>
          <td>Output speed</td>
          <td>67.8 tok/s</td>
          <td>44.7 tok/s</td>
      </tr>
      <tr>
          <td>Cost per index task</td>
          <td>$2.01</td>
          <td>$0.25</td>
      </tr>
      <tr>
          <td>Terminal-Bench 4.0</td>
          <td>42%</td>
          <td>33%</td>
      </tr>
      <tr>
          <td>SciCode</td>
          <td>59%</td>
          <td>52%</td>
      </tr>
      <tr>
          <td>GDP.pdf</td>
          <td>11%</td>
          <td>15%</td>
      </tr>
      <tr>
          <td>AA-LCR v1.1</td>
          <td>80%</td>
          <td>80%</td>
      </tr>
  </tbody>
</table>
<p>Pick the flagship when you need the top of the agentic curve and the extra 3 index points are worth an 8x price multiple — long-horizon coding, complex tool orchestration, GDPval-style deliverable work. Pick Flash when you need multimodality, an MIT license, or throughput at a tenth of the cost, and accept that you give up ground on the hardest reasoning tasks.</p>
<h2 id="vendor-numbers-vs-independent-evaluations-how-to-read-benchmark-claims">Vendor Numbers vs Independent Evaluations: How to Read Benchmark Claims</h2>
<p>Sort every GLM-5.3 benchmark number into one of three tiers before you trust it.</p>
<ol>
<li><strong>Independently run, versioned, reproducible.</strong> Artificial Analysis Intelligence Index (45 on v4.3.2), GDPval-AA v2 (1769 Elo), the akitaonrails OpenCode harness (94), Reinvently&rsquo;s Ed-o-meter (9.3 rubric). These are run by third parties with published harnesses.</li>
<li><strong>Vendor-run on a third-party harness.</strong> Z.ai&rsquo;s August 14 table — Terminal-Bench 3.0, DeepSWE, SWE-Marathon, CyberGym, ExploitBench. The harness is real; the run is the vendor&rsquo;s, and vendors choose prompts and attempts.</li>
<li><strong>Directional only.</strong> Anything labeled &ldquo;with tools,&rdquo; anything with sampling-parameter footnotes, anything whose harness version changed underneath it. HLE with tools is explicitly flagged this way in Z.ai&rsquo;s own footnotes.</li>
</ol>
<p>A useful sanity check: when a vendor number and an independent number for the same evaluation disagree, the independent one is usually lower and usually closer to what you will experience.</p>
<h2 id="verdict--who-should-use-glm-53-in-2026">Verdict — Who Should Use GLM-5.3 in 2026</h2>
<p>GLM-5.3 is the strongest open-weight model available on general agentic and professional-deliverable work, and it is not the strongest overall model. It wins on cost-adjusted performance — 1769 GDPval Elo and $0.68 per index task against $2.34 for a comparable closed model — and it loses on the hardest adversarial coding tasks, on multimodality, and on latency.</p>
<p>Deploy it if you run high-volume agentic or coding workloads where a 3x cost advantage compounds, you can absorb 67.8 tok/s and 3.32s to first token, and you are willing to build your own safety envelope. Do not deploy it if you need vision input, sub-second first-token latency, best-in-class security tooling, or turnkey self-hosting of the flagship model at 753B parameters.</p>
<p>Read any single benchmark number as directional. Read the version number next to it as mandatory.</p>
<h2 id="faq">FAQ</h2>
<p><strong>Is GLM-5.3 better than GLM-5.2?</strong>
Yes, on every benchmark Z.ai and Artificial Analysis measured. The largest absolute gains are on long-horizon agentic tasks — SWE-Marathon (19.4 to 42.5) and ExploitBench (24.4 to 54.4) — because GLM-5.3 kept the GLM-5.2 base model and improved only through post-training.</p>
<p><strong>Why did GLM-5.3 drop from 60 to 45 on Artificial Analysis?</strong>
The index changed, not the model. Artificial Analysis re-based from Intelligence Index v4.1.1 to v4.3 on September 7, 2026, removing GPQA Diamond and adding Terminal-Bench 4.0 and AutomationBench-AA. The new battery weights agentic work more heavily, which pulled the score down 15 points while leaving GLM-5.3 near the top of its class.</p>
<p><strong>What does GLM-5.3 cost?</strong>
$1.40 per 1M input tokens and $4.40 per 1M output tokens on Artificial Analysis&rsquo;s listing, with an 81% cache discount and a $0.90 blended rate at a 7:2:1 mix. Average cost per Intelligence Index task measured $2.01. GLM-5.3-Flash is dramatically cheaper at $0.15/$0.50.</p>
<p><strong>Can I self-host GLM-5.3?</strong>
Technically yes — the weights are open (published August 28, 2026 under the GLM-5.3 Licence, FP8 E4M3) — but the flagship is 753B total / 40B active parameters. Most teams cannot serve it economically. GLM-5.3-Flash at 320B/18B under an MIT license is the realistic self-hosting target.</p>
<p><strong>Does GLM-5.3 beat GPT-5.6 Sol?</strong>
Head to head, no — it trades. GLM-5.3 wins GDPval-AA v2 (1769 vs 1730 Elo) and loses Terminal-Bench 3.0 (28.3 vs 34.6), DeepSWE v1.1 (66.9 vs 72.7), and ExploitBench (54.4 vs 76.5). It wins decisively on cost, at roughly a third of the closed model&rsquo;s price per index task.</p>
]]></content:encoded></item></channel></rss>