
Cerebras CS-4 Review: The Fastest Inference Hardware Yet?
On decode-heavy, latency-bound workloads, yes: the Cerebras CS-4 is the fastest inference hardware shipping in 2026. It is not a new chip. Three doubled-clock WSE-3 Turbo wafers in a redesigned rack deliver 4,400 tokens per second per user — about 12.6x the fastest GPU service, not the advertised 30x. That one paragraph is the whole review in miniature, and the rest of this article exists because almost every number in it needs unpacking. Cerebras announced the CS-4 on 2026-08-18 at its Supernova event, shipped the first units in Q3 2026, and priced neither the box nor the rack power. The vendor headline says “up to 30x faster than GPU-based solutions.” The reproducible third-party measurement says roughly 12.6x. Both numbers are defensible depending on which workload you run and which GPU configuration you compare against — and neither is the number a buyer should use without doing arithmetic first. ...