Look at the spec table for two minutes and the story falls out on its own. Every number on the WSE-3 Turbo is exactly double the WSE-3: 125 petaflops becomes 250, 21.6 PB/s of memory bandwidth becomes 43.2, I/O goes 1.2 to 2.4 terabits. Transistor count, core count, SRAM, die area, process node: all identical. So the CS-4 that Cerebras announced on August 18 is a genuinely new rack, three wafers deep, hitting 750 sparse FP16 petaflops and 129.6 PB/s of memory bandwidth. But the silicon inside it is last year's die with the clocks pushed. That's not a knock, exactly. Doubling a 46,225 mm2 wafer's clock without melting it is hard engineering. It just means the interesting part of this launch is the rack, not the chip, and the 30x headline needs reading with care.
The short answer
Cerebras announced the CS-4 on August 18, a rack holding three WSE-3 Turbo wafers. The rack is new: modular backpacks, redesigned power delivery, direct liquid cooling. The die is not. Same 4 trillion transistors, same 900,000 cores, same 44 GB of SRAM, roughly twice the clock. The 30x-faster-than-GPUs line measures tokens per second per user on one model against unnamed GPU hardware. First shipments this quarter.
What actually shipped
A CS-4 is one rack. Inside it sit three Wafer Scale Engine 3 Turbo processors, each one a full 300 mm wafer of TSMC 5 nm silicon carved into a single chip. Cerebras rates the system at 750 PFLOPS of AI compute, 129.6 PB/s of aggregate memory bandwidth, 160.5 PB/s of on-chip fabric bandwidth and 7.2 Tbit/s of external I/O, with wafer-to-wafer latency as low as 2 microseconds over what it calls Direct Wafer Links. No switch in the path. Against a single CS-3 at 125 PFLOPS, that’s six times the listed compute in one box.
The per-wafer table is where it gets interesting. WSE-3 Turbo: 4 trillion transistors, 46,225 mm2, 900,000 AI cores, 44 GB of on-wafer SRAM, 250 PFLOPS, 43.2 PB/s of memory bandwidth, 2.4 Tbit/s of I/O. Now put that beside the WSE-3 from the CS-3 generation. Transistors, area, cores, SRAM, process node: identical, to the digit. Compute, bandwidth, fabric, I/O: exactly 2x, also to the digit.

Image: Cerebras
That pattern only comes from one thing. You don’t get a clean 2x on four unrelated metrics from an architectural change; you get it from turning the clock up. The Register put the die at roughly 1.4 GHz to 2.8 GHz, and Cerebras hasn’t published a frequency to argue with. Honestly, we’d rather they just said so. Doubling clocks across a wafer that size, with the power delivery and the thermals that implies, is a real achievement. Dressing it as a new processor generation invites exactly the scrutiny it’s now getting.
The 30x number, unpacked
Cerebras led with “up to 30 times faster than GPU-based solutions.” Here’s what sits underneath it. In the side-by-side clip published with the launch, GPT-OSS-120B generates 4,465 tokens per second per user on a CS-4, 2,308 on a CS-3, and 131 on a GPU deployment that goes unidentified. That’s a 34x ratio, rounded down to a marketing 30x.
Three things to hold onto. Tokens per second per user is a latency metric, so it rewards an architecture that keeps an entire model resident in SRAM and never touches HBM, which is the whole Cerebras thesis. It says nothing about tokens per dollar or tokens per rack when you’re serving thousands of concurrent sessions. Second, Cerebras’ own asterisk marks its petaflop figures as sparse FP16; comparing that against dense GPU numbers stretches the gap before anyone runs a model. And the GPU system stays anonymous, which for a claim this loud is a choice.
None of that makes the demo fake. We’ve seen the GPT-5.6 Sol Ultrafast tier run at speeds nothing else touches, and OpenAI says that mode runs up to 14x its standard offering for select customers on Cerebras hardware. The speed is real. The 30x framing is the vendor’s best case dressed as a general result.

Image: Cerebras. Watch the full side-by-side demo on Cerebras.
The rack is the real product
Strip the chip story away and the Nexus platform is what Cerebras actually built this year. The rack splits into three independent elements, compute, power and I/O, so each can be replaced on its own schedule. Compute arrives as a Wafer-Scale Backpack: a vertically mounted module that carries the processor along with its power conversion, its direct liquid cooling loop, its high-speed I/O and its control electronics, and plugs into the rack as one unit.
Cerebras says the design uses 50 percent fewer components than the previous generation and is 60 percent more automated to manufacture, and that power conversion now sits about 100 times closer to the processor than on a conventional GPU board. That last one is the physics that makes doubled clocks survivable. Shorter distance, less resistive loss, less voltage droop under a transient.
Efficiency is where the claim gets more defensible: up to 10x the throughput per watt of a CS-3. Even discounted heavily, a 2x clock that doesn’t cost 4x the power is the part of this launch we’d take seriously. Cerebras hasn’t published a rack power figure. The design implies roughly double the per-wafer draw, and third-party estimates land a full rack near 120 to 140 kilowatts, which would be around half a comparable GPU rack. Estimates, though. Nobody outside Cerebras has metered one.
Prefill on someone else’s silicon
The genuinely new architectural idea is disaggregated inference, and it’s an admission worth noticing. Prefill, the phase where a model chews through your input context, is compute-bound and suits GPUs fine. Decode, the token-by-token generation, is memory-bandwidth-bound and is where wafer-scale wins. So the CS-4 supports splitting the two across different hardware over standards-based RoCE v2 RDMA on plain Ethernet, with AMD’s Helios racks and AWS Trainium named as prefill partners. Cerebras claims the pairing hits 10x GPU speed with 5x the throughput of Cerebras hardware alone.
Read that as Cerebras conceding it doesn’t want the whole workload. It wants the decode half, sitting next to a GPU fleet you already own. Which is a smarter go-to-market than asking anyone to rip out their accelerators, and it’s why the RoCE v2 detail matters more than the petaflops.
Should you care yet
If you’re buying inference capacity, probably not this quarter. First shipments are promised before the quarter closes, but Cerebras disclosed no CS-4 customer agreements and no pricing, and a rack-scale system with no public price is not something you can budget against. Models over 50 trillion parameters are supported on paper, with over 1,000 tokens per second quoted on models past 10 trillion. Paper, still.
If you’re building agents, the argument is more concrete. Sean Lie, Cerebras’ CTO, framed it as reasoning headroom rather than snappiness, and that’s the right frame: at 4,000-plus tokens per second per user, a chain that would take a minute elsewhere finishes fast enough to run several times over. Whether you can rent that at a sane price is the open question, and it stays open until someone publishes a number.
The 2027 test is the one that counts. Clocks only go up once. Next generation has to be new silicon.
Sources
Cerebras, “Introducing Cerebras CS-4” (official announcement, 18 August 2026, including the spec table and the side-by-side demo video). Cerebras CS-4 product page. ServeTheHome on the WSE-3 Turbo and the CS-4 rack. The Next Web on what the launch leaves out, which carries the clock-speed reporting attributed to The Register and the rack power estimate. The Next Platform on the Nexus architecture.
Frequently asked questions
What is the Cerebras CS-4?
It is a rack-scale AI inference system Cerebras announced on 18 August 2026, the first product built on its Nexus platform architecture. One CS-4 holds three WSE-3 Turbo wafer-scale processors and is rated at 750 sparse FP16 petaflops, 129.6 PB/s of aggregate memory bandwidth, 160.5 PB/s of on-chip fabric bandwidth, 7.2 Tbit/s of external I/O and wafer-to-wafer latency as low as 2 microseconds. Cerebras says first shipments begin this quarter.
Is the WSE-3 Turbo a new chip?
Not a new design. It keeps the WSE-3's 4 trillion transistors, 900,000 AI cores, 44 GB of on-wafer SRAM and 46,225 mm2 of TSMC 5 nm silicon. What changed is clock speed: compute, memory bandwidth, fabric bandwidth and I/O each land at exactly twice the WSE-3 figure, which is the signature of a frequency bump rather than a redesign. The Register reported the die moving from roughly 1.4 GHz to 2.8 GHz. Cerebras has not published a clock speed itself.
What does the 30x faster than GPUs claim actually measure?
Tokens per second per user, on one model, against GPU systems Cerebras does not name. In the side-by-side clip published with the launch, GPT-OSS-120B runs at 4,465 tokens/sec/user on a CS-4, 2,308 on a CS-3 and 131 on the GPU box. That is a single-stream latency metric, not aggregate throughput, and Cerebras' own footnote marks its petaflop figures as sparse FP16. Treat 30x as the vendor's best case, not a general speedup.
How much does a CS-4 cost and who is buying one?
Cerebras published neither. No price, no list price per rack, no named CS-4 customer agreement in the launch materials. It also has not released an official rack power figure, though the design implies roughly double the per-wafer draw of the WSE-3 and third-party estimates put a full rack somewhere around 120 to 140 kilowatts. All of that is unconfirmed.
What is disaggregated inference on the CS-4?
Splitting the two phases of inference across different hardware. Prefill, where the model reads your input context, runs on GPUs or other accelerators; decode, the token-by-token generation, runs on the Cerebras wafers. The CS-4 supports this natively over standards-based RoCE v2 RDMA on Ethernet, with AMD Helios and AWS Trainium named as prefill partners. Cerebras claims the pairing reaches 10x GPU speed and 5x the throughput of Cerebras alone, which is a vendor number nobody has reproduced yet.