I was fortunate enough to attend Cerebras’ Supernova event this week, where the company launched its fourth-generation system, CS-4.

As someone who has grown up watching Apple live events, it was a surreal moment to be in the room as an advancement in arguably the most important race in the world unfolded.

The best moment was a live race. The same model was asked to generate an HTML periodic table on CS-4, CS-3 and a GPU system. CS-4 finished first. CS-3, already producing more than 2,300 tokens per second in the demonstration, followed shortly afterwards. The GPU was still generating when Sean Lie, CTO of Cerebras, moved on.

It was an effective demonstration of what Cerebras has become known for: tokens arriving at a speed that makes even its previous generation feel slow. As someone who’s grown up witnessing the speed of everything improve, this was very cool.

But the fastest tokens were not the most significant part of the launch. The rack behind them was.

CS-4 does not introduce a genuinely new wafer-scale processor. It uses a faster version of the same 5 nm WSE-3 found in CS-3. Cerebras calls it WSE-3 Turbo. By delivering considerably more power and cooling to the wafer, the company has approximately doubled its operating performance. Cerebras says this lifts on-wafer memory bandwidth from 21 PB/s to 43 PB/s and can deliver up to twice the token speed of CS-3.

The chip did not fundamentally change. The system around it did.

The rack has become the product.

CS-4 is the first system built on Cerebras’ new Nexus Platform Architecture. Instead of treating a rack as a collection of tightly coupled servers, cables, pumps, power supplies and networking equipment, Nexus separates it into modular power, compute and I/O assemblies.

The front of the rack contains a modular power array. At the rear sit three vertical, pluggable “backpacks,” each built around one WSE-3 Turbo. A backpack packages the wafer with power conversion, direct liquid cooling, high-speed I/O and control electronics.

This is a meaningful departure from CS-3. A full CS-4 system puts three wafers into one rack-scale product, while each wafer delivers roughly twice the performance of its predecessor. That is how Cerebras gets to its claim of six times the system-level performance: three wafers multiplied by approximately two times more performance per wafer.

More important than the arithmetic is how the system moves from factory to data centre. Cerebras says the backpack contains 50% fewer components than its prior-generation design and uses 60% more manufacturing automation. The power array can be installed in the data centre first, after which the compute backpacks are inserted on site. Cerebras claims this reduces deployment from days to hours.

They reveal what Nexus has been designed to optimise: not only tokens per second, but the path from a manufactured wafer to commissioned, revenue-generating capacity. Isn’t this the game we’re all playing?

For hyperscalers and neoclouds building inference capacity in large blocks, that distinction matters. A faster accelerator has limited commercial value while it is sitting in a factory, waiting for integration or being commissioned in a partially completed facility. Fewer parts, more automated assembly and a repeatable rack-level product can reduce several of those delays at once.

Cerebras is therefore selling something closer to a unit of deployable inference capacity than a standalone accelerator.

Double the performance, but not for free

CS-4’s performance increase comes with a substantial power requirement.

Cerebras has not yet published a CS-4 datasheet with an official rack TDP. SemiAnalysis estimates that a three-wafer CS-4 rack will draw approximately 125 to 135 kW. For comparison, Cerebras’ CS-3 datasheet specifies a maximum of 27 kW per system, or 54 kW for the two systems that fit in a rack.

That estimate implies roughly 42 to 45 kW per CS-4 wafer, compared with a maximum of 27 kW for CS-3. The WSE-3 Turbo is faster because Cerebras has engineered a system capable of pushing the existing silicon considerably harder.

This is impressive systems engineering. Delivering and removing that much power and heat across a single wafer is not trivial. It is not, however, a free efficiency gain.

Cerebras claims that CS-4 solutions can provide up to ten times more throughput per watt than CS-3. Its public launch material attributes the figure to internal benchmarking and projections, without enough methodology to reconcile it with the physical rack. SemiAnalysis reaches a more conservative conclusion: because performance and wafer power rise together, standalone rack-level performance per watt appears to improve only modestly.

Both statements may describe different operating points. Cerebras’ figure may depend on disaggregation, utilisation, batching or a specified interactivity threshold rather than raw wafer efficiency.

This distinction matters because Cerebras’ economic proposition is not simply “use less electricity.” It is that fast tokens are more valuable than slow tokens. If customers will pay a premium for dramatically higher interactivity, a rack can create more revenue even without a commensurate improvement in raw performance per watt.

The greenfield wedge

A 125 to 135 kW rack will not fit comfortably into much of the installed data-centre base. It requires high-density electrical distribution and direct liquid cooling designed to support it. For a conventional enterprise facility, that can make CS-4 difficult to adopt regardless of its inference performance.

For a new AI campus, however, the opportunity is clear. When the busways, cooling distribution units, pipework and facility water loops are still being specified, CS-4’s requirements can become design inputs rather than retrofit problems. The front-side power infrastructure can be deployed with the building, while scarce and expensive compute modules arrive later and are inserted as capacity becomes available.

Cerebras is not merely asking an operator to buy a faster inference machine. It is offering a repeatable rack envelope around which new capacity can be planned.

That does not make greenfield infrastructure a Cerebras monopoly. NVIDIA and other accelerator vendors are also driving the industry toward 100 kW-plus, liquid-cooled racks. In one sense, Cerebras may be converging with the emerging AI-factory standard rather than creating a unique facility advantage.

The sharper conclusion is that greenfield construction acts as a customer-selection mechanism. CS-4 will be most attractive to operators that can optimise the facility around high-density inference, value extreme interactivity, and deploy enough capacity for manufacturing and commissioning time to become financially material.

The power envelope narrows the market. Within that narrower market, Nexus may be a better product.

Three-part overview showing WSE-3 Turbo wafers, wafer power and cooling, and the Nexus rack-scale platform
Three WSE-3 Turbo wafers per system, with the Nexus rack-scale platform. Credit: Cerebras, August 2026

Building the upgrade path before the next chip

SemiAnalysis reports that the new system was designed from the outset for CS-4, CS-5 and CS-6. Its next wafer-scale engine is expected to arrive with CS-5 in 2027, and the company has committed to doubling performance annually, with up to 20× more throughput targeted by 2027. An impressive achievement if true.

If the physical and facility interfaces remain sufficiently stable, operators could install Nexus infrastructure today and obtain a cleaner path to subsequent Cerebras generations. That would give Cerebras something it has historically lacked: an installed-base advantage at the rack and facility level.

But this remains upgrade optionality rather than proven lock-in. Cerebras has not specified which chassis, power, cooling or I/O components will be retained through CS-5 and CS-6. A platform designed for several generations is not necessarily a promise that future compute backpacks can be dropped into every existing rack without other changes.

The answer will matter. If the same site preparation and significant portions of the rack survive multiple generations, Nexus becomes infrastructure. If each performance doubling requires a new power or cooling envelope, much of the modularity benefit will remain inside Cerebras’ factory rather than accrue to its customers.

A complete rack inside an incomplete system

Nexus simplifies Cerebras’ physical product at the same time that the company is embracing a more heterogeneous inference architecture.

Inference has two very different phases. Prefill processes an incoming prompt in parallel and is generally compute intensive. Decode generates tokens sequentially and is often constrained by how quickly model weights can be read from memory. Cerebras’ enormous on-wafer SRAM bandwidth is particularly well suited to the latter.

So CS-4 is simultaneously becoming a more complete physical product and a more specialised part of the logical system. The rack gets simpler. Orchestration across prefill, decode, KV-cache transfer, model placement and failure domains gets more complicated.

That is probably the correct strategic trade. Cerebras does not need to replace every GPU to build an important business. It needs to become the preferred engine for the part of inference where its architecture is structurally advantaged.

From fast tokens to commissioned capacity

The periodic-table demo made CS-4’s speed immediately legible. The larger ambition takes longer to see.

Cerebras is trying to turn an extraordinary piece of silicon into standardised hyperscale infrastructure. Nexus addresses the less glamorous work required to do that: power delivery, cooling, component count, factory automation, on-site installation, networking and future upgrades.

The greenfield opportunity follows from this systems approach. New AI data centres are being designed around unprecedented power densities. An operator that chooses its inference architecture early can configure the building around it, install reusable infrastructure ahead of compute delivery, and potentially carry that investment into subsequent hardware generations.

There are still important unknowns. The rack’s official power specification has not been published. The 10× throughput-per-watt claim lacks sufficient methodology. Multi-generation compatibility has been promised at the platform level but not defined at the component level. A modular rack will also place greater responsibility on Cerebras for field service, spares and multi-vendor debugging.

Even with those caveats, the direction is clear. Cerebras no longer wants to be evaluated only as the company that built the world’s largest chip or produced the fastest demo. It wants to be specified into the facility.