Cerebras CS4

Aug 19, 2026 07:28 AM - 4 weeks ago 4

Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster conclusion compared to GPUs, enhanced economics, and a elemental way todeploy hyperscale capacity. It is the architecture for frontier AI.​

Three WSE-3 Turbo per System​

Each wafer delivers up to 2x the velocity of the erstwhile generation​

More Performance per Wafer​

All caller power, cooling, and I/O unleashes moreover much capacity per wafer​

Nexus Rack-Scale Platform

Enables accelerated deployment successful hyperscale datacenters​

Up to 30x faster than GPUs​

Powered by WSE-Turbo, CS-4 delivers up to 30x faster conclusion compared to GPU systems, mounting a new record for the fastest conclusion disposable successful production.​

Higher ultrafast throughput

The CS-4 solution shifts the conclusion Pareto frontier, delivering up to 10x much throughput per watt than CS-3 while generating tokens up to 30x faster than accumulation GPU systems. The consequence is simply a strategy designed to present some throughput and interactivity.​

Frontier-ready architecture

By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers much than 1,000 tokens per 2nd connected models exceeding 10 trillion parameters, preserving interactive decode capacity astatine unprecedented scale.​

CS-4 is the first loop of the caller Cerebras Nexus Platform Architecture. It is built astir a modular conception pinch 3 foundational elements: Compute, Power, and I/O – each pinch important invention to simplify manufacturing, deployment, maintenance, and upgrades.​

Modular compute backpack design

Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is simply a self-contained assembly thatfolds the wafer, powerfulness conversion, nonstop liquid cooling, high-speed I/O, and power electronics into a compact 3D package pinch 50% less components. This creation simplifies manufacturing and reduces deployment clip from days to hours.​

High-density powerfulness delivery

With powerfulness transportation conscionable 0.5 millimeters distant from the processor - astir 100x person than the astir 50mm of accepted GPU boards - CS-4 astir eliminates board-level powerfulness loss. This enables the transportation of doubly arsenic overmuch powerfulness to the WSE-3T, enabling higher operating frequencies and faster token generation.​

Next-gen wafer I/O interface

CS-4 introduces a caller programmable I/O subsystem that doubles I/O bandwidth and reduces latency,benefitting some aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to beryllium linked wrong and across racks without a switch,for wafer-to-wafer latency as low arsenic 2 microseconds that is cardinal to interactivity for models pinch tens of trillions of parameters.​

Deploy infrastructure past compute

CS-4 separates the unchangeable power, cooling, and web furniture from its modular wafer-scale compute. The Cerebras PowerRack tin beryllium installed and facility-qualified earlier compute arrives. Compute backpacks past descent into spot and link to power, cooling, and data—reducing deployment from days to hours while simplifying work and early upgrades astatine hyperscale.​

First CS-4 shipments statesman this quarter.​
Bring the fastest AI to your information center.​

More