Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster conclusion compared to GPUs, enhanced economics, and a elemental way todeploy hyperscale capacity. It is the architecture for frontier AI.


Three WSE-3 Turbo per System
Each wafer delivers up to 2x the velocity of the erstwhile generation

More Performance per Wafer
All caller power, cooling, and I/O unleashes moreover much capacity per wafer

Nexus Rack-Scale Platform
Enables accelerated deployment successful hyperscale datacenters
Up to 30x faster than GPUs
Powered by WSE-Turbo, CS-4 delivers up to 30x faster conclusion compared to GPU systems, mounting a new record for the fastest conclusion disposable successful production.
Higher ultrafast throughput
The CS-4 solution shifts the conclusion Pareto frontier, delivering up to 10x much throughput per watt than CS-3 while generating tokens up to 30x faster than accumulation GPU systems. The consequence is simply a strategy designed to present some throughput and interactivity.
Frontier-ready architecture
By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers much than 1,000 tokens per 2nd connected models exceeding 10 trillion parameters, preserving interactive decode capacity astatine unprecedented scale.
CS-4 is the first loop of the caller Cerebras Nexus Platform Architecture. It is built astir a modular conception pinch 3 foundational elements: Compute, Power, and I/O – each pinch important invention to simplify manufacturing, deployment, maintenance, and upgrades.

Modular compute backpack design
Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is simply a self-contained assembly thatfolds the wafer, powerfulness conversion, nonstop liquid cooling, high-speed I/O, and power electronics into a compact 3D package pinch 50% less components. This creation simplifies manufacturing and reduces deployment clip from days to hours.
High-density powerfulness delivery
With powerfulness transportation conscionable 0.5 millimeters distant from the processor - astir 100x person than the astir 50mm of accepted GPU boards - CS-4 astir eliminates board-level powerfulness loss. This enables the transportation of doubly arsenic overmuch powerfulness to the WSE-3T, enabling higher operating frequencies and faster token generation.
Next-gen wafer I/O interface
CS-4 introduces a caller programmable I/O subsystem that doubles I/O bandwidth and reduces latency,benefitting some aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to beryllium linked wrong and across racks without a switch,for wafer-to-wafer latency as low arsenic 2 microseconds that is cardinal to interactivity for models pinch tens of trillions of parameters.
Deploy infrastructure past compute
CS-4 separates the unchangeable power, cooling, and web furniture from its modular wafer-scale compute. The Cerebras PowerRack tin beryllium installed and facility-qualified earlier compute arrives. Compute backpacks past descent into spot and link to power, cooling, and data—reducing deployment from days to hours while simplifying work and early upgrades astatine hyperscale.
First CS-4 shipments statesman this quarter.
Bring the fastest AI to your information center.

English (US) ·
Indonesian (ID) ·