Cerebras CS-4 Connects Three Wafers in Parallel, Claims 30x Faster Inference Than GPU

Cerebras unveiled the CS-4 inference system on August 18, 2026, using three parallel WSE-3 Turbo wafers to deliver over 1000 tokens per second on trillion-parameter models — up to 30x faster than GPU solutions, with 10x per-watt throughput gains over the CS-3.

On August 18, 2026, Cerebras released the CS-4 inference system. The system consists of three WSE-3 Turbo wafers connected in parallel. Official data shows it can output over 1000 tokens per second on models with over a trillion parameters, up to 30x faster than GPU solutions, with 10x improvement in per-watt throughput over the previous-generation CS-3.

Specific Design of the Disaggregated Architecture

The CS-4 employs the Nexus rack-level platform, reducing inter-wafer communication latency to 2 microseconds and supporting models with over 50 trillion parameters. The system delivers total compute power of 750 PFLOPs, memory bandwidth of 129.6 PB/s, and inter-wafer I/O bandwidth of 7.2 Tbps.

In AI, speed is productivity. Cerebras CS-4 delivers industry-leading speed on the largest frontier models, fundamentally changing the paradigm. — Andrew Feldman, CEO and co-founder of Cerebras

In partnership with OpenAI, it launched GPT-5.6 Sol Ultrafast mode, which is 14x faster than the standard mode. In collaboration with AMD, a disaggregated inference architecture was developed, with AMD Helios handling the Prefill phase and Cerebras handling the Decode phase, further improving combined throughput by 5x.

Practical Impact on Inference Infrastructure

The disaggregated architecture separates Prefill from Decode, theoretically improving overall throughput. The CS-4's low-latency inter-wafer interconnect addresses network latency and power distribution challenges in cross-vendor coordination.

The 10x improvement in per-watt throughput means more tokens can be generated under the same power budget. This directly impacts data center profitability, especially in agentic application scenarios that require highly interactive experiences.

  • Developers get more responsive inference services, suitable for real-time inference and multi-turn dialogue.
  • Operators get higher overall throughput under fixed power constraints.
  • New entrants can quickly scale up through modular racks, lowering upfront hardware barriers.

Analysis of Underlying Drivers

NVIDIA's long-standing dominance in the inference market stems from the CUDA ecosystem and a mature software stack. Cerebras chose the wafer-scale integration path, bypassing the multi-chip interconnect bottleneck of traditional GPUs, with the core focus on solving communication latency for large-scale models. The AMD partnership shows an attempt to build a heterogeneous solution rather than compete in isolation.

The deeper driver of this strategy is the continuous growth in frontier model parameter counts, making it difficult for a single accelerator to maintain interactive speed. The CS-4's 2-microsecond inter-wafer latency directly targets this pain point.

Independent Assessment

The CS-4 offers a mass-producible hardware option with performance data backed by concrete sources. Its collaborations with OpenAI and AMD demonstrate commercial deployment intent. When choosing deployment solutions, developers should prioritize actual API stability and total cost of ownership over peak speed marketing claims.