Cerebras launches CS-4 AI accelerator claiming up to 30x faster performance
Cerebras unveiled the fourth generation of its AI accelerator, CS-4, built on Wafer Scale Engine 3 Turbo processors, claiming up to 30x faster inference than GPU systems. The new Nexus system features a modular rack design with improved energy efficiency, ultra-low latency inter-processor communication at 2 microseconds, and native support for disaggregated inference. The platform combines high interactivity with greater throughput per watt, and first shipments are scheduled to begin in Q3 2026.
If I'm not mistaken, it's thanks to Cerebras that OpenAI can offer its AI models almost for free.
What is a Wafer Scale Engine and what performance does it deliver?
The Wafer Scale Engine is Cerebras's custom processor manufactured as an entire silicon wafer rather than traditional small chips. The third generation (WSE-3) used in CS-4 achieves lower inter-processor communication latency and higher operating frequencies, enabling token generation up to 30x faster than GPUs.
How does CS-4 compare to the previous generation CS-3?
CS-4 delivers up to 10x more token capacity and up to 2x faster performance than CS-3. Key improvements include the new Nexus platform design with modularity, more efficient power delivery (2x more power to processors), and a programmable I/O subsystem.
What is disaggregated inference and when is it useful?
Disaggregated inference splits processing into two phases: prefill (prompt processing, typically on GPU or ASIC) and decode (response generation, optimized on CS-4). This approach allows operators to combine efficient prefill with ultrafast decode and reduces energy costs.
- OpenAI and Cerebras accelerate GPT-5.6 Sol with Ultrafast Mode — cerebras.ai 75 % match
- ASUS ExpertCenter Pro ET900N G3 and Ascent GX10 bring AI supercomputer to the desktop — edgeup.asus.com 73 % match
- DeepSeek makes its cheapest model more powerful: DeepSeek-V4-Flash — huggingface.co 73 % match
- Cerebras3
- CS-4
- Wafer Scale Engine 3 Turbo
- Nexus
- AMD Helios
- AWS Trainium
- RoCE v2 RDMA