Cerebras Systems (NASDAQ: CBRS) unveiled the CS-4 on August 18, its most powerful AI accelerator to date, built on a new rack-scale architecture called Nexus. The system houses three WSE-3 Turbo wafers per rack — each a 46,225 mm² chip containing four trillion transistors and 900,000 computing cores.
The CS-4 delivers 750 petaflops of AI compute, 7.2 terabits per second of I/O bandwidth, and 129.6 petabytes per second of memory bandwidth. Cerebras claims the system achieves up to 30x faster inference than GPU-based solutions while consuming roughly 10x less power per token.
The announcement positions Cerebras as a serious challenger to Nvidia's dominance in AI inference, particularly for large language models where token throughput and latency are critical. The system can process over 1,000 tokens per second on models with 10 trillion parameters.
"We didn't build a cheaper GPU. We built a chip for how AI actually works in 2026," the company said in its launch materials.
First shipments are scheduled for Q3 2026. The CS-4 arrives as demand for AI inference hardware surges, with major cloud providers and AI companies scrambling for compute capacity to serve growing workloads.




