CS-3

Overview

Current production wafer-scale system, one WSE-3 per chassis, used for Cerebras Cloud, OpenAI Ultrafast, AWS disaggregated decode, and on-prem clusters. The 5nm WSE-3 is 46,225 mm2 with 4 trillion transistors, 900,000 AI cores, 44 GB of on-chip SRAM, and 21 PB/s of memory bandwidth, about 58 times the area of an Nvidia B200 die and thousands of times its on-chip memory bandwidth.

In volume production. Flex is expanding Milpitas lines for an expected ~7x CS-3 output increase through 2026; Sanmina and Rocket EMS are also adding lines. Still the workhorse for OpenAI, AWS, and the inference cloud while CS-4 ramps.

Key Specs

Process / size
TSMC 5nm, 46,225 mm2 (full wafer)
Transistors / cores
4 trillion transistors, 900,000 AI-optimized cores
On-chip memory
44 GB SRAM, 21 PB/s bandwidth
OpenAI Sol Ultrafast
Up to 750 output tokens/sec on GPT-5.6 Sol