Cerebras has introduced its next-generation Wafer Scale Engine (WSE-3T) and Nexus rack systems, designed to significantly enhance AI inference capabilities. The company aims to increase throughput per watt tenfold compared to its previous generation.
The WSE-3T, which stands for "Turbo," doubles the compute, memory fabric, and I/O bandwidth of the two-year-old WSE-3. This performance increase is achieved by pushing the existing wafer scale engine harder, primarily through more efficient power delivery that allows for higher operating frequencies and faster token generation.
Each WSE-3T chip delivers 250 petaFLOPS of AI compute, 44 GB of SRAM, 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity. Cerebras has also introduced new rack-scale compute architectures with its CS-4 systems, which can be equipped with up to three WSE-3T accelerators.
For inference, Cerebras' chips are now designed to function primarily as decode accelerators, partnering with Amazon Web Services (AWS) and AMD to offload compute-intensive prompt processing to their Trainium XPUs and Instinct GPUs.