Cerebras has introduced its next-generation Wafer Scale Engine (WSE) and Nexus rack systems, with the company aiming to increase throughput per watt tenfold compared to the previous generation. The new WSE-3T chip, where 'T' stands for 'Turbo', promises twice the compute, memory fabric, and I/O bandwidth of the two-year-old WSE-3.
The WSE-3T achieves these performance improvements by pushing its existing wafer scale engine harder, rather than through new silicon. This is reportedly due to innovations in power delivery, allowing twice the power through the chip for higher operating frequencies and faster token generation.
Each WSE-3T chip offers 250 petaFLOPS of AI compute, 44 GB of SRAM, 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity. For inference, Cerebras' chips now function primarily as decode accelerators, partnering with Amazon Web Services (AWS) and AMD to offload compute-intensive prompt processing to their respective Trainium XPUs and Instinct GPUs.
The company is also moving to rack-scale compute architectures with its CS-4 systems. These modular racks can be equipped with up to three 'backpack' form factors, each housing a WSE-3T chip. Each chip is equipped with 2.4 Tbps of chip-to-chip bandwidth, an increase from 1.2 Tbps, and latency has been reduced from five microseconds to two.