Facebook
Britain's News Portal
Around The Clock
BREAKING
Loading latest headlines…

Cerebras unveils WSE-3T chip and CS-4 rack systems, doubling per-chip performance

Cerebras has announced its next-generation Wafer Scale Engine (WSE-3T) and Nexus rack systems, aiming to boost throughput per watt tenfold over the previous generation. The new WSE-3T chip offers twice the compute, memory fabric, and I/O bandwidth of its predecessor.

  • The WSE-3T chip delivers 250 petaFLOPS of AI compute and 43.2 PB/s of memory bandwidth.
  • Cerebras' CS-4 rack systems can house up to three WSE-3T chips, moving to a modular architecture.
  • The WSE-3T achieves performance gains by pushing its existing wafer scale engine harder, rather than using new silicon.

Cerebras has introduced its next-generation Wafer Scale Engine (WSE) and Nexus rack systems, with the company aiming to increase throughput per watt tenfold compared to the previous generation. The new WSE-3T chip, where 'T' stands for 'Turbo', promises twice the compute, memory fabric, and I/O bandwidth of the two-year-old WSE-3.

The WSE-3T achieves these performance improvements by pushing its existing wafer scale engine harder, rather than through new silicon. This is reportedly due to innovations in power delivery, allowing twice the power through the chip for higher operating frequencies and faster token generation.

Each WSE-3T chip offers 250 petaFLOPS of AI compute, 44 GB of SRAM, 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity. For inference, Cerebras' chips now function primarily as decode accelerators, partnering with Amazon Web Services (AWS) and AMD to offload compute-intensive prompt processing to their respective Trainium XPUs and Instinct GPUs.

The company is also moving to rack-scale compute architectures with its CS-4 systems. These modular racks can be equipped with up to three 'backpack' form factors, each housing a WSE-3T chip. Each chip is equipped with 2.4 Tbps of chip-to-chip bandwidth, an increase from 1.2 Tbps, and latency has been reduced from five microseconds to two.

Why this matters: The new systems aim to address memory bandwidth bottlenecks in high-speed AI inference and offer a modular approach to AI compute infrastructure.

Related Articles

Get the news that matters.

Join thousands of readers getting the best of British news straight to their inbox.