Nvidia announced this week that its new Groq-3-based LPX racks have entered production, with early tests showing the systems generating 3,400 tokens per second in Gemma 4 31B. This speed is reportedly four times faster than rival Cerebras.
A day later, Cerebras responded by touting nearly equivalent performance from its next-gen CS-4 accelerators. However, these top-line performance figures, while making inference feel instantaneous, may not be representative of how customers will use the systems in production.
The benchmarks, conducted by Artificial Analysis, are achievable but are considered a marketing gimmick for inference-as-a-service model operators. The systems' SRAM-heavy architectures, while providing massive bandwidth, limit their ability to scale performance to a large number of users due to memory constraints.
For example, a single LPX rack with 256 LPUs or a CS4 rack with three WSE-3T accelerators may only manage a maximum batch size of 12 for a 100,000-token input length before running out of memory. This limitation is not due to compute power but insufficient memory to support larger batches without additional racks or reduced prompt sizes.
Combining GPUs with Groq or Cerebras accelerators is seen as a way to make "premium inference" more cost-effective. Nvidia has licensed Groq's technology, and Cerebras has partnered with AWS and AMD, suggesting a move towards heterogeneous compute architectures where these accelerators function as decode accelerators alongside GPUs.