Facebook
Britain's News Portal
Around The Clock
BREAKING
Loading latest headlines…

Nvidia and Cerebras AI performance claims may not reflect real-world use

Nvidia and Cerebras have announced high token generation speeds for their new AI systems, but these top-line figures may not be achievable in typical production environments.

  • Nvidia's new Groq-3-based LPX racks reportedly achieved 3,400 tokens per second in Gemma 4 31B tests.
  • Cerebras' next-gen CS-4 accelerators are touted to offer nearly equivalent performance.
  • These high performance figures are for single-request scenarios and may not scale efficiently for multiple users due to memory limitations.

Nvidia announced this week that its new Groq-3-based LPX racks have entered production, with early tests showing the systems generating 3,400 tokens per second in Gemma 4 31B. This speed is reportedly four times faster than rival Cerebras.

A day later, Cerebras responded by touting nearly equivalent performance from its next-gen CS-4 accelerators. However, these top-line performance figures, while making inference feel instantaneous, may not be representative of how customers will use the systems in production.

The benchmarks, conducted by Artificial Analysis, are achievable but are considered a marketing gimmick for inference-as-a-service model operators. The systems' SRAM-heavy architectures, while providing massive bandwidth, limit their ability to scale performance to a large number of users due to memory constraints.

For example, a single LPX rack with 256 LPUs or a CS4 rack with three WSE-3T accelerators may only manage a maximum batch size of 12 for a 100,000-token input length before running out of memory. This limitation is not due to compute power but insufficient memory to support larger batches without additional racks or reduced prompt sizes.

Combining GPUs with Groq or Cerebras accelerators is seen as a way to make "premium inference" more cost-effective. Nvidia has licensed Groq's technology, and Cerebras has partnered with AWS and AMD, suggesting a move towards heterogeneous compute architectures where these accelerators function as decode accelerators alongside GPUs.

Why this matters: The reported performance figures for AI systems may not reflect their practical utility and scalability in real-world applications, potentially influencing purchasing decisions for businesses relying on AI inference.

Related Articles

Get the news that matters.

Join thousands of readers getting the best of British news straight to their inbox.