Facebook
Britain's News Portal
Around The Clock
BREAKING
Loading latest headlines…

Nvidia's Groq 3 LPU systems achieve 3,400 tokens/second in AI benchmark

Nvidia's new LPX rack systems, featuring Groq 3 LPU technology, have demonstrated a speed of 3,400 tokens per second using Google's Gemma 4 31B model in an independent benchmark. Nvidia states this makes its platform four times faster than its nearest competitor.

  • Nvidia's LPX rack systems achieved 3,400 tokens/second with a 100,000-token input sequence in the Gemma 4 31B model.
  • Nvidia claims this performance is 4x faster than the nearest alternative platform, which Artificial Analysis' leaderboard suggests is Cerebras.
  • The Groq 3 LPUs utilise an SRAM-heavy dataflow architecture designed for high-performance inference serving.

Nvidia has provided the first performance figures for its Groq 3-based LPX rack systems, showing a significant speedup in AI inference. An independent benchmark conducted by Artificial Analysis recorded the systems processing 3,400 tokens per second (tok/s) with a 100,000-token input sequence using Google's Gemma 4 31B model.

According to Nvidia, this performance is four times faster than the closest competing platform, which Artificial Analysis' leaderboard indicates is Cerebras, achieving 882 tok/s under the same conditions.

Groq, acquired by Nvidia in late December, develops LPUs with an SRAM-heavy dataflow architecture optimised for high-performance inference serving. Unlike traditional GPUs that rely on DRAM, Groq's chips use a large pool of on-die SRAM, which offers significantly faster memory bandwidth.

Each Groq 3 LPU has 500 MB of memory, requiring Nvidia's architecture to distribute models across multiple accelerators using Ethernet. An LPX rack can be equipped with up to 256 LPUs, providing 128 GB of high-bandwidth SRAM for larger models.

While the 3,400 tok/s figure is notable, the Gemma 4 31B model is considered a best-case scenario for the hardware. The model, at 31 billion parameters, fits within a single LPX rack. Nvidia also expects further performance gains by combining its GPUs with Groq 3 LPUs in a heterogeneous inference architecture, where GPUs handle the compute-heavy prefill phase and LPUs manage the memory-bandwidth intensive decode phase.

Netherlands-based neocloud Nebius is expected to be among the first to deploy these combined systems in its datacentres.

Why this matters: Faster inference servers could lead to more responsive and capable AI agents and code assistants, potentially enhancing the performance of AI applications.

Related Articles

Get the news that matters.

Join thousands of readers getting the best of British news straight to their inbox.