Nvidia has provided the first performance figures for its Groq 3-based LPX rack systems, showing a significant speedup in AI inference. An independent benchmark conducted by Artificial Analysis recorded the systems processing 3,400 tokens per second (tok/s) with a 100,000-token input sequence using Google's Gemma 4 31B model.
According to Nvidia, this performance is four times faster than the closest competing platform, which Artificial Analysis' leaderboard indicates is Cerebras, achieving 882 tok/s under the same conditions.
Groq, acquired by Nvidia in late December, develops LPUs with an SRAM-heavy dataflow architecture optimised for high-performance inference serving. Unlike traditional GPUs that rely on DRAM, Groq's chips use a large pool of on-die SRAM, which offers significantly faster memory bandwidth.
Each Groq 3 LPU has 500 MB of memory, requiring Nvidia's architecture to distribute models across multiple accelerators using Ethernet. An LPX rack can be equipped with up to 256 LPUs, providing 128 GB of high-bandwidth SRAM for larger models.
While the 3,400 tok/s figure is notable, the Gemma 4 31B model is considered a best-case scenario for the hardware. The model, at 31 billion parameters, fits within a single LPX rack. Nvidia also expects further performance gains by combining its GPUs with Groq 3 LPUs in a heterogeneous inference architecture, where GPUs handle the compute-heavy prefill phase and LPUs manage the memory-bandwidth intensive decode phase.
Netherlands-based neocloud Nebius is expected to be among the first to deploy these combined systems in its datacentres.