AMD has launched its first full rack-scale AI compute platform, Helios, marking a significant move to challenge Nvidia's long-standing dominance in the data centre and artificial intelligence (AI) hardware market. The company claims the new system, powered by its Instinct MI455X GPUs, offers superior performance and efficiency compared to existing and upcoming Nvidia platforms, including the Vera Rubin system, particularly for AI training applications.
On paper, the Helios system, which incorporates 72 GPUs, appears to surpass Nvidia's Blackwell-based rack systems and the forthcoming Vera Rubin across several key metrics. AMD highlights that Helios provides 50% more High Bandwidth Memory (HBM4) and scale-out bandwidth, alongside an estimated 15% to 25% higher performance for AI training. While Nvidia's Vera Rubin may hold an advantage in FP4 inference workloads due to adaptive compression, Helios is projected to deliver 15% higher peak FP4 FLOPS where such compression is not applicable. This direct competition in performance at launch is a notable shift, as AMD's previous high-performance products often followed Nvidia's equivalents by a year.
The core of Helios's performance lies in its new Instinct MI455X GPU, built on AMD's 5th-generation CDNA compute architecture. This advanced GPU is a complex silicon package, integrating 24 chiplets using a combination of 2.5D and 3D packaging. Its eight compute dies are fabricated using TSMC's cutting-edge 2nm process technology, stacked upon 3nm fabric and cache dies (FCDs). These FCDs act as a cache-heavy interposer, providing 96 MB of L2 cache each and managing the 12, 36 GB HBM4 memory stacks. This intricate design allows the chip to function either as a single large GPU or two smaller ones, with support for spatial partitioning into up to eight virtual GPUs.
The MI455X also introduces significant architectural enhancements over its predecessor, the MI355X. It is designed to prioritise AI-centric datatypes like MXFP4 and MXFP8, foregoing FP64 entirely in this specific SKU to maximise die area for AI workloads. This focus is expected to deliver up to four times higher floating-point performance for AI applications compared to the MI355X. Additionally, the new generation incorporates a larger shared L2 cache and simplifies the data path by removing the last-level 'Infinity' cache, which AMD fellow Alan Smith states results in 1.5 times the aggregate bandwidth of the MI355X's Infinity cache.
For UK businesses and consumers, the emergence of a strong challenger to Nvidia's AI hardware dominance could lead to a more competitive market, potentially driving down costs and accelerating innovation in AI development. Increased competition means UK businesses investing in AI infrastructure, from cloud service providers to research institutions, may have more choice and better value for money. This could enable more organisations to access and leverage high-performance AI compute, fostering growth in various sectors.