Large-scale AI training has moved beyond single data centres, with companies like Google, Microsoft, and AWS connecting AI compute clusters across wide areas. This geographic spread is partly due to the power and space demands of new AI models, which can strain individual data centres.
Training modern AI models can require clusters with tens of thousands of GPUs. These distributed systems need to exchange results and synchronize at various points before the next stage of training can begin. This involves transferring large arrays of numerical data, such as gradients and intermediate results, between thousands of accelerators simultaneously.
Ramesh Sivakolundu of Cisco's Silicon One architecture team explained that training jobs involve a repetition of compute and synchronization. This can lead to "incast" problems where many senders converge on a single destination, potentially causing traffic to arrive faster than the network can handle. A network delay in synchronous training can become a compute delay, as GPUs cannot proceed until the network-wide exchange is complete.
Itamar Gold, director of product management at Cisco, noted that inter-site networks for coordinated AI workloads may be more constrained than networks within a single facility. Cisco estimates that connecting two 100MW AI sites could require 12,000 to 32,000 coherent optical ports at the scale-across layer, significantly more than for conventional data centre interconnects.