Large‑scale AI training has already escaped the confines of a single datacenter. Google said Gemini was trained synchronously across clusters in multiple locations; Microsoft has linked AI data centers in Wisconsin and Georgia into a distributed supercomputer; AWS has connected compute clusters for Anthropic’s Claude models; Meta has built high‑capacity interconnects for model training; and CoreWeave and Google Cloud recently announced cross‑cloud training via a private interconnect, with Azure likely to follow later this year.

This geographic spread is driven by the limits of individual facilities, which struggle to meet the compute and power demands of new models. Cisco estimates that today’s training clusters may require tens of thousands of GPUs. By 2030, the largest frontier runs could draw 4‑16 GW of power, according to Epoch AI. Distributing compute lets hyperscalers locate resources where power and space are abundant and planning constraints fewer.

How to train an LLM

At the heart of an LLM is a neural network with billions of parameters adjusted during training. A batch of data passes through the model, which makes a prediction. The system measures error and calculates weight changes. The weights are updated and the cycle repeats. Modern models distribute this work across thousands of GPUs and accelerators such as AWS Trainium and Google TPUs.

Different GPUs can process separate batches or model partitions, but they cannot operate entirely independently. At certain stages they must exchange results and synchronize before the next training step.

The data moving between accelerators is not the training text or images but large arrays of numerical data such as gradients and intermediate results. Thousands of accelerators may need to exchange and combine this data almost simultaneously.

“If you look at any training job, it is basically a repetition of compute and synchronization,” explains Ramesh Sivakolundu of Cisco’s Silicon One architecture team. “You compute the weights, communicate those weights to synchronize every GPU in the cluster, and then restart the computation. That means bursty traffic happens periodically throughout the training process.”

Thousands of accelerators can finish a phase and begin communicating at the same time, leading to incast problems where many senders converge on a single destination and traffic arrives faster than the link can handle.

The fundamental issue is synchronization. In synchronous training, one GPU group cannot ignore a slower partner and continue alone. The network‑wide exchange must complete before any participant proceeds, so network delay becomes compute delay.

“You don’t want the network to be the bottleneck in the training process,” says Sivakolundu. “The idea is to keep the GPUs fully occupied and not have the network introduce additional latency into the overall training program.”

Occasional congestion or packet loss does not halt training entirely, but a delayed or dropped flow may require retransmission, prolong collective operations, and leave other accelerators waiting. More severe failures can force a rollback to a saved checkpoint.

Different kind of interconnect

“Inside a single datacenter, you can assume a full mesh of GPU racks with the bandwidth you planned,” says Itamar Gold, director of product management at Cisco. Gold notes that with scale‑across, AI workloads may traverse a more constrained inter‑site fabric than the internal network. Traffic destined for another facility may be funneled through a narrower pipe, which can become a choke point during synchronized bursts.

Traditional DCI links independent data centers and is built for redundancy, reach, and asynchronous workload distribution. Scale‑across AI training makes the inter‑site network part of a coordinated computation, demanding sufficient bandwidth and predictable delivery.

From rack to region

Within a rack, bandwidth can reach 800 Gbps–1.6 Tbps, but distances are short. High‑speed serial lanes make a GPU cluster appear as a single compute unit. When the fabric is extended across a facility, “scale‑out” links still require up to 1.6 T. Copper is limited to short reaches; longer connections use pluggable optics, with scale‑out links spanning hundreds of meters to about 2 km.

“Scale‑across” extends a coordinated AI fabric across multiple facilities, making inter‑site network behavior critical to workload performance. Depending on architecture, links may cover metro or regional distances. Cisco modeling shows aggregate bandwidth requirements can be roughly 14 × a conventional DCI baseline.

Short‑range datacenter optics are not designed for 400‑800 Gbps signals over hundreds of kilometers. Operators need coherent optics that use sophisticated modulation and high‑performance DSP to keep high‑rate signals usable over metro and regional fiber spans, compensating for physical impairments. The number of ports also matters. Cisco estimates that connecting two 100 MW AI sites could require 12,000–32,000 coherent optical ports at the scale‑across layer, versus roughly 1,000–2,000 for conventional DCI. At 400 G or 800 G per port, this adds up to multi‑petabit aggregate capacity.

Moving that much data is also a power and space challenge. Coherent pluggables integrate the transponder function into router/switch ports, avoiding separate DWDM transponders and reducing rack space, power, and cooling—an important consideration for already power‑constrained AI factories.

The latency problem

Even with ample bandwidth and suitable optics, distance adds latency. If congestion occurs, the network needs time to signal senders to slow down. Over a local AI fabric, feedback is rapid, but across 100 km of fiber, much more data can be in flight before the sender learns of the problem. Cisco calculates that on an 800 Gbps link spanning 100 km, about 100 MB can already be transmitted during that feedback interval.

Longer feedback loops increase the value of networking silicon with deeper buffering. Shallow‑buffer switching architectures designed for ultra‑low latency inside a datacenter have less capacity to absorb traffic while congestion feedback traverses a long‑distance link.

“The reason you need deeper buffers as the distances get longer is that you don’t get feedback about the transfer of packets for a longer period of time,” says Sivakolundu. “If you don’t buffer deeply enough, you can end up dropping and retransmitting packets, increasing latency.”

Deep buffering absorbs temporary bursts or disruptions while traffic management and congestion signaling address the underlying problem. Sustained congestion still requires more capacity or a change in traffic distribution.

The same buffer helps with oversubscription. If combined upstream capacity exceeds the inter‑site link, a synchronized burst can arrive faster than the link can drain it. Buffering gives excess traffic a place to wait without being dropped while congestion controls react.

Rather than dividing buffer into fixed amounts per port, a shared buffer is a common pool of packet memory that can be allocated wherever congestion appears. This lets a busy link absorb larger bursts without dropping packets.

“The longer you’re reaching out, the more you may need to buffer,” says Gold. “You also need the flexibility to put as much buffer as possible where it is needed… Our approach is a single, fully shared buffer.”

AI training offers a predictable pattern: its collective communication is repetitive, making it more predictable than traditional DCI traffic. Cisco uses proactive congestion management to steer or schedule traffic before a predictable burst creates a problem, while retaining deep buffering as a safety net for transient congestion and unpredictable events such as link failure. The two approaches complement each other: one tries to avoid the queue, the other provides a place for traffic when a queue forms.

Cisco’s bet on co‑design

Cisco’s scale‑across architecture uses the Silicon One P200, a 51.2 Tbps programmable, deep‑buffer routing processor featured in the Cisco 8223 and Cisco N9000 platforms. These systems can pair with Cisco’s 400 G and 800 G coherent pluggables for long‑distance links, while open line systems such as Cisco Open Transport 3000 and Cisco NCS 1014 handle the optical transport layer.

Coherent pluggables generate DWDM wavelengths directly from the routing system, while the line system amplifies and carries those wavelengths across the fiber plant. For multi‑pair links, Cisco Open Transport 3000 uses a multi‑rail design that combines optical components onto one line card, reducing power per rail by 75 % and rack space by 80 %. Where a separate transport system is required, the NCS 1014 can provide 12.8 Tbps from a 1RU line card.

The P200 includes a programmable run‑to‑completion network processor and P4 tooling, allowing protocol support, telemetry, and other packet‑processing features to be developed in software rather than waiting for a new silicon generation.

For scale‑across, where operators are still testing different topologies and traffic‑management methods, the same programmable flexibility can let network functions evolve without forcing a chip replacement for each new requirement,” says Sivakolundu.

Cisco also includes hardware‑based protection for inter‑site traffic, arguing that encryption should not create another processing bottleneck for the training fabric.

A key part of Cisco’s pitch is co‑design. Because its silicon, system, and optics teams sit within the same company, requirements for port density, buffering, power, telemetry, and optical interfaces can be fed back into chip design. This avoids the system team receiving a commercial ASIC and having to work out what it can build around it.

Scale‑across has not settled on a single approved architecture. Cisco itself divides the problem into campus, metro, and regional deployments, and operators are experimenting with various combinations of networking and training techniques. Implementations vary by distance, topology, and workload design, but the common goal is maintaining coordinated AI performance across sites. The technology is still early; “We are seeing early deployments now, but it is a process and I think it will take a few years,” says Gold.

Designing a network for distributed training will not eliminate the latency of distance, but it can keep the interconnect from becoming the bottleneck. That gives hyperscalers and datacenter operators the flexibility to add compute where power and space are available while still treating multiple sites as part of the same training infrastructure.

Source link

Exit mobile version