The speedy evolution of synthetic intelligence is basically altering how we architect knowledge facilities. As AI fashions develop extra advanced, the trade is shifting focus from particular person server efficiency to the info heart’s interconnected cloth. Two components are driving this shift: increasing coaching clusters and inference workloads that now demand cluster-level efficiency.
For coaching, frontier fashions require massive numbers of GPUs, and cluster sizes now exceed the capability of a single knowledge corridor. Clusters span a number of knowledge facilities related by wide-area networks, and the infrastructure should scale to assist a whole lot of 1000’s of GPUs throughout broad geographic areas.
Inference can also be remodeling the infrastructure. Frontier fashions, even at FP4 precision, now surpass the capability of a single GPU. The push for sooner token serving is rising demand for bigger inference clusters, matching the identical coordinated, high-performance networking as coaching clusters.
Taken collectively, these adjustments make the community greater than a connectivity layer. The community is changing into the system-level cloth that determines how a lot of the AI infrastructure can be utilized, how rapidly jobs full, and the way predictably inference might be served.
The community is now the system
A number of years in the past, GPU compute energy was the first bottleneck for AI mannequin coaching. As distributed coaching has scaled, that constraint has shifted decisively from compute to community — GPU communication now determines general cluster effectivity. As an example, Meta’s manufacturing knowledge reveals that in large-scale Deep Neural Community coaching runs, community overhead accounts for as much as 60% of complete coaching iteration time — a share that will increase with cluster measurement.
This is the reason we take into consideration the following part of AI networking as a continuum. Scale-up connects accelerators inside a server or rack, the place proprietary applied sciences resembling NVLink and rising approaches resembling UALink, have centered on extraordinarily low latency and excessive bandwidth. Scale-out connects racks and pods into bigger coaching clusters, the place InfiniBand has traditionally been a standard alternative for high-performance materials. Scale-across connects clusters, storage, front-end networks, and knowledge facilities, the place Ethernet is already the operational basis.
At scale, for coaching and inference alike, the community issues as a lot because the compute itself. The query is not whether or not AI wants specialised networking conduct. It does. The actual query is whether or not we ship that conduct by means of a patchwork of proprietary materials, or by means of one widespread Ethernet basis that may develop throughout the entire continuum.
Why Ethernet turns into the widespread basis
Proprietary networking options have lengthy dominated high-performance computing, however they introduce vendor lock-in and restrict scalability throughout numerous {hardware}. InfiniBand nonetheless has a task in loads of AI deployments, however the route of the trade isn’t in query — Ethernet is changing into the predominant networking expertise for AI infrastructure. Embracing Ethernet places you on the suitable working mannequin from day one: open, interoperable, and constructed to scale throughout many domains.
Cisco is championing an “Ethernet-first” technique for AI for 3 core causes:
- Open Requirements and Interoperability: Ethernet allows organizations to combine parts from a number of distributors. This flexibility is important for future-proofing knowledge facilities as AI {hardware} evolves.
- Unmatched Scalability: InfiniBand’s proprietary cloth administration struggles above ~tens of 1000’s of GPUs, requiring advanced workarounds as clusters develop. Ethernet has no such ceiling — hyperscalers have already leveraged many years of mature switching structure and standards-based tooling to validate Ethernet-based clusters at a whole lot of 1000’s of GPUs throughout a number of knowledge facilities.
- Funding Safety — With a Studying Curve: Ethernet builds on acquainted infrastructure — present switching platforms, administration tooling, and a broad engineering expertise pool. That basis issues. However AI cloth operations shouldn’t be a straight extension of enterprise networking. RoCEv2 and RDMA introduce new failure modes; congestion administration (PFC, ECN, buffer tuning) requires cautious calibration to keep away from GPU stalls; and telemetry at hundred-thousand-GPU scale calls for purpose-built tooling. Expertise switch partially, not totally. The benefit over InfiniBand is a extra open, composable operational mannequin.
That working mannequin issues as a result of no two AI environments look alike. Coaching desires ultra-low latency and predictable collective communication. Inference desires QoS that accounts for load, location, and value. A multi-site deployment desires fault tolerance, tenant isolation, and deterministic telemetry stretched throughout a a lot larger failure area. Ethernet provides you one basis that may flex to all these necessities — as a substitute of sewing collectively a separate expertise island for every one.
What Ethernet should ship for AI
To earn its place because the widespread AI cloth, Ethernet should deal with what makes AI site visitors totally different. This site visitors is synchronized, bursty, and costly to stall. Fall behind on the community, and GPUs sit idle. Let congestion unfold, and job completion occasions stretch out. Take too lengthy to heal a failure, and enormous jobs lose effectivity.
First up: clever load balancing. AI materials should unfold site visitors throughout many paths with out sacrificing single-flow efficiency, maintaining tempo with fashionable NIC bandwidth and placing the entire topology to work. Weighted adaptive routing, multipath transport, source-routed and path-aware forwarding — these all serve the identical purpose: react to hotspots quick, with out introducing instability.
Second: congestion management and dependable supply. Which means quick congestion detection, exact notification, and restoration that doesn’t throw away helpful work. Packet trimming, native hyperlink restore, selective retransmission, ordered and unordered retransmission, header optimization — none of those are standalone options. They’re all doing the identical job: maintaining AI site visitors shifting when the material is underneath strain.
Third: isolation and repair assurance. AI clusters more and more run a number of tenants and a number of jobs aspect by aspect, and a fault or noisy neighbor in a single mustn’t ever degrade one other’s efficiency. Delivering that assure with out heavy per-job configuration — particularly as workloads transfer off InfiniBand — is what separates a cloth that merely connects GPUs from one that may be trusted to run manufacturing AI at scale.
That is precisely the place requirements like UEC, ESUN, and Multipath Dependable Connection (MRC) earn their maintain. They’re defining how Ethernet picks up the AI-specific conduct it wants — congestion management, multipath operation, dependable transport, path consciousness, telemetry, interoperability — with out giving up the openness that made Ethernet the suitable alternative to start with.
Ethernet plus P4 programmability: The multiplying issue
In AI, networking requirements are evolving quickly. New protocols resembling UEC Transport and MRC are being developed to deal with challenges in AI and ML site visitors, together with congestion management, environment friendly use of material bandwidth, packet ordering, and telemetry.
New requirements resembling these usually require capabilities in networking that may solely be met within the new ASIC era which is usually out there eighteen months later at finest.
Traditionally, this assumption made sense. ASICs are constructed to a hard and fast specification, and as soon as set, adjustments will not be potential. If a typical was not included within the unique design, it can’t be supported by the chip.
AI is difficult this mannequin.
AI workload necessities are evolving at an unprecedented tempo. UEC and MRC will not be minor updates; every introduces vital new capabilities required on the switching ASIC degree. These adjustments are arriving sooner than conventional silicon improvement cycles can assist.
This presents a major problem for purchasers constructing infrastructure right this moment. Delaying an AI buildout to attend for brand spanking new {hardware} shouldn’t be possible. The price of delay, together with misplaced coaching runs, decreased competitiveness, and idle capital, is substantial.
Cisco’s Silicon One was designed to deal with this problem.
Since Silicon One is programmable in P4: it’s not restricted to the preliminary set of functions envisioned when the ASIC was designed. P4 allows engineers and clients to outline packet processing in software program, separating community logic from bodily {hardware}. When a brand new normal emerges, resembling a revised UEC congestion response or new MRC capabilities, we will ship these updates in software program on present {hardware}, usually inside weeks or months relatively than ready for the following product cycle.
That’s the multiplying issue. Requirements set the route for the ecosystem, however P4 programmability decides how briskly clients see the profit on actual infrastructure. It additionally means customer-specific conduct — scheduler-aware coverage, topology-specific routing, tenant isolation — doesn’t have to attend on a fixed-function silicon roadmap.
The place Cisco Silicon One matches in
Cisco Silicon One sits proper on the intersection of high-performance Ethernet, rising AI networking requirements, and P4 programmability. That’s not a coincidence — AI networks want each efficiency and flexibility directly: efficiency to maintain GPUs fed, adaptability to maintain up with requirements and buyer necessities which might be nonetheless very a lot in movement.
We’ve got demonstrated this functionality a number of occasions throughout actual, production-relevant options:
- Packet Trimming: Quite than dropping packets outright throughout congestion occasions, packet trimming preserves the header whereas discarding the payload, permitting receivers to selectively request retransmission of solely the lacking knowledge. This considerably reduces pointless full-flow retransmissions and improves throughput underneath load—delivered on present Silicon One {hardware} by means of a P4 software program replace, with no silicon adjustments required.
- Full MRC Help: Multipath Dependable Connection introduces a complete suite of load balancing and congestion management mechanisms purpose-built for AI and ML site visitors patterns. As a result of Silicon One is P4-programmable, we have been capable of implement the whole MRC functionality set—together with its multipath load balancing and congestion response algorithms—with out ready for a brand new ASIC era.
- Weighted Adaptive Routing: AI workloads generate extremely bursty, uneven site visitors that may quickly create hotspots throughout a cloth. Weighted Adaptive Routing dynamically distributes flows throughout out there paths based mostly on real-time congestion metrics, assigning weights to steer site visitors away from congested hyperlinks and maximize cloth utilization. Delivering this functionality on present {hardware} requires solely a P4 software program replace.
- Multi-tenant and Multi-job Isolation: Many of the AI clusters, except for foundational mannequin coaching, assist a number of tenants and a number of jobs inside every tenant. Imposing tenant- and job-level isolation insurance policies to forestall cross-communication is a essential service that the community operator should present. As clients migrate from InfiniBand to Ethernet, supporting an environment friendly answer that minimizes configuration and community churn at any time when a tenant and a job are scheduled onto the cluster turns into a key differentiator.
MRC is an efficient illustration of why Cisco’s SRv6 funding pays off right here. Its switch-side necessities — SRv6 uSID forwarding, packet trimming, deterministic path-pinned telemetry — line up with capabilities we’ve already constructed by means of SRv6 and programmable Silicon One forwarding. And since that forwarding conduct is programmable, each these capabilities and customer-specific extensions can maintain evolving {hardware} you’ve already deployed, because the spec matures.
This isn’t a theoretical benefit; it’s the distinction between telling a buyer “we assist that right this moment” and “we’ll have silicon for that in 12 to 18 months.” In AI infrastructure, this distinction is essential.
The broader level is that programmability is important. Given the speedy evolution of AI networking requirements, it’s the solely viable architectural method. Persevering with to construct rigid ASICs to a hard and fast specification and counting on market stability is more and more troublesome to justify as new protocols are launched.
The trail ahead
The way forward for AI relies upon not solely on server silicon but in addition on the material connecting these servers. As we enter the period of huge, multi-rack clusters, the trade wants a strong, versatile networking basis.
That basis comes right down to a single, open constructing block — Ethernet — versatile sufficient to deal with three distinct scaling challenges directly:
- Scale-up, ultra-optimized: inside the rack, Ethernet should match the uncooked, low-latency efficiency of devoted scale-up materials between GPUs.
- Scale-out, performant and dependable: throughout racks and pods, it should maintain full throughput and dependable supply as coaching clusters scale out to tens of 1000’s of GPUs.
- Scale-across, fault-tolerant and QoS-aware: throughout knowledge facilities and geographies, it should protect job isolation and predictable efficiency as 1000’s of GPUs coaching clusters — and more and more, inference clusters — span the huge space community.
As Ethernet evolves, it solves for all three — with out giving up the open, standards-based ecosystem that makes it the suitable long-term alternative for AI infrastructure.
Cisco is dedicated to delivering this basis. By prioritizing open requirements, high-performance silicon, and clever automation, we guarantee tomorrow’s infrastructure can assist right this moment’s breakthroughs.
To be clear, this isn’t Ethernet as a substitute of innovation. It’s Ethernet because the open basis innovation builds on — multiplied by P4 programmability and delivered in platforms like Cisco Silicon One — so AI networks can evolve simply as quick because the workloads driving on them.
Extra assets:
