AI Data Center Networking Stack | Switches, Coherent Optics & SONiC Fabric | EPS Global
For Neoclouds, Sovereign Data Centers & GPU Cluster Operators

Networking for AI factories. Delivered.

EPS Global assembles the complete networking stack for AI: optics, switches, fabric software, intelligence, cabling, power and cooling into one validated solution. Anchored by 25 years of Coherent optical distribution. Deployed from 28 global locations.

Book a Technical Consultation

A pre-sales engineer is assigned within one business day.

By submitting, you agree to be contacted by EPS Global about your inquiry. We will not share your details with third parties.

25yrs
Coherent authorized partnership
28
Global stocking locations
14
Validated technology partners
1
Single source of procurement
The Problem

AI factories don't stall on GPUs. They stall on the network.

A neocloud building an AI factory has to assemble power, cooling, structured cabling, optical interconnect, open switches, AI fabric software, cluster management, and network intelligence — each from a different specialist.

"Integration debt" — months managing 10+ vendors before the first GPU collective even runs.

Each vendor optimizes for their own layer. None integrate the whole. The result: mismatched delivery schedules, validation rework, and finger-pointing when something doesn't talk to something else.

Challenges you might be facing:

  • 10+ separate vendor relationships per build phase
  • Optical supply chain risk at neocloud volumes
  • Fabric OS lock-in or DIY SONiC integration
  • No software intelligence above the fabric
  • Power, cooling, OOB out of scope
  • Cabling design and labelling done in-house
  • Onsite install at five different timelines
The Solution: Validated Stack from EPS Global

The Full Ecosystem. One commercial relationship.

Every layer of the sovereign data center networking stack from in-rack DAC and AECs to network intelligence - supplied, validated, and delivered through a single EPS Global engagement.

Broadcom
Broadcom SONiC
Coherent
Optics & AOC · 25-yr partnership
Celestica
Network switching
Maia Edge
AI cluster management
Aria Networks
Network intelligence
Hedgehog
SONiC fabric
Aviz Networks
SONiC fabric
Amphenol
DAC & structured cabling
Vertiv
Power, cooling, OOB
Proficium
Design & install services
Coherent
Optics & AOC · 25-yr partnership
Celestica
Network switching
Edgecore
Network switching
UfiSpace
Network switching
Maia Edge
AI cluster management
Aria Networks
Network intelligence
Hedgehog
SONiC fabric
Aviz Networks
Aviz Networks
Amphenol
DAC & structured cabling
Vertiv
Power, cooling, OOB
Proficium
Design & install services
The Architecture

Multiple layers, validated end-to-end.

Hover or tap any layer to explore the partner technologies, reference specs, and 1,024-GPU POD quantities behind each tier of the stack.

Loading interactive architecture diagram...
Reference Build

What a 1,024-GPU POD looks like in BoM.

Indicative quantities for a representative POD on the EPS Global stack — three-tier Clos topology, 2:1 spine, single-source procurement.

52
Network switches
4 spine · 16 leaf · 32 ToR
~700
Coherent optical modules
400G / 800G / 1.6T
~256
Coherent AOC
Inter-rack runs
~1,024
Amphenol DAC
In-rack server-to-ToR
~950–1,000 Coherent optics and AOCs (Active Optical Cables) per POD — the single largest line in the BoM by unit count, reflecting why a 25-year authorized Coherent partnership is the anchor of this stack. Plus 2× Vertiv rPDUs per rack, in-rack CDU 121 cooling, and Avocent ACS8000 OOB management.
Network Architecture

Scale Up, Scale Out, Scale Across.

AI workloads don't scale linearly. They require distinct networking strategies at the node, cluster, and multi-site levels. Here is how our validated stack supports your growth without bottlenecks.

Fig. 04 — Scale out · the POD fabric02 / 03
ONE RACK — 32 GPUs OF 1,024 400G UPLINKS → LEAF TIER (02) ToR — DS4000 · 32 × 400G GPU SERVER · 4 × GPUGPU SERVER · 4 × GPUGPU SERVER · 4 × GPUGPU SERVER · 4 × GPUGPU SERVER · 4 × GPUGPU SERVER · 4 × GPUGPU SERVER · 4 × GPUGPU SERVER · 4 × GPU SCALE-UP DOMAIN LOSSLESS SERVER ↔ ToR DAC ≤3m — AMPHENOL × 32 PER RACK VERTIV — rPDU × 2 · CDU 121 100 kW+ · LIQUID LOOP POWER & THERMALS SPECIFIED WITH THE NETWORK — NOT AFTER IT SPINE — 4 × DS5000 · 64 × 800G SP-01SP-02SP-03SP-04 800G OSFP LEAF — 16 × 800G (6 SHOWN) 400G / AOC TOR — 32 × DS4000 (6 SHOWN) DAC 1,024 GPUs · 32 RACKS · 2:1 SPINE ▢ THE RACK FROM 01 EVERY COLLECTIVE WAITS FOR ITS SLOWEST PATH — THE RED ONE SETS STEP TIME SITE A1,024-GPU POD SITE B1,024-GPU POD SITE C1,024-GPU POD ▢ THE POD FROM 02 ONE TRAINING DOMAIN — SYNCHRONOUS ACROSS SITES COHERENT DCI — 400G / 800G ZR / ZR+ · METRO & REGIONAL ARIA — TRAFFIC ENG · UNIFIED TELEMETRY
Three-tier Clos, full bisection — select a scale below to move the camera.

Scale Up: Intra-Rack Density

Maximizing GPU compute density within the individual rack. As GPU power requirements surge, the network must deliver ultra-low latency from the server to the Top-of-Rack (ToR) switch.

  • High-Speed Interconnects: Lossless, short-reach copper and optical connectivity for server-to-ToR links.
  • High-Radix Switching: Maximizing port density per rack unit to support dense GPU configurations.
  • Advanced Infrastructure: High-capacity power distribution and liquid cooling to support next-gen silicon thermals.

Scale Out: The AI Cluster Fabric

Expanding to 1,024+ GPUs requires a lossless, non-blocking Clos topology. This is where traditional networks fail and AI-specific Ethernet fabrics take over.

  • Next-Gen Optics: 800G and 1.6T transceivers leveraging advanced photonics for massive bandwidth.
  • Lossless Ethernet: Open networking leaf/spine architectures running optimized fabric operating systems.
  • Workload Optimization: Tuning RoCEv2 / RDMA traffic to accelerate GPU collective operations.

Scale Across: Distributed AI

Connecting multiple sovereign data centers to act as a unified, distributed AI factory. Essential for massive training runs that exceed the power capacity of a single facility.

  • Data Center Interconnect (DCI): High-capacity, long-haul optical transport for metro and regional links.
  • Network Intelligence: SDN-driven traffic engineering and capacity planning across distributed sites.
  • Unified Management: Seamless orchestration and telemetry across geographically dispersed GPU clusters.
Training Infrastructure

The fabric sets job completion time.

Training traffic looks nothing like a general-purpose data center — and the network, not the GPU, decides how fast a job finishes.

Collectives like all-reduce and all-to-all are synchronous: thousands of GPUs exchange gradients in lock-step, and the entire step waits for the last packet on the slowest path. East-west traffic dominates, flows are few and enormous, and a single ECMP hash collision or congested uplink shows up directly as idle GPU time across the whole job.

That means tail latency, not average latency, is the number that matters — and packet loss is unacceptable, because a retransmit stalls the collective. At training scale, the network is a scheduling problem you solve in hardware, topology and fabric tuning before the first job runs.

EPS Global designs, supplies and validates against these requirements as a system: switching, optics, cabling, fabric OS and the physical layer underneath it, proven together before they ship.

What a training fabric must deliver

  • Non-blockingThree-tier Clos with spine capacity scoped to your oversubscription target
  • LosslessRoCEv2 with PFC and ECN tuned for GPU collectives, so RDMA never sees a drop
  • 800G spineHigh-radix Tomahawk 5 switching with the buffer depth large flows demand
  • Qualified opticsTransceivers and AOC validated against the specific ASIC and NIC combination
  • TelemetryPer-flow visibility to find the congested link before it shows up in step time
  • Power & thermalsRack power and liquid cooling specified with the network, not after it
Inference Infrastructure

Built for the demands of inference

Inference workloads are latency-sensitive, high-concurrency and multi-tenant, and you run them close to your users. For an inference business, the network is cost-per-token: idle GPUs and dropped packets are capacity you've paid for and can't bill. EPS Global builds and validates the stack that keeps that capacity earning.

Right-sized PODs

From sub-256-GPU clusters to multi-tenant serving racks, the bill of materials matches the deployment. We size the same eight validated layers to the workload in front of you, so your capital buys serving capacity.

Low-latency, multi-tenant fabric

Lossless Ethernet with tenant isolation for high-concurrency token serving. Many customers share one fabric at predictable latency, the operating model GPU-as-a-Service depends on. Your GPUs stay saturated and cost-per-token stays predictable.

Deploy close to demand

Run inference close to its users. EPS Global sources and stocks the same stack in-region across 28 locations, so capacity sits next to the people it serves. Shorter paths cut latency and egress, improving your unit economics region by region.

Inference is a regional workload. With stock in 28 locations, you can deploy wherever sovereignty requires.

Talk to a pre-sales engineer →
EPS Global SwitchLab

Your fabric runs on our bench before it runs your jobs.

Celestica DS5000 800G switch on the bench at the EPS Global SwitchLab, Indianapolis
EPS Global SwitchLab, Indianapolis — Celestica DS5000 on the bench.

An AI fabric has no tolerance for surprises at install. In the EPS Global SwitchLab, engineers burn in switching hardware, install and configure the SONiC fabric OS, and validate optics and cable assemblies against the specific ASIC and NIC combination in your design — before anything ships. Racks-ready equipment arrives with labelled cabling and powers up the way it left the bench.

The optical layer gets particular attention because it's the largest line in the BoM: twenty-five years of authorized Coherent distribution means supply depth at neocloud volumes, and qualification discipline at 800G, where insertion-loss budgets leave no slack.

Optical anchor depth

Volume and supply chain depth at the layer where a training POD consumes nearly a thousand units.

Validated interoperability

Switching, optics and SONiC proven together in the lab — not assembled for the first time on your floor.

Choice, not lock-in

Hedgehog or Aviz on the same hardware; the fabric OS decision never strands the switching investment.

One engagement

Vertiv power and cooling, Amphenol cabling and Proficium install scoped in the same BoM and the same PO.

The EPS Global Edge

What a hardware-only distributor can't do for you.

EPS Global is a value-added distributor — not a box-mover. The neocloud stack is assembled, validated, and delivered as a single engagement.

Optical anchor depth

Twenty-five years of authorized Coherent distribution. Volume pricing and supply chain depth that general distributors cannot match — critical when optics is ~700 units per POD.

Strategic software layer

Maia Edge and Aria Networks are new EPS Global relationships at the AI cluster management and network intelligence layers — differentiation no hardware-only VAD can replicate.

Choice, not lock-in

Hedgehog or Aviz Networks for the SONiC fabric OS — both validated on Celestica hardware. EPS Global engineering advises on the right fit for your operations model.

Complete physical infrastructure

Vertiv closes power, cooling, rack enclosures, and OOB management within the same EPS Global engagement. Amphenol covers DAC and structured fibre.

Professional services wrapped in

Proficium provides infrastructure design, custom cable manufacturing, and global install. Their BoM flows directly into EPS Global procurement.

Global fulfilment

28 EPS Global stocking locations across Europe, North America, and Asia Pacific support initial POD builds and ongoing replenishment of optics, AOC, PDUs, and cabling.

From the Lab & the Field

Unboxings, vendor interviews, and architectural deep-dives.

Hardware in the rack, optics on the bench, and conversations with the engineers who build the silicon.

Julie Eng, Coherent CTO, on Co-Packaged Optics and pluggable transceivers
Vendor Interview

Optics in the AI Era — Coherent's CTO on the Future of CPO and Pluggables

Coherent's CTO Julie Eng breaks down Co-Packaged Optics, the future of pluggable transceivers, and what AI-driven demand means for next-generation data center networking.

Supercharge Your AI Data Center: 1.6T OSFP Coherent Transceiver Unboxed
Unboxing

Supercharge Your AI Data Center: 1.6T OSFP Coherent Transceiver Unboxed

EPS Global Systems Engineer Wellyson Mota gives us an exclusive look inside the switch lab in Dublin to demonstrate this state-of-the-art 1.6T OSFP transceiver.

Neoclouds, AI infrastructure and white box networking — Celestica and EPS Global
Industry Deep-Dive

Neoclouds, AI Infrastructure & the Rise of White Box Networking — Celestica and EPS Global

How Celestica is powering the world's largest hyperscalers and neoclouds, as industry experts unpack the hardware, software, and open networking ecosystems driving today's AI infrastructure revolution.

Celestica DS5000 800G data center switch unboxing
Unboxing

Unboxing Celestica's High-performance 800G Data Center Switch — DS5000

Explore the Celestica DS5000 800G switch — a 64-port, 51.2 Tbps powerhouse built on Broadcom's Tomahawk 5 ASIC.

Alan Fagan and Matt Free from Aria Networks
Podcast

Building the AI Factory: Solving the GPU Bottleneck with "Deep Networking"

The AI infrastructure market is exploding, but legacy networks are quietly bottlenecking billion-dollar GPU clusters. How do you stop wasting compute and turn your network into a revenue multiplier?

Edgecore AIS800-64D 800G AI/ML networking switch unboxing
Unboxing

Next-Gen 800G Networking for AI & ML — Unboxing the Edgecore AIS800-64D

A deep-dive of the AIS800-64D, exploring how its 51.2 Tbps Tomahawk 5 architecture and flexible 800G port configurations accelerate AI and ML workloads at scale.

Marc Austin, Hedgehog CEO, on building AI networks at half the cost
Vendor Interview

Build Better AI Networks for Half the Cost — In Conversation with Hedgehog CEO Marc Austin

Learn how open networking and disaggregated hardware can deliver superior AI infrastructure performance at half the cost of proprietary solutions.

Seth Haller, EPS Global Director of Sales, on Supply Chain as a Service
EPS Global Insight

"Supply Chain as a Service" — Seth Haller on EPS Global's Philosophy for Customer Success

Seth Haller, Director of Sales for Western US at EPS Global, shares insights from his 36-year career on solving the persistent supply chain problem.

Frequently Asked Questions

Common questions from sovereign data center operators.

A GPU cluster at this scale requires a non-blocking spine-leaf fabric built on high-radix open Ethernet switches (e.g. Celestica DS3000/SN5600 class), 800G or 1.6T optics at the spine tier, 400G at the leaf, 100G–400G DAC or AOC at the leaf-to-NIC edge, and a SONiC-based fabric OS for automation and telemetry. Above the fabric you need an AI cluster management layer to schedule workloads across GPUs and track utilisation in real time.
Buyers commonly evaluate Celestica, Edgecore, and Delta white-box switches running SONiC for AI fabric builds. Celestica's DN-series and SN-series platforms offer the port density, buffer depth, and ASIC choice (Broadcom, Marvell) suited to 400G, 800G, and 1.6T spine-leaf at GPU-cluster scale, and are available through distribution with validated SONiC interoperability tested against Hedgehog and Aviz Networks fabric OS builds.
Both run on the same Celestica hardware. Aviz SONiC suits GPU-as-a-Service operators who need rich multi-tenant orchestration and day-2 operations tooling. Hedgehog is purpose-built for open networking automation in tightly coupled AI training clusters. The right choice depends on whether you're running dedicated training infrastructure or a shared GPU cloud — pre-sales engineering can scope this against your operational model.
For in-rack and top-of-rack connections (≤100 m), 100G–400G VCSEL-based direct-detect transceivers (SR4, SR8) are standard. For leaf-to-spine links, 400G or 800G OSFP/QSFP-DD is the current design point; 1.6T (2×800G) ports are now available on leading spine ASICs and should be factored into any build with a 3–5 year horizon. For inter-rack, end-of-row, and any campus or DCI links over 500 m, coherent optics — 400G/800G ZR or ZR+ — provide the reach and density required. Coherent transceivers typically represent the largest single line item in a GPU POD bill of materials, so distributor stocking depth and lead time matter significantly.
Short-reach switch-to-switch and switch-to-NIC connections are typically served by 400G or 800G DAC up to 3–5 m or Active Optical Cable (AOC) up to 30–100 m. At 800G and above, connector and insertion-loss budgets are tighter — cable assembly qualification against the specific ASIC and transceiver combination matters. Structured cabling with pre-labelled MPO trunk assemblies is used for longer intra-cage runs and future-proof patching. Cabling is routinely under-specified at planning stage — cable type, reach, polarity, and labelling scheme should be locked in alongside the switch and optics BoM.
Active Electrical Cables (AECs) are narrow-gauge copper cables with active signal-conditioning silicon — retimers, and FEC where needed — built into the connector at each end. That active electronics restores signal integrity across the run, letting copper hold its low power, light weight, and tight bend radius at 112G-per-lane speeds, where passive DAC runs out of reach at roughly 2–3 m. A typical AEC carries 400G or 800G up to ~7 m with far less bulk than the equivalent DAC and lower power than an AOC. In an AI fabric they fill the gap between passive copper and optics — server-to-ToR and ToR-to-leaf runs too long for DAC but too short to justify transceivers — and a stable copper link that stays up keeps RDMA flows lossless and GPUs saturated. EPS Global scopes AEC versus DAC versus AOC against the switch, NIC, and reach map in your design.
A 1,024-GPU cluster at H100/H200 class typically demands 30–60 kW per rack depending on GPU density, requiring high-density PDUs (3-phase, 22–44 kW rated), precision cooling (in-row or rear-door), and out-of-band management for every rack. Power and cooling should be specified concurrently with the networking BoM — vendors like Vertiv supply integrated rack power, cooling, and OOB that are validated against the switch and cabling stack.

Ready to spec your sovereign data center networking stack?

A pre-sales engineer is assigned within one business day of consultation request.

Need Help?

We have local language and currency support in each of our 28 locations, ensuring you always have access to friendly customer support to deliver your hardware solutions regardless of your location.