Session brief
How to Build Your Own AI Data Center in 2025 — Paul Gilbert, Arista Networks
Paul Gilbert , Arista Networks
Overview
This talk focuses on the infrastructure required to build and operate AI data centers, specifically addressing the networking challenges and solutions for training and inference. It highlights the significant differences between traditional data center networking and the demands of AI workloads, emphasizing the need for high bandwidth, low latency, and specialized traffic management. The core thesis is that building effective AI data centers requires a fundamental shift in network design and implementation to handle the unique, high-intensity demands of GPUs.
Who should watch
- AI Engineers
- Network Engineers
- Infrastructure Architects
- Product Managers involved in AI deployments
- Builders facing challenges with AI model training and inference performance
- Those needing to understand the hardware and networking requirements for large-scale AI
Key takeaways
- AI networks require dedicated, isolated infrastructure with extremely high bandwidth, often operating at 400 Gbps and moving towards 800 Gbps and beyond, with no oversubscription on the backend network.
- GPU communication is highly synchronized and bursty, demanding specialized network protocols like RoCE v2 (PFC and ECN) for congestion control and lossless data transfer to maintain job completion times.
- Power and cooling requirements for AI servers are significantly higher than traditional servers, necessitating new rack designs (100-200 KW) and water cooling solutions.
- Traffic patterns in AI networks are predominantly east-west (GPU-to-GPU communication), unlike traditional north-south traffic, requiring careful load balancing strategies that consider bandwidth utilization rather than just IP/port tuples.
- Visibility and telemetry are crucial for diagnosing issues, with tools like AI agents on GPUs communicating with switches to correlate network and GPU performance problems.
- Network upgrades need to be non-disruptive, with capabilities like smart system upgrades allowing code updates without taking switches offline.
- The Ultra Ethernet Consortium is developing new standards to enhance Ethernet capabilities for AI workloads, focusing on offloading more intelligence to NICs.
Notable quotes
*The backend network depending on the model that you train the gpus will actually work at 400 GB.*
*We have no over subscription in the network and from from our point of view if you look at what one of these servers can put on the network you know just a h100 is 8 400 gig gpus and 4 400 gig is 4.8 terabytes which is and that's just one server.*
*If one of the big problems that's we've always add is Optics and transceivers and Doms which is the rates and the loss between them and the cables Etc and when you start building these networks with thousands of gpus you will have a lot of cable problems and you will have a lot of GPU problems.*
Unofficial community note. Prefer the recording for nuance.