As AI models reach unprecedented sizes, xPU cluster performance relies heavily on the scale-up network that tightly coordinates thousands of accelerators for distributed training workloads. Scaling up the xPU domain intensifies intra-domain communication demands, requiring higher bandwidth and lower latency than scale-out fabrics can provide. The relentless scaling of these clusters has pushed traditional copper interconnects to their physical limits, creating significant bottlenecks in the scale-up domain. To address the immense I/O demands of tightly coupled AI accelerators, the industry is exploring innovative Optical PHY architectures, from traditional "narrow-and-fast" approaches to emerging "wide-and-slow" interfaces as defined by the OCI MSA. Advancing rack-level optical connectivity through technologies such as expanded-beam connectors and optical backplanes is equally essential to realize true resource disaggregation. This workshop examines the key technologies required to break free from copper, focusing on advanced Optical PHY designs and rack-scale optical integration shaping the next generation of AI infrastructure.
Organizers
-
Siamak Amiralizadeh
Meta Platforms Inc.,, United States
-
Paraskevas Bakopoulos
NVIDIA Corporation, Greece
-
Guangwei Cong
National Inst. Of Advanced Industrial Science and Technology, Japan
-
Yuyang Wang
Univ. of Connecticut, United States
-
Scott Yin
Ghent University, Belgium