Teresa A Nick, Philippe Lewicki, Jeffrey Breugelmans, Shane Drexler, Imre Knausz, Nikhil Jain, Hakki Refai, Husam Alissa
AI workloads demand increasing energy, scalability and transparency from data center infrastructure. Modern distributed AI models use clusters of processors to perform collective matrix operations based on 1940s pairwise designs, which remain foundational in tensor cores and scale in complexity with the number of processors. To eliminate this bottleneck in distributed AI clusters, here we present Parallel Photonic Integration (PPI), a collective computation architecture purpose-built for hyperscale AI. PPI replaces sequential electronic operations with fully parallel, photonic-simulcast computation, which enables multi-node matrix operations while also delivering real-time model transparency without added cluster load. Replacing network switches enables PPI to integrate into existing AI frameworks and perform light-based multi-matrix calculations, as shown in our early-stage proof-of-concept. Predictive modeling of AllReduce across large GPU clusters running Transformer workloads shows that PPI can reduce per-operation energy by over 50% compared with switched-fabric interconnects, while reducing AllReduce latency by over 100x at frontier scale. These results demonstrate that PPI can provide a scalable, energy-efficient, and transparent foundation for next-generation AI.