dayliyreport

Search

AI

Meta and Oracle adopt NVIDIA Spectrum-X for AI Data Centers

·5 min read
Advertisement

Meta and Oracle are significantly enhancing their artificial intelligence data centers by integrating NVIDIA's cutting-edge Spectrum-X Ethernet networking switches. This strategic adoption is poised to revolutionize the efficiency of AI training and accelerate the deployment of large-scale AI systems across massive computing clusters. The move highlights a growing industry need for robust and scalable networking solutions capable of handling the immense computational demands of advanced AI models.

NVIDIA's CEO, Jensen Huang, underscored the transformative power of trillion-parameter models, describing modern data centers as "giga-scale AI factories." He likened Spectrum-X to the "nervous system" that seamlessly connects millions of graphics processing units (GPUs), enabling the development and training of the most complex AI models to date.

Oracle's plan involves leveraging Spectrum-X Ethernet in conjunction with its Vera Rubin architecture to construct sophisticated AI factories. Mahesh Thiagarajan, Executive Vice President at Oracle Cloud Infrastructure, explained that this new configuration will facilitate more efficient connections among millions of GPUs. This enhancement will ultimately enable customers to train and deploy novel AI models at an accelerated pace, providing a significant competitive advantage.

Similarly, Meta is expanding its AI infrastructure by incorporating Spectrum-X Ethernet switches into its proprietary Facebook Open Switching System (FBOSS). Gaya Nagarajan, Meta's Vice President of Networking Engineering, stressed the importance of an open and efficient next-generation network. Such a network is crucial for supporting increasingly larger AI models and ensuring the delivery of services to billions of global users.

The emphasis on building flexible AI systems is a key driver behind these collaborations. Joe DeLaere, who oversees NVIDIA's Accelerated Computing Solution Portfolio for Data Centre, highlighted that flexibility is paramount as data center complexity escalates. He noted that NVIDIA's MGX system provides a modular design, allowing partners to customize combinations of CPUs, GPUs, storage, and networking components. This modularity also fosters interoperability, ensuring that the same design can be utilized across various hardware generations, thus offering flexibility, faster market entry, and readiness for future advancements.

Addressing the critical challenge of power efficiency in large-scale AI operations, NVIDIA is pursuing a comprehensive "from chip to grid" strategy. This involves close collaboration with power and cooling providers to maximize performance per watt. Innovations include a shift to 800-volt DC power delivery, which minimizes heat loss, and the introduction of power-smoothing technology to mitigate electrical grid spikes, potentially reducing peak power needs by up to 30 percent and allowing for greater compute density within existing infrastructure. NVIDIA's upcoming Vera Rubin architecture, set for commercial release in late 2026, will integrate seamlessly with Spectrum-X networking and MGX systems to power the next generation of AI factories.

NVIDIA's MGX system plays a pivotal role in data center scalability, supporting both scale-up connectivity with NVLink and scale-out growth with Spectrum-X Ethernet. Gilad Shainer, the company's Senior Vice President of Networking, pointed out that MGX can unify multiple AI data centers into a single, cohesive system, which is essential for massive distributed AI training operations. This capability allows for high-speed connections across regions, enabling companies like Meta to link sites via dark fiber or additional MGX-based switches.

The adoption of Spectrum-X by Meta underscores the growing importance of open networking. While Meta will use FBOSS as its network operating system, Spectrum-X supports other operating systems like Cumulus, SONiC, and Cisco's NOS through various partnerships. This broad compatibility offers hyperscalers and enterprises the flexibility to standardize their infrastructure using systems best suited to their specific environments, expanding the overall AI ecosystem.

Spectrum-X Ethernet is uniquely engineered for AI workloads, such as training and inference, delivering up to 95 percent effective bandwidth—a significant improvement over traditional Ethernet. This specialized design, combined with adaptive routing and telemetry-based congestion control, eliminates network bottlenecks and ensures stable, high-performance operation. These features facilitate faster training and inference speeds while allowing multiple workloads to run concurrently without interference, maximizing the return on GPU investments for hyperscalers like Meta.

Related Articles