Microsoft Azure Unveils the Revolutionary NDv6 GB300 VM Series for AI Workloads
Microsoft Azure has recently made headlines with the launch of its new NDv6 GB300 VM series. This groundbreaking development introduces the industry’s first supercomputing-scale production cluster of NVIDIA GB300 NVL72 systems, specifically crafted to handle the most demanding AI inference workloads by OpenAI.
Unmatched Supercomputing Power
The NDv6 GB300 series is a technological marvel, incorporating over 4,600 NVIDIA Blackwell Ultra GPUs linked via the NVIDIA Quantum-X800 InfiniBand networking platform. This remarkable setup demonstrates Microsoft’s commitment to pushing the boundaries of AI workloads, providing an enormous scale of compute necessary for achieving superior inference and training performance. This achievement reflects years of collaboration between NVIDIA and Microsoft, underscoring their commitment to engineering robust AI infrastructures.
Optimizing AI Data Centers
Nidhi Chappell, the corporate vice president of Microsoft Azure AI Infrastructure, emphasizes that the launch of the GB300 NVL72 production cluster is not merely about powerful hardware; it’s about optimizing every component of a modern AI data center. Their partnership focuses on delivering a dependable infrastructure that empowers customers like OpenAI to deploy next-generation AI solutions at unprecedented scales and speeds.
Deep Diving into the NVIDIA GB300 NVL72
Central to Azure’s NDv6 GB300 VM series is the cutting-edge NVIDIA GB300 NVL72 system. Each rack is engineered to house 72 NVIDIA Blackwell Ultra GPUs and 36 NVIDIA Grace CPUs, creating a powerful unit designed to accelerate training and inference for expansive AI models.
This system boasts an astounding 37 terabytes of fast memory and 1.44 exaflops of FP4 Tensor Core performance per VM, offering a unified memory space critical for managing complex reasoning models and agentic AI systems. What sets the NVIDIA Blackwell Ultra apart is its full-stack AI platform, including innovative communication libraries and technologies designed for peak training performance.
Outstanding Performance Metrics
The capabilities of the NVIDIA GB300 NVL72 system are further validated by recent MLPerf Inference v5.1 benchmarks, where these systems achieved unprecedented performance using the NVFP4 format. Notably, they delivered up to 5x higher throughput per GPU for the 671-billion-parameter DeepSeek-R1 reasoning model when compared to the previous NVIDIA Hopper architecture.
The Innovative Networking Fabric
Connecting over 4,600 Blackwell Ultra GPUs into a singular supercomputing entity wasn’t a simple feat. Microsoft Azure relies on a sophisticated two-tiered NVIDIA networking architecture that enhances both scale-up performance within the individual racks and scale-out performance across the entire cluster.
Within each GB300 NVL72 rack, the fifth-generation NVIDIA NVLink Switch fabric enables an impressive 130 TB/s of direct, all-to-all bandwidth. This design transforms each rack into a unified accelerator, providing a shared memory pool crucial for handling large, memory-intensive AI models.
Scaling beyond the individual racks requires leveraging the advanced capabilities of the NVIDIA Quantum-X800 InfiniBand platform. Designed specifically for trillion-parameter-scale AI, it offers remarkable bandwidth of 800 Gb/s per GPU, ensuring seamless communication across all 4,608 GPUs.
Advanced Features for Maximum Efficiency
The architectural finesse of Azure’s cluster extends beyond its raw capability; it incorporates advanced features like NVIDIA Quantum-X800’s adaptive routing and congestion control. These elements work in synergy with NVIDIA’s Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) v4, significantly enhancing the efficiency of large-scale training and inference operations.
Pioneering Future AI Innovations
Building the world’s first large-scale production NVIDIA GB300 NVL72 cluster pushed the limits of data center design. From custom liquid cooling systems to reengineered power distribution and software stacks, every layer was reimagined to meet the demands of this supercomputing power.
The NDv6 GB300 VM series marks a pivotal moment in AI infrastructure, with Azure aiming to deploy hundreds of thousands of NVIDIA Blackwell Ultra GPUs. This innovative landscape promises to be the breeding ground for future breakthroughs in AI technologies, setting the stage for more advanced solutions.
For further insights and developments, you can explore this exciting announcement in detail on the Microsoft Azure blog.
Inspired by: Source

