AMD announced that its first rack-scale AI system, Helios, will be supplied to Microsoft Azure to drive frontier model AI inference workloads, while also supporting Azure AI services and customer applications. Helios is expected to cost between $5 million and $5.5 million, and AMD's stock rose 8% on the day. Meta and OpenAI have confirmed customer support, and eight of the top ten global AI companies run workloads on its Instinct GPUs.
Fact Restoration
On July 20, 2026, AMD and Microsoft expanded their long-term partnership, with Microsoft becoming the first hyperscaler to publicly commit to large-scale deployment of Helios. Helios integrates AMD Instinct MI455X GPUs, sixth-generation EPYC Venice CPUs, Pensando networking technology, and the ROCm software stack. Each compute tray is equipped with four Instinct GPUs driven by a single EPYC CPU, along with up to 12 networking chips based on Pensando technology. AMD plans to begin shipping to customers, including Microsoft, in the second half of 2026. Microsoft Azure will add two new virtual machine series based on Venice CPUs: Azure HDv2 targets agentic AI and data pipeline workloads, while Azure HXv2 targets semiconductor design use cases.
Mechanism Breakdown
Helios, as AMD's first rack-scale system, adopts a full-stack proprietary hardware architecture with the goal of delivering the best total cost of ownership (TCO) and the lowest per-token cost. AMD is responsible for the silicon foundation, while Microsoft packages these components within Azure cloud operations, developer services, security, and AI platform services. The partnership extends their collaboration on Surface PCs, Xbox, and the 2023 MI300X GPU to now cover the full spectrum of GPUs, CPUs, networking, and software. Microsoft CEO Satya Nadella stated that this move provides customers with the performance, scale, and choice needed to build and run next-generation AI applications. AMD's data center business lead, Forrest Norrod, emphasized a system-level approach rather than selling individual components, covering training, inference, fine-tuning, and agentic workloads.
Industry Impact
In terms of the competitive landscape, Helios directly competes with NVIDIA's Grace Blackwell and Vera Rubin systems. NVIDIA currently holds over 95% of the data center GPU market share, while AMD accounts for approximately 4.5%. If AMD's share rises to 20% to 25%, it would correspond to hundreds of billions of dollars in revenue. On the supply chain side, AMD's data center business revenue grew 57% year-over-year in the first quarter of 2026, and the company plans to achieve tens of billions of dollars in data center AI revenue starting in 2027, with the majority coming from Helios. Developers can use Azure Foundry Managed Compute to host production AI workloads, while enterprise users gain an additional infrastructure option beyond NVIDIA. Meta has committed to using up to 6 gigawatts of AMD GPUs over time and has already deployed 1 gigawatt on Helios racks; OpenAI, Oracle, and Tata Consultancy Services have also made significant deployment commitments.
Comparison and Precedent
Compared to NVIDIA's second-generation rack-scale system Vera Rubin, Futurum Group estimates Helios costs between $5 million and $5.5 million, higher than Vera Rubin's $3.5 million to $4 million. AMD CEO Lisa Su previously stated that Helios has significant advantages over NVIDIA's rack systems in inference performance, memory bandwidth, and memory capacity. Counterpoint Research analyst Neil Shah noted that AMD Helios chip performance is comparable to NVIDIA's GPUs and CPUs, but the key lies in software and optimization, with the CUDA ecosystem still being a maturity gap.
Strategic Judgment
The early deployment performance of Helios will determine whether it can transition from an alternative under capacity shortages to a technology leader. Signals to watch include actual per-token inference cost data after shipping in the second half of 2026, and the utilization rates of Azure HDv2 and HXv2 instances. AMD is likely to achieve its data center AI revenue target by 2027, but the software ecosystem gap remains the biggest variable.
When selecting technology, developers should prioritize testing ROCm-optimized workloads on Azure with Helios support, focusing on compatibility with existing CUDA code and total cost of ownership. Enterprise users should compare Helios and NVIDIA systems based on per-token cost relative to their inference scale, while also monitoring the actual performance of Microsoft's new Venice CPU instances in agentic AI scenarios. Deployment feedback from Meta and OpenAI will serve as important reference indicators.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接