On August 26, 2026, AWS and NVIDIA jointly announced that they will deploy an additional 2 million NVIDIA GPUs across AWS's global infrastructure (including AI factories) between 2027 and 2028, spanning three generations of architecture — Blackwell Ultra, Rubin, and Rubin Ultra — for workloads including agentic AI, scientific computing, enterprise automation, and physical AI. According to NVIDIA's official newsroom, this is another major upgrade following the announcement at NVIDIA GTC 2026 in March of this year of over 1 million additional GPUs starting in 2026 — the two plans are separated by only a few months, with the increment doubled and the time window extended by two years.
Both companies were direct about the immediate trigger for the upgrade: actual customer demand has exceeded all previous forecasts. NVIDIA CEO Jensen Huang said in the joint statement that AI compute demand "has exceeded every prediction," and characterized the expansion as a full-stack upgrade extending the partnership from GPUs to "CPUs, networking, open models, and software." Looking at the schedule, the first batch of 1 million GPUs has a deployment window of 2026, while the second batch of 2 million extends to 2027–2028. Absorbing this volume within three years is a considerable commitment for both the supply chain and engineering deployment.
From "Buying Chips" to "Co-Building an Infrastructure System"
The substantive change in this partnership is not just the GPU count, but the reconfiguration of the entire technology stack. According to the joint announcement, the scope of collaboration has expanded to include CPUs (NVIDIA Vera), networking (NVIDIA Spectrum), custom high-bandwidth memory (NVHBM), open models (NVIDIA Nemotron), and robotics platforms (Jetson, Omniverse, Isaac). This is a paradigm shift from "purchasing accelerator chips" to "co-designing an AI infrastructure system," covering the full depth from underlying silicon to upper-layer software models.
At the architecture integration level, the next-generation Trainium chips from Annapurna Labs, AWS's chip design subsidiary, will support the NVIDIA NVLink Fusion interconnect protocol and connect to NVIDIA's custom high-bandwidth memory NVHBM and scale-up architecture, enabling mixed deployment of AWS's self-developed AI chips and NVIDIA GPUs within the same rack-level system.
This design breaks the binary opposition between the "self-developed chip route" and the "externally purchased GPU route." Trainium handles specific training tasks, while NVIDIA GPUs take on inference and general-purpose computing, with both interconnected at high speed via NVLink Fusion and sharing a high-bandwidth memory pool. For large AI training clusters, this heterogeneous mix means compute can be allocated based on workload characteristics rather than forcing all tasks through a single chip type. This reflects a realistic industry assessment: for top-tier AI workloads, the switching costs accumulated by NVIDIA's software ecosystem (CUDA, NIM, Nemotron, etc.) are sufficient to make even cloud providers with self-developed chip capabilities choose deeper integration over substitution.
Quantifiable Performance Improvements
The joint announcement disclosed several sets of specific performance data. Amazon EC2 G7 instances powered by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs deliver 4.6x higher AI inference performance and 2.1x higher graphics performance compared to the previous-generation G6 instances. At the data processing level, NVIDIA cuDF introduced to Amazon EMR delivers processing speeds up to 3.7x faster than pure CPU configurations, with a 30% improvement in price-performance ratio. Amazon OpenSearch's GPU vector indexing speed has improved by up to 9x, with costs reduced to approximately one-quarter of the original.
Behind these three sets of numbers lies a common logic: GPU acceleration is penetrating from model training and inference into stages previously dominated by CPUs, such as data pipelines and index construction. Looking at the distribution of compute consumption, future AI systems' GPU dependency will exist not only in the "model running" stage, but will also extend across the entire data preparation and retrieval pipeline.
Government AI Factories: Compute Enters the Highest Security Tier
The partnership also includes a separate government business line: the two companies plan to build AI factories for U.S. federal government and national security workloads, deploying 100,000 NVIDIA GPUs on AWS secure infrastructure to support applications at Impact Level 6 (IL6) and higher security levels.
IL6 is the U.S. Department of Defense's baseline security classification for classified information processing. Being able to deploy NVIDIA Blackwell-class GPUs at this level means AI inference capabilities are formally entering government systems that require the highest level of information security controls. This is a clear milestone in commercial cloud computing's penetration into the government security domain, and will serve as a benchmark for future government AI procurement.
Physical AI: From Data Centers to Warehouse Robots
Physical AI is another new direction in this partnership. Amazon Robotics is adopting NVIDIA Jetson (edge computing modules), Omniverse (physics simulation platform), and Isaac (robot development framework) for robot simulation, synthetic data generation, training, path optimization, and validation, advancing Amazon's warehouse automation and next-generation robotics R&D.
This directly links Amazon's own logistics operations to the compute expansion. AWS is both the seller and the consumer of compute — the large-scale deployment of the warehouse robot network itself constitutes internal demand supporting this expansion. The commercialization of physical AI, from data center training to warehouse floor deployment, is forming a complete vertical chain within the AWS-NVIDIA partnership framework.
Ripple Effects on the Industry Landscape
For AI developers and enterprise users, the direct implication of this expansion is that wait times and supply constraints for accessing NVIDIA's latest-generation GPU compute on AWS are expected to ease during 2027–2028. Cloud quotas for Blackwell-series GPUs are currently generally tight across major cloud providers; if the 2 million GPU addition plan is delivered on schedule, it will substantially fill this supply-demand gap.
On the software front, NVIDIA Nemotron open models will continue to be available through Amazon Bedrock and Amazon SageMaker, allowing enterprise users to use these models through AWS standard invocation interfaces without having to maintain inference infrastructure themselves. For enterprise AI teams moving from pilots to large-scale deployment, this lowers the barrier to using top-tier open models in a managed environment.
AWS CEO Matt Garman said in the statement that customers want the freedom to choose the tools best suited to their AI workloads while ensuring all technologies work seamlessly together, "which is exactly why we are working deeply with NVIDIA to make AWS the best platform for running NVIDIA AI technology, optimizing infrastructure performance across networking, security, and deployment."
The Trajectory of a 16-Year Partnership and This Structural Shift
Jensen Huang positioned this expansion as a continuation of a 16-year partnership. Looking at the timeline, AWS and NVIDIA started with GPU virtualization and early EC2 GPU instances, progressed through large-scale Volta and Turing deployments during the deep learning boom, and have now reached today's full-stack integration of Blackwell and Rubin — the depth of collaboration advancing with each architecture generation.
The biggest structural change this time is that the partnership has evolved from a simple division of labor — "AWS provides the compute site + NVIDIA provides the chips" — to deep co-design at the levels of chip interconnect (NVLink Fusion access to Trainium), memory architecture (NVHBM), government security compliance (IL6), and robotics platforms (Isaac). This depth of coupling makes the cost of either party switching to an alternative partner significantly higher in the future.
Validation Variables for the Expansion Plan
The pace of delivering the 2 million GPUs is the most critical variable. The Rubin and Rubin Ultra architectures are still in the early stages of volume production ramp-up; whether the supply chain can support this pace of expansion will be the core uncertainty over the next 12 to 18 months. From the 1 million GPUs at GTC 2026 to this additional 2 million, the slope of the demand curve has already exceeded the supply side's initial forecasts — the next question is whether production can keep up.
The deep integration of Trainium with NVLink Fusion signals a strategic adjustment in AWS's self-developed chip approach: no longer pursuing Trainium as a complete replacement for NVIDIA GPUs, but rather allowing the two to coexist complementarily. This is a clear signal to Intel, AMD, and other vendors seeking to enter the AI data center chip market — single-point competition at the chip level is no longer sufficient to reshape the landscape; deep binding of software ecosystems and interconnect protocols is the true moat at this stage.
The 100,000-GPU deployment for government AI factories is expected to become a landmark reference for U.S. government AI procurement and may trigger other cloud providers to follow suit in the secure compute market. This is a relatively unoccupied market segment — once the supply side establishes infrastructure standards, subsequent government AI application deployments will revolve around this framework. The change in AWS's share of NVIDIA's major customer revenue in the next fiscal quarter, along with the actual volume production timeline for Rubin-architecture GPUs, will be the two figures that validate whether the expansion plan is progressing on schedule.
© 2026 Winzheng.com 赢政天下 | 转载请注明来源并附原文链接