Skip to content
30 August 2026

AWS and NVIDIA Expand AI Infrastructure with 2 Million More GPUs

AWS and NVIDIA are set to deploy 2 million additional NVIDIA GPUs across AWS's global infrastructure by 2028, enhancing AI capabilities for various industries.

AWS and NVIDIA Expand AI Infrastructure with 2 Million More GPUs

The landscape of artificial intelligence is evolving rapidly, and at the forefront of this transformation are Amazon Web Services (AWS) and NVIDIA. In a groundbreaking announcement, the two tech giants revealed plans to deploy an additional 2 million NVIDIA GPUs across AWS’s global infrastructure by 2028. This expansion is part of a broader strategic collaboration aimed at accelerating AI development and deployment on an unprecedented scale.

This initiative builds on a long-standing partnership between AWS and NVIDIA, which has spanned 16 years of joint innovation. The companies have consistently pushed the boundaries of AI capabilities, from launching the world’s first GPU-accelerated cloud instance to offering the widest range of NVIDIA GPU solutions available today. As AI workloads continue to scale at a rapid pace, this expanded collaboration is poised to meet the growing demands of customers across various sectors.

Expanding AI Compute Capacity

At the NVIDIA GTC 2026 conference, AWS initially announced plans to add more than 1 million NVIDIA GPUs starting in 2026. However, the demand for AI infrastructure has surpassed these expectations, prompting the companies to scale up their commitments. The new plan involves deploying an additional 2 million NVIDIA GPUs, including the latest Blackwell Ultra, Rubin, and Rubin Ultra models, across AWS’s global infrastructure by 2028.

This massive expansion will power a diverse range of customer workloads, from agentic AI and scientific discovery to enterprise automation and physical AI. The increased capacity will also support the deployment of NVIDIA Blackwell capacity, including the RTX PRO 4500 Blackwell Server Edition GPUs for Amazon EC2 G7 instances. These instances deliver significant performance improvements, offering 4.6x AI inference performance and 2.1x graphics performance compared to previous-generation G6 instances.

Enhancing AI Infrastructure with NVIDIA Vera CPUs

In addition to expanding GPU capacity, AWS and NVIDIA are collaborating to bring NVIDIA Vera CPU-based infrastructure to AWS. Vera CPUs are designed to support agentic AI workloads that require high-performance CPU compute alongside accelerated infrastructure. This addition provides customers with yet another option to configure their AI infrastructure, aligning with AWS’s strategy of offering the broadest possible set of compute choices.

Vera CPUs are particularly suited for the CPU work behind agentic AI and reinforcement learning, including code execution, tool use, sandboxing, analytics, data pipelines, and orchestration. As both a host CPU for accelerated systems and a standalone CPU for AI factory workloads, Vera ensures that GPUs are kept fed, agents remain responsive, and training loops continue to move efficiently.

Advancing AI with NVLink Fusion and NVHBM

At the re:Invent 2026 conference, AWS announced support for NVIDIA NVLink Fusion high-speed chip interconnect technology in next-generation Trainium chips. This technology, combined with NVIDIA’s new custom high-bandwidth memory (NVHBM), enables Annapurna Labs to tap into NVIDIA’s custom memory technology and scale-up architecture. The integration of NVLink Fusion and NVHBM enhances performance and efficiency for AI workloads while seamlessly integrating Trainium and GPUs within a common rack-scale architecture.

This heterogeneous AI infrastructure allows for more efficient data processing and vector indexing on platforms like Amazon EMR and Amazon OpenSearch. By leveraging NVIDIA cuDF and cuVS CUDA-X libraries, customers can achieve faster, more cost-efficient analytics and AI applications. Furthermore, the collaboration extends to robotics workloads, with Amazon Robotics adopting NVIDIA’s physical AI platform to speed innovation in warehouse automation and next-generation robots.

The expanded collaboration between AWS and NVIDIA is set to revolutionize AI development and deployment. By deploying 2 million additional NVIDIA GPUs and enhancing AI infrastructure with Vera CPUs, NVLink Fusion, and NVHBM, the companies are poised to meet the surging demand for AI capabilities. This initiative will empower customers across various industries to build and deploy AI solutions at an unprecedented scale, driving innovation and efficiency in the AI era.

Author

Beatrice Mitchell

Beatrice Mitchell, Manchester-rooted and classically elegant, famously commissioned a rebuttal series after a controversial council planning meeting in Stockport, insisting on community testimony. Holds a firm editorial line on accountability and narrative fairness, and collects vintage city planning maps as an idiosyncratic hobby.