The landscape of artificial intelligence is evolving rapidly, and at the forefront of this transformation are Amazon Web Services (AWS) and NVIDIA. These tech giants have announced a significant expansion of their collaboration, aiming to deploy an additional 2 million NVIDIA GPUs across AWS data centers by 2028. This initiative is part of a broader strategy to scale AI infrastructure, encompassing everything from accelerator capacity to high-bandwidth interconnects.
The partnership, which has spanned 16 years, is designed to provide customers with the flexibility to choose the best tools for their AI workloads while ensuring seamless integration. This expansion is not just about increasing the number of GPUs but also about enhancing the
Scaled GPU Deployments and Enhanced EC2 Instances
The multi-year rollout will integrate a variety of NVIDIA GPUs including the Blackwell Ultra, Rubin, and Rubin Ultra, into the AWS Global Infrastructure. These GPUs are specifically designed to power large AI factories handling distributed training, scientific simulation, enterprise automation, and agentic workflows.
AWS is also expanding its current-generation Blackwell fleet by introducing the NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs in Amazon EC2 G7 instances. These instances deliver up to 4.6 times higher AI inference performance and 2.1 times higher graphics throughput compared to previous-generation EC2 G6 instances. AWS is the first major cloud provider to offer instances accelerated by the RTX PRO 4500, setting a new standard in AI performance.
To maintain cluster-wide throughput and low latency across these high-density nodes, AWS and NVIDIA are collaborating on NVIDIA Spectrum networking optimizations tailored for large GPU training clusters. This ensures that the infrastructure can handle the most demanding AI workloads efficiently.
Innovations in CPU and Memory Technology
Addressing the compute requirements of multi-step agentic AI workloads, AWS and NVIDIA are collaborating to bring NVIDIA Vera CPU-based infrastructure to AWS. Vera is purpose-built for the next generation of AI, providing a high-performance CPU option alongside accelerated infrastructure. This complements AWS’s strategy of offering the broadest possible choice of compute, from its own custom silicon to partner CPUs and accelerators.
In the interconnect and memory domain, NVIDIA and Amazon’s Annapurna Labs are extending their partnership around NVIDIA NVLink Fusion. Originally introduced for next-generation AWS Trainium processors, NVLink Fusion is now being paired with NVIDIA’s custom high-bandwidth memory (NVHBM) technology. This combination gives Trainium access to faster, more power-efficient memory, allowing Annapurna Labs to combine Trainium silicon and NVIDIA GPUs within a shared, scale-up rack architecture.
Secure AI Factories for Government Workloads
For government-sector requirements, AWS and NVIDIA are co-developing dedicated AI factories equipped with 100,000 GPUs, deployed across isolated, secure AWS infrastructure. These environments are engineered to support sensitive federal agency and defense workloads, providing compliance and operational security for classifications at Impact Level 6 (IL6) and above.
The expanded alliance continues to leverage existing core technical integrations across the AWS stack. All NVIDIA GPU and Trainium-based EC2 instances rely on the AWS Nitro System for offloaded virtualization and hardware security, paired with Elastic Fabric Adapter (EFA) networking to deliver line-rate, low-latency scale-out interconnectivity.
Additionally, the NVIDIA Nemotron open model family remains natively integrated into AWS, offered as fully managed, serverless endpoints on Amazon Bedrock and as customizable foundation models on Amazon SageMaker. This integration supports a wide range of AI applications, from data processing to robotics automation.



