Artificial intelligence projects often face a significant hurdle when moving from the laboratory to real-world deployment: managing AI inference costs at an enterprise scale. To address this financial bottleneck, NVIDIA and Google Cloud have introduced a deeply integrated hardware and software roadmap. Announced at the Google Cloud Next conference in Las Vegas, the partnership aims to redefine how companies deploy and operate complex models.
The collaboration tackles the growing financial and energy demands of running deployed AI models. By co-designing the underlying technology, the two tech giants have created an infrastructure that drastically reduces AI inference costs while improving processing speed. The companies also revealed dedicated silicon and secure computing environments to support highly regulated industries and complex industrial applications.
Tackling the Expense of Running Models at Scale
While training an artificial intelligence model happens only a handful of times, inference occurs millions of times a day as models answer questions, analyze documents, or run applications. This ongoing operational expense has forced many companies to limit their projects. To solve this, Google Cloud and NVIDIA introduced the new A5X bare-metal instances, which operate on NVIDIA Vera Rubin NVL72 rack-scale systems.
By optimizing the hardware and software layers together, the A5X instances deliver up to ten times lower inference cost per token compared to previous generations. The system also achieves ten times higher token throughput per megawatt, offering a significant boost in energy efficiency.
Preventing processing delays across thousands of chips requires massive bandwidth. The A5X architecture resolves this by pairing NVIDIA ConnectX-9 SuperNICs with Google Virgo networking technology. This setup allows operations to scale up to 80,000 NVIDIA Rubin GPUs within a single site and up to 960,000 GPUs across multi-site deployments.
Google Splits Silicon Focus with TPU 8i and TPU 8t
As the industry shifts its focus toward running applications rather than just training them, Google announced a major change to its proprietary silicon lineup. For the first time, Google is splitting its eighth-generation Tensor Processing Units into two distinct models designed for different tasks.
The TPU 8i is built specifically to handle reasoning and inference. It triples on-chip SRAM to 384 MB and increases high-bandwidth memory to 288 GB. This design bridges the gap between processing speed and data access, delivering an 80% improvement in performance per dollar for inference tasks.
Meanwhile, the TPU 8t focuses entirely on training massive models. A single superpod connects 9,600 chips to deliver 121 exaflops of compute power and two petabytes of shared memory. This allows developers to shrink model training times from months down to weeks.
Unlocking Innovation for Heavily Regulated Industries
Strict data sovereignty rules and privacy concerns frequently stall machine learning initiatives in sectors like finance and healthcare. To meet these compliance demands, Google Gemini models running on NVIDIA Blackwell and Blackwell Ultra GPUs are entering preview on Google Distributed Cloud.
This deployment incorporates NVIDIA Confidential Computing, a hardware-level security protocol. It ensures that prompts and fine-tuning data remain completely encrypted within a protected environment, preventing cloud infrastructure operators from viewing or altering the underlying data. Additionally, Google introduced a public cloud preview of Confidential G4 VMs equipped with NVIDIA RTX PRO 6000 Blackwell GPUs, giving multi-tenant environments access to these same cryptographic protections.
Fueling Agentic Systems and Physical Factory Automation
Building autonomous systems requires vast computing power and specialized software. NVIDIA Nemotron 3 Super is now available on the Gemini Enterprise Agent Platform, providing developers with tailored tools for agentic tasks. To ease the operational burden of long reinforcement learning cycles, the companies launched Managed Training Clusters to automate infrastructure management and failure recovery.
The partnership also extends into the physical world. NVIDIA Omniverse libraries and the open-source NVIDIA Isaac Sim framework are now available through the Google Cloud Marketplace. Developers use these tools to build physically accurate digital twins and train robotic simulation pipelines before deploying them to factory floors. Major industrial software providers, including Cadence and Siemens, are already running their solutions on this accelerated infrastructure.
Real-World Adoption by Global Technology Leaders
Early adopters are already translating this advanced hardware into quantifiable returns. OpenAI leverages large-scale inference on NVIDIA GB300 and GB200 NVL72 systems on Google Cloud to handle its most demanding workloads, including operations for ChatGPT.
Snap migrated its large-scale data pipelines to GPU-accelerated Spark on Google Cloud to cut the extensive costs tied to A/B testing. In the pharmaceutical sector, Schrödinger uses NVIDIA accelerated computing to compress complex drug discovery simulations from weeks into just hours.
The developer ecosystem surrounding these tools is expanding rapidly. Over 90,000 developers have joined the joint NVIDIA and Google Cloud community in just one year, signaling a fast-paced shift toward more accessible, high-performance computing.
