By using this site, you agree to our Privacy Policy and Terms of Use.
Accept
VellaTimesVellaTimesVellaTimes
  • News
    NewsShow More
    Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
    Espresso Extraction Science: The Finer Grind Flaw
    May 18, 2026
    A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
    Amazon Alexa for Shopping Replaces Rufus AI Assistant
    May 18, 2026
    Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
    OpenAI acquires Weights.gg to boost voice AI tools
    May 18, 2026
    Federal agents standing outside a modern university biology laboratory building at dusk during an active investigation.
    US Arrests Chinese Scientists for Smuggling Biological Materials
    May 18, 2026
    A dramatically lit modern corporate courtroom with futuristic technology elements, representing a high-stakes artificial intelligence legal trial.
    Elon Musk OpenAI Lawsuit Exposes Clashes Over AI Safety
    May 18, 2026
  • Technology
    TechnologyShow More
    Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
    OpenAI acquires Weights.gg to boost voice AI tools
    May 18, 2026
    A polished silicon wafer rests on a surface inside a modern semiconductor manufacturing facility.
    Samsung Strike Threatens Global AI Chip Production
    May 18, 2026
    A glowing computer screen displaying the text GPT-5.5 Instant in a modern, high-tech office environment with soft blue and purple lighting.
    GPT-5.5 Instant: OpenAI’s New Default ChatGPT Model
    May 10, 2026
    Wide view of a modern AI data center with server racks, glowing fiber-optic cables, and semiconductor hardware in the foreground.
    AI Infrastructure Spending Drives Nvidia, AMD Shares
    May 10, 2026
    A glowing computer monitor displaying lines of code and digital network graphics in a modern tech office setting.
    Airbnb AI Coding: 60% of New Software Now Generated by AI
    May 9, 2026
  • AI
    AIShow More
    A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
    Amazon Alexa for Shopping Replaces Rufus AI Assistant
    May 18, 2026
    A dramatically lit modern corporate courtroom with futuristic technology elements, representing a high-stakes artificial intelligence legal trial.
    Elon Musk OpenAI Lawsuit Exposes Clashes Over AI Safety
    May 18, 2026
    A high-tech global map visualization showing glowing digital connections across different continents, representing the worldwide adoption of artificial intelligence.
    Global AI Adoption in 2026: Trends and Growing Divide
    May 10, 2026
    A modern smartphone displaying an artificial intelligence chat interface used for online shopping and product comparison.
    Alibaba Qwen AI Taobao Integration Launches Agentic Shopping
    May 10, 2026
    A split-screen illustration showing a high-tech modern office using advanced AI tools contrasted against an older, dimly lit workspace.
    Global AI Adoption Surges But Rich-Poor Divide Widens
    May 9, 2026
  • Science
    ScienceShow More
    Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
    Espresso Extraction Science: The Finer Grind Flaw
    May 18, 2026
    Federal agents standing outside a modern university biology laboratory building at dusk during an active investigation.
    US Arrests Chinese Scientists for Smuggling Biological Materials
    May 18, 2026
    Header image of a quantum communication lab setup with fiber-optic equipment, a telecom quantum dot device, and interferometer components used for long-distance quantum key distribution.
    Quantum Key Distribution Reaches 120 km With Quantum Dots
    May 10, 2026
    Abstract geometric representation of glowing quantum paraparticles interacting within a three-dimensional mathematical grid in deep blue and gold tones.
    Quantum Paraparticles Exist: New Math Challenges Physics
    May 10, 2026
    A large expedition cruise ship is navigating rough ocean waters under a cloudy sky.
    Global Authorities Respond to Andes Hantavirus Outbreak on MV Hondius Cruise Ship
    May 9, 2026
  • World
    WorldShow More
    Allu Arjun Commitment to Ethical Brand Partnerships
    Exploring Allu Arjun’s Commitment to Ethical Brand Partnerships
    December 18, 2023
    Orry aka Orhan Awatramani
    Orhan Awatramani ‘Orry’ Biography, Lifestyle and Rise to Fame
    December 8, 2023
    Alia Bhatt Latest Deepake Video Victim
    Alia Bhatt becomes latest victim of Deepfake Videos, Obscene Video goes Viral
    November 28, 2023
    Napoleon Movie Review
    Napoleon Movie Review: A Historical Epic by Ridley Scott Reviewed
    November 25, 2023
  • Bookmarks
Search
Category
  • News
  • Technology
  • AI
  • Science
  • World
Company
  • About Us
  • Contact Us
  • Fact Checking Policy
  • Terms & Conditions
  • Privacy Policy
  • Copyright Policy
Resources
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
© 2022 VellaTimes • All Rights Reserved.
Reading: Google and NVIDIA Launch Integrated Architecture to Cut AI Inference Costs
Share
Notification Show More
Font ResizerAa
VellaTimesVellaTimes
Font ResizerAa
  • News
  • Technology
  • AI
  • Science
  • World
Search
  • Explore
    • News
    • Technology
    • AI
    • Science
    • World
  • Useful Links
    • About Us
    • Contact Us
    • Fact Checking Policy
    • Terms & Conditions
    • Privacy Policy
    • Copyright Policy
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
© 2022 VellaTimes • All Rights Reserved.
News

Google and NVIDIA Launch Integrated Architecture to Cut AI Inference Costs

Sameer Katoch
Last updated: 26/04/2026
Sameer Katoch
Share
6 Min Read
Advanced Google Cloud and NVIDIA server racks glowing with blue and green lights inside a modern, high-performance data center.

Artificial intelligence projects often face a significant hurdle when moving from the laboratory to real-world deployment: managing AI inference costs at an enterprise scale. To address this financial bottleneck, NVIDIA and Google Cloud have introduced a deeply integrated hardware and software roadmap. Announced at the Google Cloud Next conference in Las Vegas, the partnership aims to redefine how companies deploy and operate complex models.

Contents
Tackling the Expense of Running Models at ScaleGoogle Splits Silicon Focus with TPU 8i and TPU 8tUnlocking Innovation for Heavily Regulated IndustriesFueling Agentic Systems and Physical Factory AutomationReal-World Adoption by Global Technology Leaders

The collaboration tackles the growing financial and energy demands of running deployed AI models. By co-designing the underlying technology, the two tech giants have created an infrastructure that drastically reduces AI inference costs while improving processing speed. The companies also revealed dedicated silicon and secure computing environments to support highly regulated industries and complex industrial applications.

Tackling the Expense of Running Models at Scale

While training an artificial intelligence model happens only a handful of times, inference occurs millions of times a day as models answer questions, analyze documents, or run applications. This ongoing operational expense has forced many companies to limit their projects. To solve this, Google Cloud and NVIDIA introduced the new A5X bare-metal instances, which operate on NVIDIA Vera Rubin NVL72 rack-scale systems.

By optimizing the hardware and software layers together, the A5X instances deliver up to ten times lower inference cost per token compared to previous generations. The system also achieves ten times higher token throughput per megawatt, offering a significant boost in energy efficiency.

Preventing processing delays across thousands of chips requires massive bandwidth. The A5X architecture resolves this by pairing NVIDIA ConnectX-9 SuperNICs with Google Virgo networking technology. This setup allows operations to scale up to 80,000 NVIDIA Rubin GPUs within a single site and up to 960,000 GPUs across multi-site deployments.

Google Splits Silicon Focus with TPU 8i and TPU 8t

As the industry shifts its focus toward running applications rather than just training them, Google announced a major change to its proprietary silicon lineup. For the first time, Google is splitting its eighth-generation Tensor Processing Units into two distinct models designed for different tasks.

The TPU 8i is built specifically to handle reasoning and inference. It triples on-chip SRAM to 384 MB and increases high-bandwidth memory to 288 GB. This design bridges the gap between processing speed and data access, delivering an 80% improvement in performance per dollar for inference tasks.

Meanwhile, the TPU 8t focuses entirely on training massive models. A single superpod connects 9,600 chips to deliver 121 exaflops of compute power and two petabytes of shared memory. This allows developers to shrink model training times from months down to weeks.

Unlocking Innovation for Heavily Regulated Industries

Strict data sovereignty rules and privacy concerns frequently stall machine learning initiatives in sectors like finance and healthcare. To meet these compliance demands, Google Gemini models running on NVIDIA Blackwell and Blackwell Ultra GPUs are entering preview on Google Distributed Cloud.

This deployment incorporates NVIDIA Confidential Computing, a hardware-level security protocol. It ensures that prompts and fine-tuning data remain completely encrypted within a protected environment, preventing cloud infrastructure operators from viewing or altering the underlying data. Additionally, Google introduced a public cloud preview of Confidential G4 VMs equipped with NVIDIA RTX PRO 6000 Blackwell GPUs, giving multi-tenant environments access to these same cryptographic protections.

Fueling Agentic Systems and Physical Factory Automation

Building autonomous systems requires vast computing power and specialized software. NVIDIA Nemotron 3 Super is now available on the Gemini Enterprise Agent Platform, providing developers with tailored tools for agentic tasks. To ease the operational burden of long reinforcement learning cycles, the companies launched Managed Training Clusters to automate infrastructure management and failure recovery.

The partnership also extends into the physical world. NVIDIA Omniverse libraries and the open-source NVIDIA Isaac Sim framework are now available through the Google Cloud Marketplace. Developers use these tools to build physically accurate digital twins and train robotic simulation pipelines before deploying them to factory floors. Major industrial software providers, including Cadence and Siemens, are already running their solutions on this accelerated infrastructure.

Real-World Adoption by Global Technology Leaders

Early adopters are already translating this advanced hardware into quantifiable returns. OpenAI leverages large-scale inference on NVIDIA GB300 and GB200 NVL72 systems on Google Cloud to handle its most demanding workloads, including operations for ChatGPT.

Snap migrated its large-scale data pipelines to GPU-accelerated Spark on Google Cloud to cut the extensive costs tied to A/B testing. In the pharmaceutical sector, Schrödinger uses NVIDIA accelerated computing to compress complex drug discovery simulations from weeks into just hours.

The developer ecosystem surrounding these tools is expanding rapidly. Over 90,000 developers have joined the joint NVIDIA and Google Cloud community in just one year, signaling a fast-paced shift toward more accessible, high-performance computing.

TAGGED: agentic AI, AI inference costs, Confidential Computing, Google Cloud, Nvidia, TPU 8i, Vera Rubin NVL72
Share This Article
Facebook Twitter Whatsapp Whatsapp Telegram Copy Link
By Sameer Katoch
As the Founder of VellaTimes and an avid traveler, I'm passionate about the daily news events happening globally. With over five years of experience in the writing field, I am committed to delivering top-notch news that satisfies your daily news intake.
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *


Most Read

Amazon OpenAI investment talks: up to $50 billion

January 30, 2026

Google Gemini Workspace Features: Powerful AI Upgrades

March 18, 2026

OpenAI Shuts Down Sora Video App to Focus on Robotics

March 30, 2026

Researchers Bypass No-Cloning Theorem for Secure Quantum Cloud Storage

March 10, 2026

Ola CEO Bhavish Aggarwal unveils new Indian AI Chat App ‘Krutrim’

November 28, 2023

Ofcom WhatsApp probe examines Meta data responses in UK

January 24, 2026

Related News

Close-up of a silver espresso machine extracting a fresh shot of coffee into a glass cup in a softly lit cafe setting.
News

Espresso Extraction Science: The Finer Grind Flaw

Nisha Pradhan Nisha Pradhan May 18, 2026
A smartphone resting on a wooden desk displaying an AI-powered Amazon search bar in a modern home office setting.
News

Amazon Alexa for Shopping Replaces Rufus AI Assistant

Sameer Katoch Sameer Katoch May 18, 2026
Wide news-style image showing an OpenAI office scene with screens displaying audio waveforms and voice technology graphics
News

OpenAI acquires Weights.gg to boost voice AI tools

Rakesh Paul Rakesh Paul May 18, 2026

About Us

VellaTimesVellaTimesVellaTimes

VellaTimes is a leading news portal that covers the latest trending news in technology, lifestyle, entertainment, automobiles, travel, and sports.

Explore

  • News
  • Technology
  • AI
  • Science
  • World

Useful Links

  • About Us
  • Contact Us
  • Fact Checking Policy
  • Terms & Conditions
  • Privacy Policy
  • Copyright Policy

Subscribe Us

Subscribe to our newsletter for the Latest News and Top Stories!

© 2022 VellaTimes • All Rights Reserved.
  • Home
  • Web Stories
  • Bookmarks
  • Interests
  • Disclaimer
  • Sitemap
adbanner
AdBlocker Detected
Our site is an advertising supported site. Please whitelist us to support our work.
Okay, I'll Whitelist