Google introduced Ironwood, its seventh-generation Tensor Processing Unit, on April 9, 2025, describing it as 'our most powerful, capable and energy efficient TPU yet, designed to power thinking, inferential AI models at scale,' according to Amin Vahdat, the company's VP/GM for ML, Systems and Cloud AI. The chip is now available commercially on Google Cloud under the name TPU7x, with Google's documentation noting it shares similarities with the earlier TPU v5p while scaling to pods of 9,216 chips.
Despite being pitched primarily at inference, Google published a developer's guide in March 2026 for training large models on Ironwood, reflecting demand for the chip across both training and inference workloads.
Why it matters
TPUs are Google's answer to a broader industry shift toward custom silicon designed for a company's own AI stack, rather than relying solely on merchant GPUs from Nvidia and AMD. Amazon has done the same with Trainium, and Microsoft is developing its own Maia chips. For Google Cloud customers, having a purpose-built inference chip is a differentiator against providers whose infrastructure is built around general-purpose GPUs.
How it works
Google's documentation describes Ironwood as custom-designed for large-scale AI workloads with 'a significant performance improvement over previous TPU generations,' available through Google Kubernetes Engine for customers running inference and training pipelines. The chip's pod-scale design, connecting thousands of chips, is intended to let very large models run inference across a single coordinated pool rather than being split across incompatible smaller clusters.
Evidence
In an April 6, 2026 post co-authored by Google Distinguished Engineer David Patterson, Google said Ironwood TPUs deliver 3.7x carbon efficiency gains, part of the company's stated effort to publish transparent metrics on the environmental impact of its AI infrastructure. That figure is a company-reported efficiency comparison rather than an independently audited benchmark.
The competing read
Google's efficiency and performance claims for Ironwood come from its own blog posts and documentation, and the company has an incentive to promote TPUs as it competes for cloud AI workloads against Nvidia-based offerings from AWS, Microsoft Azure and Oracle. Independent, cross-vendor benchmarks comparing Ironwood directly against Nvidia's Blackwell or Rubin GPUs on identical workloads were not part of the sourced material here, so the comparative performance claims should be treated as Google's own framing.
What happens next
Google has continued rolling out developer tooling for Ironwood through 2026, including guidance for both training and inference, suggesting the company intends TPU7x to be a general workhorse rather than a narrowly inference-only part despite its original positioning. How much of Google's own AI workload, including Gemini, shifts further onto Ironwood versus external GPU capacity will be a signal of the chip's real-world cost advantage.
