Google Aims to Flip the Script on AI Inference with New Ironwood TPUs
Google Cloud may not be the first name that comes to mind when you think “high-end chipmaker.” But with its seventh-generation tensor processing unit (TPU) dubbed “Ironwood” becoming available this month, it’s a name that perhaps should be.
AI inference is the name of the game these days as the world’s biggest companies look to monetize the exceptional capabilities demonstrated by foundation models. However, AI training and AI inference are different workloads, so the world’s top chipmakers are developing optimized chips suited for each.
Google Cloud is hoping to garner as much of that market for AI inference as it can with its Ironwood TPU, which was developed specifically for AI inference workloads and which will soon be available in its public cloud.
Ironwood, which Google Cloud first unveiled at its Next 25 conference in April, boasts impressive stats. Each TPU is capable of delivering 4.6 petaFLOPS of FP8 performance, which is 4x what Google crammed into its sixth-generation TPU, “Trillium.” Ironwood’s FP8 performance exceeds that of Nvidia’s B200 GPU, which delivers 4.5 petaFLOPS of FP8 performance. However, it trails the 5 petaFLOPS delivered by the GB200 Blackwell GPU.
Google Cloud will make its seventh-generation TPU, dubbed “Ironwood,” available in late November 2025
Each Ironwood TPU boasts 192 GB of high-bandwidth memory (HBM3E), which is 6x as much as Trillium. Google says it can move data in and out of the TPU at the rate of 7.2 TBps per chip, which is 4.5x what Trillium had to offer.
These individual stats are impressive, but it’s the team stats where Ironwood really shines. Google Cloud is connecting TPUs using its proprietary’s Inter Chip Interconnect (ICI) technology, which boasts 1.2 TBps of bidirectional throughput.
Google is offering two Ironwood pods: a 256-chip pod, and a 9,216-chip pod. That 9,216-chip pod features some really gaudy status. According to Google, it will boast 1.77 PB of shared HBM, all of which is accessible as a single unified high-speed fabric. All told, this AI hypercluster could theoretically deliver 42.5 FP8 ExaFLOPS of peak performance.
However, Google can scale Ironwood even further utilizing its Optical Circuit Switching (OCS) technology. According to Mark Lohmeyer, the VP & GM, Compute and AI Infrastructure, Google can use OCS and its Jupiter datacenter network technology to connect hundreds of these 9,216-chip pods together into a single logical unit that spans hundreds of thousands of TPUs. Now that is the sort of scale that could be a gamechanger.
Google Cloud is offering its “Ironwood” TPUs in 9,612 pods
Not surprisingly, Google is using Ironwood to serve its own Gemini models, as well as a host of other Google services that leverage Gemini, such as YouTube, search, and Gmail. Lohmeyer boasts that 9 of the top 10 AI labs are using Google Cloud to train or run their AI.
Google Cloud has several other innovations up its sleeve with its Ironwood clusters. It debuted its third-generation liquid cooling design to keep the chips cool. While Google Cloud hasn’t released the thermal design power (TDP) profile of individual chips, it says a full contingent of 9,612 chips will draw about 10 megawatts. That would suggest that each chips draws around 1,000 watts per chip.
It rolled out its Cloud Storage Anywhere Cache as part of Google Cloud Storage, which enables customers to cache data right next to the accelerator (TPU or GPU), which the company says can reduce read latency by up to 96%.
The hyperscaler has been working to develop open source software, too. Lohmeyer touted Google’s co-founding of LLM-d to provide Kubernentes-native distributed inference, which allows more efficient scaling of AI inference workloads. And it made a contribution to the vLLM project to help make it easier to move workloads between TPUs and GPUs.