Google Reveals Ironwood Chip to Accelerate AI Inference at Scale
Google introduces Ironwood, its seventh-generation TPU designed specifically for AI inference, promising massive computing power and energy efficiency for cloud customers. The new chip intensifies the ongoing battle among tech giants to build custom AI hardware.
Google unveils Ironwood, its seventh-generation TPU AI accelerator chip, at its Cloud Next conference. Unlike previous versions, Ironwood is the first TPU purpose-built specifically for inference, which means it focuses on running AI models rather than just training them. Google plans to release the chip later this year to Google Cloud customers in two different cluster sizes of 256 chips and 9,216 chips.
The new hardware delivers impressive specifications, including 4,614 TFLOPs of peak computing power and 192GB of dedicated RAM per chip with 7.4 Tbps bandwidth. Google includes an enhanced specialized core called SparseCore to handle heavy data workloads like product recommendations and search rankings. The chip's architecture minimizes data movement to reduce latency and save power, making it Google's most energy-efficient TPU to date.
Ironwood enters a fiercely competitive AI accelerator market currently led by Nvidia. Tech giants like Amazon and Microsoft are aggressively pushing their own custom silicon, such as Amazon's Trainium and Inferentia chips and Microsoft's Maia 100. Google integrates Ironwood into its AI Hypercomputer infrastructure to offer businesses a powerful alternative for running advanced, thinking AI models at a massive scale.