This page documents production updates to Cloud TPU. You can periodically check this page for announcements about new or updated features, bug fixes, known issues, and deprecated functionality.
You can see the latest product updates for all of Google Cloud on the Google Cloud page, browse and filter all release notes in the Google Cloud console, or programmatically access release notes in BigQuery.
To get the latest product updates delivered to you, add the URL of this page to your feed reader, or add the feed URL directly.
June 01, 2026
Generally available: Compute Engine supports Google's custom-developed accelerator Tensor Processing Unit (TPU), providing a converged experience across AI accelerators on Google Cloud. You can use the Compute Engine instance API and managed instance group (MIG) API to create and manage TPU VMs. You can perform standard VM configurations such as using a custom OS or configure boot disk size. Compute Engine APIs support the creation and management of TPU slices across all consumption options, enabling small-scale experimentation and large-scale training and inference workloads.
For more information, see TPU resources in Compute Engine.
April 27, 2026
Generally available: Cloud TPU now offers TPU availability in AI zones. To learn more, see About AI zones.
March 31, 2026
Generally available: TPU7x is generally available (GA). TPU7x is the first release within the Ironwood family, Google Cloud's seventh generation TPU. TPU7x supports large-scale AI training and inference, providing performance and cost-effectiveness for demanding workloads such as large language (LLMs), mixture of experts (MoEs), and diffusion models. For more information, see the TPU7x (Ironwood) documentation.
November 24, 2025
Preview: TPU7x is available in Preview. TPU7x is the first release within the Ironwood family, Google Cloud's seventh generation TPU. TPU7x supports large-scale AI training and inference, providing performance and cost-effectiveness for demanding workloads such as large language models (LLMs), mixture of experts (MoEs), and diffusion models. For more information, see the TPU7x (Ironwood) documentation.
May 22, 2025
Public preview: You can request Cloud TPUs using future reservations in calendar mode. This mode, powered by the Dynamic Workload Scheduler, lets you check TPU availability up to 120 days in advance and request capacity based on your schedule. You can use calendar mode to reserve TPUs for 1 to 90 days. Requesting a short-term reservation with calendar mode is a good fit for training and experimentation workloads that require precise start times and have a defined duration. For more information, see Request a short-term reservation using calendar mode.
Public preview: You can enable reservation sharing for Cloud TPU. This feature lets you share a reservation across multiple projects. You can also share a reservation with Vertex AI for training or serving workloads. For more information, see Share a Cloud TPU reservation.
March 31, 2025
Flex-start for Cloud TPU, powered by Dynamic Workload Scheduler, is available in Preview. Flex-start is a flexible and cost-effective consumption option for AI workloads. Flex-start enables you to dynamically provision TPUs for up to 7 days using the queued resources API, without long-term reservations. This option is ideal for quick experimentation, small-scale testing, dynamic inference provisioning, and model fine-tuning. For more information about Flex-start for Cloud TPU, see Request Cloud TPUs using Flex-start.
December 16, 2024
This Release Note announces General Availability of Trillium AKA v6e. Trillium is the 6th generation and latest Cloud TPU. It is fully integrated with our AI Hypercomputer architecture to deliver compelling value to our Google Cloud Platform AI customers.
We used Trillium TPUs to train the new Gemini 2.0, Google's most capable AI model yet, and now enterprises and startups alike can take advantage of the same powerful, efficient, and sustainable infrastructure. Today, Trillium is generally available for Google Cloud customers and this week we will be delivering our first large tranches of Trillium capacity to some of our biggest Google Cloud Platform customers.
Here are some of the key improvements that Trillium delivers over the prior generations, v5e and v5p:
Over 4x improvement in training performance.
Up to 3x increase in inference throughput.
A 67% increase in energy efficiency.
An impressive 4.7x increase in peak compute performance per chip.
Double the High Bandwidth Memory (HBM) capacity.
Double the Interchip Interconnect (ICI) bandwidth.
100,000 Trillium chips per Jupiter network fabric with 13 Petabits/sec of bisection bandwidth, capable of scaling a single distributed training job to hundreds of thousands of accelerators.
Trillium provides up to 2.1x increase in performance per dollar over Cloud TPU v5e and up to 2.5x increase in performance per dollar over Cloud TPU v5p in training dense LLMs like Llama2-70b and Llama3.1-405b.
GKE integration enables seamless AI workload orchestration using Google Compute Engine MIGs including XPK for faster iterative development.
Multislice training with Trillium scales from one to hundreds of thousands of chips across pods using DCN.
Training and serving fungibility enables use of same Cloud TPU quota for both training and inference.
Support for collection scheduling with collection SLOs being defended.
Full-host VM support to enable inference support for larger models (70B+ parameters).
Official Libtpu releases that guarantees stability across all three frameworks (Jax/Pytorch-XLA/Tensorflow).
These enhancements enable Trillium to excel across a wide range of AI workloads, including:
Scaling AI training workloads like LLMs including dense and Mixture of Experts (MoE) models
Inference performance and collection scheduling
Embedding-intensive models acceleration
Delivering training and inference price-performance
November 01, 2024
You can now request Cloud TPUs as queued resources in the Google Cloud Console. Queuing your request for TPU resources can help alleviate stockout issues. If the resources you request are not immediately available, your request is added to a queue until the request succeeds or you delete it. You can also specify a time range in which you want to fulfill the resource request. For more information, see Manage queued resources.
Creating a Multislice TPU environment is now available in the Google Cloud Console. You can use Multislice to run training jobs using multiple TPU slices within a single Pod or on slices in multiple Pods. You must use a queued resource request to create a Multislice environment. For more information, see Cloud TPU Multislice overview.
March 11, 2024
Cloud TPU now supports TensorFlow 2.16.1. For more information see the TensorFlow 2.16.1 release notes.
December 04, 2023
Cloud TPU now supports TensorFlow 2.14.1. For more information see the TensorFlow 2.14.1 release notes.
November 13, 2023
Cloud TPU now supports TensorFlow 2.15.0, which adds support for PJRT. For more information see the TensorFlow 2.15.0 release notes.
October 05, 2023
Cloud TPU now supports TensorFlow 2.13.1. For more information see the TensorFlow 2.13.1 release notes.
September 27, 2023
Cloud TPU now supports TensorFlow 2.14.0. For more information see the TensorFlow 2.14.0 release notes.
August 29, 2023
You can now create Cloud Tensor Processing Unit (TPU) nodes in Google Kubernetes Engine (GKE) to run AI workloads, from training to inference models. GKE manages your cluster by automating TPU resource provisioning, scaling, scheduling, repairing, and upgrading. GKE provides TPU infrastructure metrics in Cloud Monitoring, TPU logs, and error reports for better visibility and monitoring of TPU node pools in GKE clusters. TPUs are available with GKE Standard clusters. GKE supports TPU v4 in version 1.26.1.gke-1500 and later, and supports TPU v5e in version 1.27.2-gke.1500 and later. To learn more, see TPUs in GKE introduction.
July 21, 2023
Cloud TPU now supports TensorFlow 2.12.1. For more information see the TensorFlow 2.12.1 release notes.
July 10, 2023
Cloud TPU now supports TensorFlow 2.13.0. For more information see the TensorFlow 2.13.0 Release Notes.