Skip to main content
Google Cloud Documentation
Technology areas
  • AI and ML
  • Application development
  • Application hosting
  • Compute
  • Data analytics and pipelines
  • Databases
  • Distributed, hybrid, and multicloud
  • Industry solutions
  • Migration
  • Networking
  • Observability and monitoring
  • Security
  • Storage
Cross-product tools
  • Access and resources management
  • Costs and usage management
  • Infrastructure as code
  • SDK, languages, frameworks, and tools
/
Console
  • English
  • Deutsch
  • Español
  • Español – América Latina
  • Français
  • Indonesia
  • Italiano
  • Português
  • Português – Brasil
  • עברית
  • 中文 – 简体
  • 中文 – 繁體
  • 日本語
  • 한국어
Sign in
  • Cloud TPU
Start free
Overview Guides Reference Samples Support Resources
Google Cloud Documentation
  • Technology areas
    • More
    • Overview
    • Guides
    • Reference
    • Samples
    • Support
    • Resources
  • Cross-product tools
    • More
  • Console
  • Discover
  • Introduction to Cloud TPU
  • TPU architecture
  • TPU versions
    • TPU7x (Ironwood)
    • TPU v6e
    • TPU v5p
    • TPU v5e
    • TPU v4
  • Regions and zones
  • Cloud TPU resources in Compute Engine
  • JAX AI stack
  • Get started
  • Set up a Google Cloud project
  • Quickstarts
    • Create a TPU instance
    • Create a multi-host TPU slice
  • Plan your Cloud TPU resources
  • Reserve TPUs
    • About TPU reservations
    • Request a reservation for up to 90 days (in calendar mode)
    • Request a future reservation for one year or longer
    • Share a reservation
    • Consume a reservation
  • Create TPUs
    • TPU creation overview
    • Create a TPU VM instance
    • Create TPU Flex-start VMs
    • Create TPU Spot VMs
    • TPU instances in MIGs
    • Create a MIG for a multi-host TPU slice
    • Create a MIG for single-host TPU slices
  • Configure TPUs
  • Configure networking and access
  • TPU OS images
  • Use a custom OS image
  • Encrypt a TPU VM boot disk with a CMEK
  • Use a cross-project service account
  • Manage storage
    • Storage options for TPU VMs
    • Storage best practices
    • Attach storage disks
    • Connect to Cloud Storage buckets
  • Manage TPUs
  • Manage TPU resources with Compute Engine
  • Manage maintenance events in managed capacity mode
  • Manage TPUs in All Capacity mode
    • All Capacity mode overview
    • Request an All Capacity mode reservation
    • View the topology and health of All Capacity mode TPUs
    • Report and repair faulty hosts in All Capacity mode
    • Manage maintenance events in All Capacity mode
  • Multislice training
  • Schedule TPU collections for inference workloads
  • Run workloads
  • Train a model using TPU7x
  • Run inference on Cloud TPU
  • Train on Cloud TPU slices
  • Scale a model on TPUs
  • Work with image datasets
    • Convert an image classification dataset for use with Cloud TPU
    • Download, pre-process and upload the ImageNet dataset
    • Download, pre-process and upload the COCO dataset
  • Optimize performance
  • Cloud TPU performance guide
  • Improve your model's performance with bfloat16
  • TPU7x (Ironwood) performance optimizations
  • Monitor and troubleshoot TPUs
  • Monitor TPU VMs
  • Monitor TPU health
  • Monitor TPU goodput
  • TPU monitoring library
  • Monitor with tpu-info CLI
  • Troubleshoot PyTorch models
  • Troubleshoot JAX models
  • ML Diagnostics platform
    • Overview
    • Set up GKE
    • Get started with the SDK
    • Get started with the CLI
    • Use ML Diagnostics with MaxText
    • View machine learning runs
    • Monitor workloads
  • Profile TPUs
  • Profile TPU VMs
  • Profile Multislice environments
  • Profile PyTorch/XLA workloads
  • Cloud TPU API
  • Discover
    • TPU software versions
    • TPU versions
      • TPU v3
      • TPU v2
  • Get started
    • Set up a Google Cloud project
    • Create TPUs
    • Consume a reservation
    • Run JAX on Cloud TPU VM
    • Run PyTorch on Cloud TPU VM
    • Run JAX on Cloud TPU slices
    • Run PyTorch on Cloud TPU slices
  • Configure TPUs
    • Connect a TPU to a shared VPC network
    • Connect to a TPU VM without a public IP address
    • Configure networking and access
    • Encrypt a TPU VM boot disk with a CMEK
    • Use a cross-project service account
    • Attach durable block storage to a TPU VM
    • Connect to Cloud Storage buckets
    • Mount a Filestore instance on a TPU VM
  • Manage TPUs
    • Manage TPU resources
    • Manage queued resources
    • Request TPU Flex-start VMs
    • Manage TPU Spot VMs
    • Prepare for maintenance events
    • Autocheckpoint
    • View maintenance notifications
    • Manually start host maintenance
    • Preemptible TPUs
  • Run workloads
    • Train a model using v6e
    • Train a model using v5e
    • Scale ML workloads using Ray
    • Run TPU applications in a Docker container
  • Monitor and troubleshoot TPUs
    • Troubleshoot TPU VMs
    • Monitor TPU VMs
    • Dashboards for monitoring and logging
    • Troubleshoot TensorFlow models
    • Cloud TPU error glossary
    • Cloud TPU audit logs
  • Tutorials and notebooks
    • Train ResNet with PyTorch
    • MaxDiffusion inference on v6e
    • Notebooks
  • AI and ML
  • Application development
  • Application hosting
  • Compute
  • Data analytics and pipelines
  • Databases
  • Distributed, hybrid, and multicloud
  • Industry solutions
  • Migration
  • Networking
  • Observability and monitoring
  • Security
  • Storage
  • Access and resources management
  • Costs and usage management
  • Infrastructure as code
  • SDK, languages, frameworks, and tools
  • Home
  • Documentation
  • AI and ML
  • Cloud TPU
  • Guides
Stay organized with collections Save and categorize content based on your preferences.

Cloud TPU performance guide

Your first step when troubleshooting TPU performance is to profile your model. For more information on capturing a performance profile, see Profile your model.