Skip to main content
Google Cloud Documentation
Technology areas
  • AI and ML
  • Application development
  • Application hosting
  • Compute
  • Data analytics and pipelines
  • Databases
  • Distributed, hybrid, and multicloud
  • Industry solutions
  • Migration
  • Networking
  • Observability and monitoring
  • Security
  • Storage
Cross-product tools
  • Access and resources management
  • Costs and usage management
  • Infrastructure as code
  • SDK, languages, frameworks, and tools
/
Console
  • English
  • Deutsch
  • Español – América Latina
  • Français
  • Indonesia
  • Italiano
  • Português – Brasil
  • עברית
  • 中文 – 简体
  • 中文 – 繁體
  • 日本語
  • 한국어
Sign in
  • AI Hypercomputer
Start free
Overview Guides Resources
Google Cloud Documentation
  • Technology areas
    • More
    • Overview
    • Guides
    • Resources
  • Cross-product tools
    • More
  • Console
  • Discover
  • Overview
  • Choose your accelerator infrastructure
  • Choose an orchestrator and deployment option
  • Accelerators
    • About GPU accelerators
    • General GPUs
    • Clustered GPUs
  • Performance-optimized infrastructure
    • Networking services
      • GPU networking overview
      • Network services for deployments
      • Networking best practices
    • Storage services
  • Open software
    • OS and Docker images
  • Choose a consumption option
  • Cluster management
    • Cluster management overview
    • Cluster management configurations
  • Terminology
  • Get started
  • Plan and create your AI infrastructure
  • Choose between general GPUs and clustered GPUs
  • Design your AI infrastructure with Gemini
  • Obtain capacity and quota
    • Overview
    • Reserve capacity
    • View reserved capacity
  • Quickstarts
    • Create a fully managed Slurm cluster with A4 VMs
    • Create a self-managed Slurm cluster with A4 VMs
  • Deploy infrastructure
  • Deployment options overview
  • Compact placement policy and workload policy overview
  • Deploy AI-optimized VMs and clusters
    • Create GKE clusters
      • Create an AI-optimized GKE cluster with default configuration
      • Create a custom AI-optimized GKE cluster which uses A4X Max
      • Create a custom AI-optimized GKE cluster which uses A4X
      • Create a custom AI-optimized GKE cluster which uses A4 or A3 Ultra
      • Create GKE Standard clusters which use A3 Mega or A3 High
      • Create GKE Autopilot clusters which use A3 Mega or A3 High
    • Create Slurm clusters
      • Create a fully managed cluster
      • Create a self-managed cluster
    • Create an instance
      • Create A4X Max
      • Create A4X
      • Create A4 or A3 Ultra
      • Create A3 High or A3 Mega
      • Create A3 Edge
      • Create A2
      • Create G4
      • Create G2
      • Create N1 with GPUs
    • Create instances in bulk
      • Create A4X Max
      • Create A4X
      • Create A4 or A3 Ultra
      • Create A3 High or A3 Mega
      • Create A3 Edge
      • Create A2
      • Create G4
      • Create G2
      • Create N1 with GPUs
    • Create a managed instance group (MIG)
      • Create A4X Max
      • Create A4X
      • Create A4 or A3 Ultra
      • Create A3 High or A3 Mega
      • Create A3 Edge
      • Create A2
      • Create G4
      • Create G2
      • Create N1 with GPUs
  • Run workloads
  • Run workloads with Pathways on Cloud
    • Introduction to Pathways on Cloud
    • Create a GKE cluster with Pathways
    • Run a batch workload with Pathways
    • Run an interactive workload with Pathways
    • Perform multihost inference using Pathways
    • Resilient training with Pathways
    • Port JAX workloads to Pathways
    • Troubleshoot Pathways on Cloud
  • Schedule GKE workloads
    • Schedule workloads with Topology Aware Scheduling (TAS)
    • Enable node health prediction
  • Use Kueue to orchestrate workloads
  • AI workload tutorials
    • Overview
    • GPU
      • Run inference with vLLM on GKE
        • DeepSeek V3.1
        • DeepSeek V3.2-Speciale
        • Gemma 3
        • GPT-OSS
        • Llama 4
        • Qwen3
      • Run fine-tuning
        • Gemma 3 on a GKE cluster
        • Gemma 3 on a multi-host GKE cluster
        • Gemma 3 on a Slurm cluster
        • Gemma 3 for vision tasks on GKE
        • Llama 4 on a Slurm cluster
        • Mixtral-8x7b on a Slurm cluster
      • Run training
        • Qwen2 on a Slurm cluster
    • TPU
      • Serve workloads
        • Serve Qwen2-7B with vLLM on TPUs
        • Serve Qwen2-7B-Instruct with vLLM on TPUs
        • Serve Qwen3-8B-Base with vLLM on TPUs
        • Serve Llama-3.1-8B with vLLM on TPUs
      • Run fine-tuning
        • Run supervised fine-tuning on a TPU VM with MaxText
        • Run multi-host supervised fine-tuning on Qwen3-14b using MaxText
        • Run RL training on a v6e-8 TPU VM