• About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
Saturday, August 15, 2026
mGrowTech
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
No Result
View All Result
mGrowTech
No Result
View All Result
Home Google Marketing

Run Ray on TPU, Part 1: The foundations

Josh by Josh
July 20, 2026
in Google Marketing
0
Run Ray on TPU, Part 1: The foundations


header

TL;DR: If you already scale Python with Ray on GPUs, your code can now run on TPU (Tensor Processing Unit, Google’s AI accelerator chip) with fully supported official APIs you already know. The task-and-actor model, a JaxTrainer, the same Ray Serve deployment just pointed at TPUs orchestrated by Google Kubernetes Engine (GKE).

As of Ray 2.55, Google Cloud TPUs are a first-class accelerator in Ray. This means that TPUs are now in Ray’s release pipelines with official pre-built images and support across the core libraries, instead of the old “experimental” path where you built your own containers and leaned on community help. In this “Run Ray on TPU” series, you will learn how a TPU slice is just another accelerator Ray schedules onto (Part 1), then walk through each library (Part 2).

Ray and TPU: A quick introduction

Ray is a distributed-computing framework: you write Python, and Ray runs it across a cluster as tasks (stateless functions) and actors (stateful workers). To Ray, a TPU is just another schedulable resource, the way a GPU is. You ask for it, Ray places your work on it.

But there is one thing to keep in mind, and then we move on.

TPU chips are wired together into a fixed group called a slice: several host machines (VMs) whose chips share a dedicated high-speed link called the ICI (Inter-Chip Interconnect). A multi-host model has to land on one whole slice, or its workers can’t reach each other and the job just hangs.

If you think in GPUs, picture a slice as a single multi-GPU box where the fast interconnect (NVLink) only exists inside the box. Split your workers across two boxes with no cable between them and the collective operations, the all-reduce steps that synchronize gradients, never finish. Training just hangs. A TPU slice behaves the same way: the ICI is that cable, and it only reaches the chips of one slice.

fig1 (2)

That’s the whole reason “Ray on TPU” needs anything special. Something has to guarantee all your workers land on one intact slice. On GPUs you barely think about it; on TPUs it’s crucial, and Ray and GKE (Google Kubernetes Engine, Google’s managed Kubernetes) handle it for you.

One more word you’ll see everywhere is topology: it’s the shape of a slice, written like 4×4 for a 16-chip slice. You ask for a topology, not a chip count.

Once you understand slicing and topology with TPU, the existing Ray stack and your development process remains unchanged and it runs on TPU slices that GKE provisions. The diagram below is the whole system in one picture: on the left, the code you write (the Ray libraries you already use); in the middle, the Ray Core layer that reserves whole slices; on the right, the GKE managed layer that provisions the hardware and labels it so Ray can find slice boundaries.

fig2 (2)

GKE provisions a slice and labels its hosts, Ray Core reads those labels to reserve the whole slice at once, and your library call sits on top, declaring a topology and nothing more. No hand-written placement code anywhere. The rest of this part walks the bottom two layers, GKE then Ray Core while Part 2 covers the Ray AI libraries.

How GKE orchestrates Ray on TPU

You run Ray on TPU through GKE using the Ray Operator add-on.

# Autopilot (fully managed nodes)
gcloud container clusters create-auto CLUSTER \
  --enable-ray-operator --location=LOCATION

# or Standard (you manage node pools)
gcloud container clusters create CLUSTER \
  --addons=RayOperator --location=LOCATION

Shell

That single flag installs two things that matter for TPU. The first, KubeRay, is the Kubernetes operator that turns RayCluster, RayService, and RayJob YAML into running Ray clusters; it’s the same KubeRay you’d use with GPUs. The second is the TPU-specific part: the Ray TPU webhook, which stamps every TPU host with labels like ray.io/tpu-slice-name so Ray can tell which machines are wired into the same slice. That label is the thread the whole system pulls on.

From there, you ask for TPUs in a manifest the same way you’d ask for any node, with a nodeSelector for the generation and topology and the chip count as a resource. A multi-host slice adds one field, numOfHosts.

# inside a RayCluster workerGroupSpec
nodeSelector:
  cloud.google.com/gke-tpu-accelerator: tpu-v6e-slice   # the TPU generation
  cloud.google.com/gke-tpu-topology: "4x4"              # the slice shape
# ... and request chips via the google.com/tpu resource limit
numOfHosts: 4   # multi-host: how many host VMs make up this slice

Shell

GKE provisions the slice, the webhook labels it, Ray reads the labels. You write Python. Once the add-on is up, you’ll see the KubeRay operator pod running, and applying that manifest brings up a head pod plus one worker pod per host in the slice. The cluster step of the get-started example provisions all of this with Terraform.

What actually keeps your workers together is a Ray Core primitive sitting just above this layer, the slice placement group, and that’s where the rest of the guide starts.

Ray Core on TPU

Ray Core is the base layer, the task-and-actor engine and scheduler everything else sits on. Its TPU support lives in the public ray.util.tpu API, and there’s really one function to know: slice_placement_group(). It takes that “keep my workers on one intact slice” guarantee from earlier and turns it into a single call, reserving a whole slice atomically (all hosts or none) by matching on the webhook labels.

from ray.util.tpu import slice_placement_group
from ray.util.scheduling_strategies import PlacementGroupSchedulingStrategy

# Reserve one whole v6e 4x4 slice (16 chips across 4 hosts), atomically
spg = slice_placement_group(topology="4x4", accelerator_version="v6e")
ray.get(spg.placement_group.ready(), timeout=600)

@ray.remote(resources={"TPU": 4})
def worker(rank, world): ...

tasks = [
    worker.options(
        scheduling_strategy=PlacementGroupSchedulingStrategy(
            placement_group=spg.placement_group)
    ).remote(rank=i, world=spg.num_hosts)
    for i in range(spg.num_hosts)
]

Python

It is important to highlight that you rarely call slice_placement_group yourself. The Ray AI libraries (Data, Train, Serve) call it for you, so in practice you declare a topology and they handle the slice. You’d only reach for slice_placement_group() directly when you’re writing a custom distributed workload that isn’t Train, Serve, or Data. One caveat worth knowing: the API is public but marked alpha (@PublicAPI(stability="alpha")), so it’s usable today but the surface can still shift between releases.

That’s the foundation. Next are the libraries.

You now have the whole mental model: a slice has to stay intact, GKE provisions and labels it, and Ray Core reserves it as a unit so you never hand-write placement code. Everything you actually build sits on top of that and reuses it.

In Part 2, we will explore how you can use Ray AI libraries on TPU for serving LLMs with vLLM, feeding slices with Ray Data and training with JaxTrainer.

Additional resources

For now, thanks for reading! And if you have any additional questions or feedback, feel free to reach out on socials (LinkedIn, X).

Happy building!



Source_link

READ ALSO

Pixel 11 Pro launches with the best Gemini Intelligence, yet

You can now turn off Google Gemini’s visible watermarks

Related Posts

Pixel 11 Pro launches with the best Gemini Intelligence, yet
Google Marketing

Pixel 11 Pro launches with the best Gemini Intelligence, yet

August 15, 2026
You can now turn off Google Gemini’s visible watermarks
Google Marketing

You can now turn off Google Gemini’s visible watermarks

August 15, 2026
Pixel 11’s Magic Capture helps you take better photos
Google Marketing

Pixel 11’s Magic Capture helps you take better photos

August 15, 2026
Official Pixel 11 wireless charger stand drops to lowest-ever $35
Google Marketing

Official Pixel 11 wireless charger stand drops to lowest-ever $35

August 15, 2026
Google’s best new camera feature is only for the Pixel 11 series
Google Marketing

Google’s best new camera feature is only for the Pixel 11 series

August 15, 2026
HeyGen x Google Cloud: Bringing Avatar IV to TPUs
Google Marketing

HeyGen x Google Cloud: Bringing Avatar IV to TPUs

August 14, 2026
Next Post
How AI Is Changing Search Rankings

How AI Is Changing Search Rankings

POPULAR NEWS

Trump ends trade talks with Canada over a digital services tax

Trump ends trade talks with Canada over a digital services tax

June 28, 2025
15 Trending Songs on TikTok in 2025 (+ How to Use Them)

15 Trending Songs on TikTok in 2025 (+ How to Use Them)

June 18, 2025
Communication Effectiveness Skills For Business Leaders

Communication Effectiveness Skills For Business Leaders

June 10, 2025
Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

November 4, 2025
App Development Cost in Singapore: Pricing Breakdown & Insights

App Development Cost in Singapore: Pricing Breakdown & Insights

June 22, 2025

EDITOR'S PICK

Mistral AI Releases Mistral Small 4: A 119B-Parameter MoE Model that Unifies Instruct, Reasoning, and Multimodal Workloads

Mistral AI Releases Mistral Small 4: A 119B-Parameter MoE Model that Unifies Instruct, Reasoning, and Multimodal Workloads

March 17, 2026
What is AISO & How Does it Work?

What is AISO & How Does it Work?

May 13, 2026
Take Charge of AI: Mastering the Art of Prompt Engineering

Take Charge of AI: Mastering the Art of Prompt Engineering

June 11, 2025
How to reach a wider audience in 2025

How to reach a wider audience in 2025

August 2, 2025

About

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow us

Categories

  • Account Based Marketing
  • Ad Management
  • Al, Analytics and Automation
  • Brand Management
  • Channel Marketing
  • Digital Marketing
  • Direct Marketing
  • Event Management
  • Google Marketing
  • Marketing Attribution and Consulting
  • Marketing Automation
  • Mobile Marketing
  • PR Solutions
  • Social Media Management
  • Technology And Software
  • Uncategorized

Recent Posts

  • This Beautifully Weird Necklace Is Secretly a USB Drive
  • Pixel 11 Pro launches with the best Gemini Intelligence, yet
  • CRM & In-App Messaging Automation to Boost Retention
  • Climb Scary Squishy Dumpling Tower Code Roblox
  • About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions