Video diffusion models are powerful tools for generating high-quality videos, but they’ve traditionally been slow and resource-intensive. Google’s latest advancements in optimizing these models on TPUs (Tensor Processing Units) make video generation faster and more efficient. If you’re just starting to explore AI-powered video creation, this development simplifies the process and reduces the technical barriers to entry.
Why Video Diffusion Models Are Slow
Video diffusion models generate videos by iteratively removing noise from a random starting point. This process involves two main bottlenecks:
- Many Denoising Steps: Each frame requires multiple steps to refine, which adds up quickly.
- High Computational Cost: Each step involves complex calculations, especially for high-resolution videos.
For example, generating an 81-frame HD (720p) video can involve sequence lengths ranging from 50,000 to 400,000 tokens. Scaling up to 2K (1440p) quadruples the sequence length, making the process even slower.
Tip
If you’re working with video diffusion models, start with lower resolutions to reduce computational load.
The Role of Attention in Video Diffusion
Self-attention is a key component of video diffusion models, but it’s also a major source of latency. Full attention computes interactions between every pair of tokens, which scales quadratically with sequence length. For instance, in a 720p video, attention accounts for 55.5% of per-layer latency. At 2K, this jumps to 88.2%.

Speeding Up Attention with Sparsity
Google’s solution leverages sparse attention, which skips unimportant interactions to reduce computation. Video diffusion models exhibit structured attention patterns:
- Spatial Heads: Focus on patches within the same or nearby frames.
- Temporal Heads: Track a small spatial region across many frames.
By dynamically profiling and routing attention heads to either spatial or temporal masks, Google’s Sparse VideoGen (SVG) approach preserves quality while reducing latency.

Implementing Sparse Attention on TPUs
Google’s optimizations include custom JAX and Pallas Splash Attention kernels. These kernels skip empty tiles, avoid masking full tiles, and align the sparse mask with hardware execution. Key improvements include:
| Optimization Step | Latency (ms) | Speedup Over Dense |
|---|---|---|
| Naive sparse traversal (B1) | 96.37 | -22% |
| Full/boundary tile specialization (B2) | 54.12 | 31% |
| Tile-aligned sparse traversal (B3) | 32.76 | 2.40x |
These optimizations demonstrate that sparse attention requires both a suitable mask and an efficient execution strategy.
Dynamic Routing and Token Layouts
To handle dynamic masks and token layouts, Google introduced a pipeline that:
- Profiles attention heads at inference time.
- Routes each head to a spatial or temporal mask.
- Permutes tokens for efficient temporal access.
This approach reduces latency by 38% in distributed inference scenarios.
What This Means for Beginners
Google’s advancements make video diffusion models more practical for beginners by:
- Reducing Latency: Faster generation means quicker experimentation and iteration.
- Lowering Costs: Efficient use of TPUs reduces hardware requirements.
- Simplifying Workflows: Dynamic routing and sparse attention minimize manual tuning.
Want to try all of this hands-on? Start with the free Vibe Coding 101 course.
Source
Based on Google Developers Blog’s announcement, “Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs”. Written for people learning to build with these tools.