Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs
To address the quadratic latency bottleneck of self-attention in high-resolution video diffusion models, developers implemented Sparse VideoGen (SVG) to dynamically route attention heads to highly structured spatial or temporal sparse masks. Translating this algorithmic sparsity into physical…
Sources
- T1Accelerating Spatio-Temporal Attention for Video Diffusion on TPUsGoogle — The Keyword / AI / Research / DeepMind / Developers