Current diffusion fashions allow high-quality video era, however undergo from sluggish runtimes. The massive transformer-based backbones utilized in these fashions are bottlenecked by spatiotemporal consideration. On this paper, we establish {that a} important fraction of token-to-token connections constantly yield negligible scores throughout varied inputs, and their patterns usually repeat throughout queries. Thus, the eye computation in these circumstances could be skipped with little to no impact on the outcome. This statement continues to carry for connections amongst native token blocks. Motivated by this, we introduce CalibAtt, a training-free methodology that accelerates video era through calibrated sparse consideration. CalibAtt performs an offline calibration cross that identifies block-level sparsity and repetition patterns which can be steady throughout inputs, and compiles these patterns into optimized consideration operations for every layer, head, and diffusion timestep. At inference time, we compute the chosen input-dependent connections densely, and skip the unselected ones in a hardware-efficient method. In depth experiments on Wan 2.1 14B, Mochi 1, and few-step distilled fashions at varied resolutions present that CalibAtt achieves as much as 1.58× end-to-end speedup, outperforming present training-free strategies whereas sustaining video era high quality and text-video alignment.
- † Tel Aviv College
- ** Work carried out whereas at Apple
