TY - GEN
T1 - TI2V-Zero
T2 - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
AU - Ni, Haomiao
AU - Egger, Bernhard
AU - Lohit, Suhas
AU - Cherian, Anoop
AU - Wang, Ye
AU - Koike-Akino, Toshiaki
AU - Huang, Sharon X.
AU - Marks, Tim K.
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., 'a woman is drinking water.'). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a 'repeat-and-slide' strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation.
AB - Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., 'a woman is drinking water.'). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a 'repeat-and-slide' strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation.
UR - https://www.scopus.com/pages/publications/85213994352
UR - https://www.scopus.com/pages/publications/85213994352#tab=citedBy
U2 - 10.1109/CVPR52733.2024.00861
DO - 10.1109/CVPR52733.2024.00861
M3 - Conference contribution
AN - SCOPUS:85213994352
SN - 9798350353006
T3 - Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
SP - 9015
EP - 9025
BT - Proceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
PB - IEEE Computer Society
Y2 - 16 June 2024 through 22 June 2024
ER -