StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text #7731

clarencechen · 2024-04-21T08:39:56Z

Model/Pipeline/Scheduler description

Text-to-video diffusion models enable the generation of high-quality videos given text prompts, making it easy to create diverse and individual content. However, existing approaches mostly focus on short video generation (typically 16 or 24 frames), requiring hard cuts when naively extended to the case of long video synthesis. StreamingT2V, enables autoregressive generation of long videos of 80, 240, 600, 1200 or more frames with smooth transitions. The key components are:

A ControlNet-like module which conditions the current generation on frames extracted from the previous chunk, using a cross-attention mechanism to integrate its features into the UNet's skip residual features.
An IP-Adapter-like module which extracts high-level scene and object features from a fixed anchor frame in the first video chunk and is mixed into the prompt embedding features before executing spatial cross-attention.
A SDEdit-based video refinement stage with randomized chunk sampling of overlapped frames per denoising timestep.

Open source status

The model implementation is available.
The model weights are available (Only relevant if addition is not a scheduler).

Provide useful links for the implementation

Code: https://github.com/Picsart-AI-Research/StreamingT2V/-
Weights: https://huggingface.co/PAIR/StreamingT2V
Point of Contact Author: @hpoghos

DN6 · 2024-04-22T10:38:14Z

Sounds like a nice addition. I think we can open it up to the community to work on. Or would you like to work on it @clarencechen?

dg845 · 2024-09-15T23:06:38Z

Hi, I'd like to try working on this if the maintainers still think this is a good addition to the library :).

yiyixuxu · 2024-09-17T21:20:39Z

hi @dg845
it's been a while! Welcome back!
this one I think it'd go into the community folder, to begin with

a-r-r-o-w · 2024-11-20T00:15:22Z

Hi @dg845! We did indeed plan to support StreamingT2V but other things took priority. Would love to have this if you find time to PR - thanks! I'm familiar with the codebase so would love to be of help in any way

a-r-r-o-w added community-examples Good second issue contributions-welcome labels Aug 30, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text #7731

StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text #7731

clarencechen commented Apr 21, 2024 •

edited

Loading

DN6 commented Apr 22, 2024

dg845 commented Sep 15, 2024

yiyixuxu commented Sep 17, 2024

a-r-r-o-w commented Nov 20, 2024

StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text #7731

StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text #7731

Comments

clarencechen commented Apr 21, 2024 • edited Loading

Model/Pipeline/Scheduler description

Open source status

Provide useful links for the implementation

DN6 commented Apr 22, 2024

dg845 commented Sep 15, 2024

yiyixuxu commented Sep 17, 2024

a-r-r-o-w commented Nov 20, 2024

clarencechen commented Apr 21, 2024 •

edited

Loading