Video generation models as world simulators
What changed
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- Fine-tune video and image models at scale with NVIDIA NeMo Automodel and ๐ค Diffusers
- Sora 2 System Card
- Fine-tuning GPT-3 to scale video creation
Sources
- Video generation models as world simulators (openai-blog)primary