Video generation models as world simulators

Practical AI: Tools, Models & Frameworkstransformer

What changed

We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes.

Why it matters

A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.

How it compares

Related prior coverage to compare against:

  • Fine-tune video and image models at scale with NVIDIA NeMo Automodel and ๐Ÿค— Diffusers
  • Sora 2 System Card
  • Fine-tuning GPT-3 to scale video creation

Sources