showlab/Awesome-Video-Diffusion
A curated list of recent diffusion models for video generation, editing, and various other applications.
README
Awesome Video Diffusion
A curated list of recent diffusion models for video generation, editing, restoration, understanding, nerf, etc.
(Source: Make-A-Video, Tune-A-Video, and Fate/Zero.)
Table of Contents
- Open-source Toolboxes and Foundation Models
- Evaluation Benchmarks and Metrics
- Commercial Product
- Video Generation
- Efficient Video Generation
- Controllable Video Generation
- Character Customization
- Motion Customization
- Long Video / Film Generation
- Video Generation with 3D/Physical Prior
- Video Editing
- Human or Subject Motion
- Video Enhancement and Restoration
- Audio Synthesis for Video
- Talking Head Generation
- Reinforcement Learning for Video Generation
- Policy Learning
- Virtual Try-On
- 3D
- 4D
- Game Generation
- AI Safety
- Rendering with Virtual Engine
- Open-World Model
- Video Understanding
- Healthcare and Biology
- Other Applications
- Code-rendered Video Generation
Open-source Toolboxes and Foundation Models
+ Omni-Rewriter
Open agentic prompt-expansion harness for image/video generation (schema → validate → bounded repair → dialect render; H3/Seedance/Seedream/Qwen-Image). Expand ≠ generate.
+ NanoI2V
A step-by-step teaching series for building an Image-to-Video model from scratch in PyTorch, covering 3D VAEs, DiT, Flow Matching, RoPE, and conditioning.
+ Helios: Real Real-Time Long Video Generation Model
+ Waver: Wave Your Way to Lifelike Video Generation
+ FastVideo: A unified inference and post-training framework for accelerated video generation
+ LightX2V: Light Video Generation Inference Framework
+ Wan-Video
+ SkyReels-V2
+ DiffSynth-Studio
+ Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
+ Cosmos
+ HunyuanVideo: A Systematic Framework For Large Video Generative Models
+ Allegro
+ Mochi 1
+ Movie Gen: A Cast of Media Foundation Models
+ Pyramidal Flow Matching for Efficient Video Generative Modeling
+ Show-1
+ I2VGen-XL (image-to-video / video-to-video)
-9cf)
-9cf)
+ text-to-video-synthesis-colab
+ VideoCrafter: A Toolkit for Text-to-Video Generation and Editing
+ ModelScope (Text-to-video synthesis)
+ Diffusers (Text-to-video synthesis)
+ Wunjo CE (Video Generation and Editing)
Evaluation Benchmarks and Metrics
+ LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference (Oct., 2025)
+ Stable Cinemetrics: Structured Taxonomy and Evaluation for Professional Video Generation (Sep., 2025)
+ OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation (Mar., 2025)
+ VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness (Mar., 2025)
+ Impossible Videos (Mar., 2025)
+ MEt3R: Measuring Multi-View Consistency in Generated Images (Jan., 2025)
+ Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation (Dec., 2024)
+ Evaluation Agent, Efficient and Promptable Evaluation Framework for Visual Generative Models (Dec., 2024)
+ Frechet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos (Jun., 2024)
+ T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation (Jun., 2024)
+ ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation (NeurIPS, 2024)
+ PEEKABOO: Interactive Video Generation via Masked-Diffusion (CVPR, 2024)
+ T2VScore: Towards A Better Metric for Text-to-Video Generation (Jan., 2024)
+ StoryBench: A Multifaceted Benchmark for Continuous Story Visualization (NeurIPS, 2023)
+ VBench: Comprehensive Benchmark Suite for Video Generative Models (Nov., 2023)
+ FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation (Nov., 2023)
+ EvalCrafter: Benchmarking and Evaluating Large Video Generation Models (Oct., 2023)
+ Evaluation of Text-to-Video Generation Models: A Dynamics Perspective (Jul., 2024)
+ VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models (May., 2024)
+ Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers (CVPR, 2024)
+ ReLight My NeRF: A Dataset for Novel View Synthesis and Relighting of Real World Objects (CVPR, 2023)
Commercial Product
+ Dream Machine (Luma AI)
+ Riffkit
+ [SEE