Open-source AI models for video generation

Video models create short clips from a text prompt or animate a still image. Teams use them for ad creatives, social media content and scene prototypes. These models are hardware-hungry, so check GPU memory, clip length and resolution, and the license terms up front.

37 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
VideoGGUF2025–2026

MAGI

Sand AI · China

Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.

  • Long videos with continuation
  • Video with sound
  • Animating images
Sizes
4.5B – 114B-A6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025–2026

SANA-Video

NVIDIA · USA

NVIDIA's lightweight, fast video model. Produces 720p clips on a single GPU; a 4-step version enables quick generation.

  • Quick clips for social media
  • Bulk video generation
  • Video from an image
Sizes
2B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025–2026

Wan

Alibaba · China

Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).

  • Short promo videos
  • Animating product photos
  • Videos for social media
Sizes
1,3B – 14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationRU2022–2026

Kandinsky

Sber (Kandinsky Lab) · Russia

Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.

  • Images from Russian-language descriptions
  • Short promo videos from text or a photo
  • Instruction-based image editing
Sizes
2B – 19B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2024–2026

LTX-Video / LTX-2

Lightricks · Israel

A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.

  • Ad videos with sound
  • Video from a product photo
  • Voiced scenes for social media
Sizes
2B – 22B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2026

MiniMax H3 (Hailuo)

MiniMax · China

Open weights of MiniMax's Hailuo video model. A large 33B model that makes video from text and images, but needs several server GPUs.

  • Cinematic ad videos
  • Video from text and images
  • Complex scenes with motion
Sizes
33B + 32B encoder
Hardware
from: Cluster
Commercial use with conditionsDetails
Robotics2025–2026

NVIDIA Cosmos

NVIDIA · USA

"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.

  • Synthetic video for training robots and self-driving vehicles
  • Testing scenarios in simulation
  • Robot control (Policy versions)
Sizes
2B – 65B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2025–2026

SCAIL

Zhipu AI (Z.ai) and Tsinghua University · China

Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.

  • Transferring an actor's motion to a character
  • Replacing a character in a finished video
  • Animating mascots and illustrations
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Photo editing2023–2026

BRIA RMBG

BRIA AI · Israel

BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.

  • Cutting products out onto a white background
  • Staff and expert photos without background
  • Background removal in video
Sizes
44M – 220M
Hardware
from: Laptop
Non-commercial onlyDetails
3D2025–2026

HY-World (HunyuanWorld)

Tencent · China

Generates whole 3D worlds and scenes from text or an image that you can walk through. The second version builds a scene from video and photos.

  • 3D scenes for games and virtual tours
  • Backgrounds and environments for video production
  • Draft locations for simulations
Sizes
set of several models
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2025–2026

Matrix-Game

Skywork AI (Kunlun) · China

An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.

  • Game world prototypes without an engine
  • Interactive demos and simulations
  • Generating data to train agents
Sizes
1.8B – 17B
Hardware
from: 1 GPU
Commercial use allowedDetails
Photo editing2025–2026

MatAnyone

S-Lab, Nanyang Technological University · Singapore

Cuts a person out of video with a precise alpha mask, including hair and edges, without a green screen. Needs a first-frame mask, for example from SAM.

  • Background replacement in video without chroma key
  • Cutting out a person for editing and effects
  • Preparing videos for advertising and social media
Sizes
about 35M
Hardware
from: Laptop
Non-commercial onlyDetails
Music and sound2025–2026

ThinkSound / PrismAudio

Alibaba Tongyi (FunAudioLLM) · China

Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.

  • Audio for video based on the scene
  • Editing individual sounds in a track
  • Sound effects from a description
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2026

daVinci-MagiHuman

SII-GAIR and Sand.ai · China

Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.

  • Presenter video from a script
  • Ad videos with a talking character
  • Training videos with a narrator
Sizes
15B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2025–2026

SkyReels

Skywork AI (Kunlun Tech) · China

Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.

  • Long videos with continuation
  • Video with one character from a reference
  • Talking avatar from a voice
Sizes
1.3B – 19B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
VideoGGUF2026

MOVA

OpenMOSS / MOSI · China

Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.

  • Short clips with speech and sound from a description
  • Ad scenes with dialogue
  • Video prototypes for storyboards
Sizes
32B-A18B
Hardware
from: 1 GPU
Commercial use allowedDetails
Deepfake detection2024–2025

VideoSeal

Meta · USA

A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.

  • Marking video created or processed by AI
  • Finding your own mark in re-uploaded clips
  • Protecting ad materials from being reused as someone else's
Sizes
a mark of 96 to 1024 bits
Hardware
from: Laptop
Commercial use allowedDetails
VideoGGUF2024–2025

HunyuanVideo

Tencent · China

Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.

  • Video from a text script
  • Animating images
  • Base for fine-tuning your own video models
Sizes
8.3B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2025

Ovi

Character.AI · USA

Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.

  • Short clips with talking characters
  • Animating an image with voice-over
  • Ad scene prototypes
Sizes
11B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2025

LongCat-Video

Meituan · China

A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.

  • Long videos
  • Video from a photo
  • Continuing an existing video
Sizes
13.6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and sound2025

HunyuanVideo-Foley

Tencent Hunyuan · China

Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.

  • Foley and sound effects for video
  • Sound for AI-generated ads
  • Sound design for short videos
Sizes
not stated on the model card (weights about 10 GB; XL version with memory offloading)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2025

Hunyuan-GameCraft

Tencent Hunyuan · China

Turns a single image into a controllable game-scene video: the camera moves on keyboard commands. Minimum 24 GB of GPU memory, 80 GB recommended.

  • Interactive video prototypes of game locations
  • Camera walkthrough videos of a scene
  • Level demos for pitches
Sizes
based on HunyuanVideo
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Deepfake detection2023–2025

DeepfakeBench

The Chinese University of Hong Kong, Shenzhen (SCLBD) · China

Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.

  • First-pass check of a submitted video or selfie
  • Comparing several detectors on your own data
  • Fine-tuning a detector for your own flow of applications
Sizes
Xception- and EfficientNet-class detectors, tens of millions of parameters
Hardware
from: Laptop
Non-commercial onlyDetails
Photo editing2025

SeedVR / SeedVR2

ByteDance Seed · China

ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.

  • Upscaling photos and video to 2K–4K
  • Restoring old videos and photos
  • Enhancing user photos before publishing
Sizes
3B – 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2025

LatentSync

ByteDance · China

Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.

  • Dubbing videos into another language with lip sync
  • Editing lines in finished video without reshooting
  • Talking avatars for training courses
Sizes
requires 8–18 GB of VRAM
Hardware
from: Laptop
Commercial use with conditionsDetails
Virtual try-on2024–2025

CatVTON / CatV2TON

Sun Yat-sen University and Pixocial · China

A lightweight try-on model that runs on a regular GPU. There is a mask-free version and CatV2TON, which also tries clothes on in video.

  • Trying a garment on a customer's photo
  • Draft product cards on a model
  • Try-on in a short video
Sizes
899M
Hardware
from: Laptop
Non-commercial onlyDetails
Video2024–2025

Open-Sora

HPC-AI Tech · Singapore

A fully open video generation project: weights, code and training recipe. Version 2.0 at 11B makes video from text and from an image.

  • Video from a text description
  • Animating images
  • Training your own video model
Sizes
up to 11B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025

Step-Video

StepFun · China

A large 30B video model producing clips of up to 204 frames. Needs server hardware, but is open under MIT.

  • Video from a description
  • Animating images
Sizes
30B
Hardware
from: Cluster
Commercial use allowedDetails
Photo editing2024–2025

BEN2

Prama LLC · USA

A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.

  • Cutting out products and people from photos
  • Background removal in video
  • Preparing photos for a catalog
Sizes
about 95M
Hardware
from: Laptop
Commercial use allowedDetails
Music and sound2024

MMAudio

University of Illinois and Sony AI · USA / Japan

Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.

  • Sound effects for silent video
  • Sound for clips from AI generators
  • Draft sound design for editing
Sizes
size not stated on the model card
Hardware
from: Laptop
Non-commercial onlyDetails
Video2024

CogVideoX

Zhipu AI (Z.ai) and Tsinghua University · China

A 2–5B video model that runs on a single gaming GPU. A popular base for research and add-ons.

  • Short clips from text
  • Animating images
  • Video fine-tuning experiments
Sizes
2B – 5B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
VideoGGUF2024

Mochi

Genmo · USA

An open 10B video model with realistic motion. At release it was among the strongest open models; no updates now.

  • Video from a description
  • Short ad scenes
Sizes
10B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2022–2024

RIFE (Practical-RIFE)

hzwer (Zhewei Huang) and co-authors · China

Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.

  • Increasing video frame rate
  • Smooth slow-motion video
  • Smoothing clips from AI generators
Sizes
lightweight model (size not stated on the model card)
Hardware
from: Laptop
Commercial use allowedDetails
VideoNot maintained2023–2024

AnimateDiff

Shanghai AI Lab and CUHK · China

A module that brings Stable Diffusion image models to life, turning them into short animations. One of the first open video technologies.

  • Short animations in brand style
  • Animated covers and banners
  • Animated stickers
Sizes
motion module on top of SD 1.5 / SDXL
Hardware
from: Laptop
Commercial use allowedDetails
VideoNot maintained2023

Stable Video Diffusion

Stability AI · UK

Stability AI's first open video model: turns a photo into a 2–4 second clip. Now behind newer models in quality.

  • Animating product photos
  • Short video intros
  • Animating illustrations
Sizes
about 1.5B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Photo editingNot maintained2022–2023

InSPyReNet / transparent-background

Taehoon Kim (POSTECH) · South Korea

A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.

  • Batch background removal from photos
  • Replacing the background with a color or blur
  • Background removal in video
Sizes
small (based on Swin-B)
Hardware
from: Laptop
Commercial use allowedDetails
AvatarsNot maintained2020

Wav2Lip

IIIT Hyderabad · India

The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.

  • Quick dubbing tests
  • Comparison with newer lip-sync models
  • Educational and research projects
Sizes
small model, 96-pixel face
Hardware
from: Laptop
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment