Video models create short clips from a text prompt or animate a still image. Teams use them for ad creatives, social media content and scene prototypes. These models are hardware-hungry, so check GPU memory, clip length and resolution, and the license terms up front.
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.
Synthetic video for training robots and self-driving vehicles
Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.
BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.
An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.
S-Lab, Nanyang Technological University · Singapore
Cuts a person out of video with a precise alpha mask, including hair and edges, without a green screen. Needs a first-frame mask, for example from SAM.
Background replacement in video without chroma key
Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
Marking video created or processed by AI
Finding your own mark in re-uploaded clips
Protecting ad materials from being reused as someone else's
Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
The Chinese University of Hong Kong, Shenzhen (SCLBD) · China
Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.
First-pass check of a submitted video or selfie
Comparing several detectors on your own data
Fine-tuning a detector for your own flow of applications
Sizes
Xception- and EfficientNet-class detectors, tens of millions of parameters
ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.
Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.
Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.
Increasing video frame rate
Smooth slow-motion video
Smoothing clips from AI generators
Sizes
lightweight model (size not stated on the model card)
A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.
The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.