Video2025–2026
Alibaba · China
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
- Short promo videos
- Animating product photos
- Videos for social media
- Sizes
- 1,3B – 14B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Meituan · China
Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.
- News or course presenter videos
- Promo videos with a talking character
- Singing and voice-over
- Sizes
- based on LongCat-Video 13.6B
- Hardware
- from: 1 GPU
Avatars2024–2026
Fudan University · China
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
- Presenter video from a photo and audio
- Long training videos
- Live avatar
- Sizes
- about 1B – 5B
- Hardware
- from: 1 GPU
Avatars2026
SII-GAIR and Sand.ai · China
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
- Presenter video from a script
- Ad videos with a talking character
- Training videos with a narrator
- Sizes
- 15B
- Hardware
- from: 1 GPU
VideoGGUF2025–2026
Skywork AI (Kunlun Tech) · China
Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.
- Long videos with continuation
- Video with one character from a reference
- Talking avatar from a voice
- Sizes
- 1.3B – 19B
- Hardware
- from: 1 GPU
Avatars2024–2026
Ant Group · China
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
- Presenter video from a photo and voice
- Avatar with gestures for presentations
- Voiced characters
- Sizes
- up to 1.3B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Alibaba (Quark) · China
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
- Live avatar for customer dialogue
- Endless broadcasts with a presenter
- Interactive characters
- Sizes
- 14B
- Hardware
- from: 1 GPU
Avatars2025
Meituan · China
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
- Video dubbing with matched facial expressions
- Dialogue between two characters from audio
- Long videos with a presenter
- Sizes
- 14B
- Hardware
- from: 1 GPU
Avatars2025
ByteDance · China
Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.
- Dubbing videos into another language with lip sync
- Editing lines in finished video without reshooting
- Talking avatars for training courses
- Sizes
- requires 8–18 GB of VRAM
- Hardware
- from: Laptop
AvatarsGGUF2025
Tencent · China
Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.
- Presenter video from a photo and audio
- Scenes with several speakers
- Cartoon characters
- Sizes
- about 13B
- Hardware
- from: 1 GPU
Avatars2024–2025
Tencent Music (Lyra Lab) · China
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
- Dubbing video into another language
- Live avatar in a video chat
- Editing speech in a finished video
- Sizes
- under 1B
- Hardware
- from: Laptop
AvatarsNot maintained2024
Kuaishou (Kling) · China
Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.
- Animating portraits
- Transferring an actor's expressions to a character
- Mascot animation
- Sizes
- under 1B
- Hardware
- from: Laptop
AvatarsNot maintained2023
Xi'an Jiaotong University and Tencent AI Lab · China
An older lightweight talking-head model: one photo plus audio becomes a video. Runs on weak hardware, but quality is noticeably below newer models.
- Talking photo for greetings
- Simple voiced avatars
- Sizes
- under 1B
- Hardware
- from: Laptop
AvatarsNot maintained2020
IIIT Hyderabad · India
The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.
- Quick dubbing tests
- Comparison with newer lip-sync models
- Educational and research projects
- Sizes
- small model, 96-pixel face
- Hardware
- from: Laptop