Hallo
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
- Presenter video from a photo and audio
- Long training videos
- Live avatar
- Sizes
- about 1B – 5B
- Hardware
- from: 1 GPU
Synthesia and HeyGen are used for training videos, instructions and internal news: build an avatar once, then change only the script. Open models do the same on your side: a photo or short clip plus a voice track becomes a talking presenter, and separate models re-sync the lips of existing footage to a new audio track. Scripts and employees' faces never reach a third-party service, and you can ship as many videos as you need, including several language versions over one source clip. The limits are visible: an avatar is convincing in a calm talking-head shot, while gestures, head turns and long runtimes are weaker, and you need a GPU plus a separate voice track. The legal side matters more than the technical one here: a person's face and voice may only be used with their written consent.
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
These five are listed as commercially usable (MIT and Apache 2.0), whereas the classic Wav2Lip is non-commercial and LatentSync ships with the non-commercial InsightFace detector, which a business has to replace. Written consent for the use of a person's face and voice is mandatory.
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.