Open-source alternative to Synthesia and HeyGen: video avatars on your own server

Synthesia and HeyGen are used for training videos, instructions and internal news: build an avatar once, then change only the script. Open models do the same on your side: a photo or short clip plus a voice track becomes a talking presenter, and separate models re-sync the lips of existing footage to a new audio track. Scripts and employees' faces never reach a third-party service, and you can ship as many videos as you need, including several language versions over one source clip. The limits are visible: an avatar is convincing in a calm talking-head shot, while gestures, head turns and long runtimes are weaker, and you need a GPU plus a separate voice track. The legal side matters more than the technical one here: a person's face and voice may only be used with their written consent.

Updated 22 Sep 2026Find a model in 4 questions

What to use instead

Avatars2024–2026

Hallo

Fudan University · China

A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.

  • Presenter video from a photo and audio
  • Long training videos
  • Live avatar
Sizes
about 1B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2024–2026

EchoMimic

Ant Group · China

Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.

  • Presenter video from a photo and voice
  • Avatar with gestures for presentations
  • Voiced characters
Sizes
up to 1.3B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2025

MultiTalk / InfiniteTalk

Meituan · China

Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.

  • Video dubbing with matched facial expressions
  • Dialogue between two characters from audio
  • Long videos with a presenter
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
AvatarsGGUF2025–2026

Live Avatar

Alibaba (Quark) · China

A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.

  • Live avatar for customer dialogue
  • Endless broadcasts with a presenter
  • Interactive characters
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2024–2025

MuseTalk

Tencent Music (Lyra Lab) · China

Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.

  • Dubbing video into another language
  • Live avatar in a video chat
  • Editing speech in a finished video
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails

These five are listed as commercially usable (MIT and Apache 2.0), whereas the classic Wav2Lip is non-commercial and LatentSync ships with the non-commercial InsightFace detector, which a business has to replace. Written consent for the use of a person's face and voice is mandatory.

Other alternatives

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment