LatentSync
Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.
- Developer
- ByteDance, China
- First release
- Jan 2025
- Latest release
- Jun 2025
- Sizes
- requires 8–18 GB of VRAM
- License
- Commercial use with conditionsWeights are OpenRAIL++ (commercial use with usage restrictions), code is Apache 2.0; the built-in InsightFace face detector is non-commercial and must be replaced for business use
- Running
- On your own serverNeeds a GPU
- Industries
- Media and production, Marketing and content, Education
What it does
- Dubbing videos into another language with lip sync
- Editing lines in finished video without reshooting
- Talking avatars for training courses
Where it is used
Hardware requirements
Versions
- LatentSync 1.6 (512 пикселей)
- LatentSync 1.5
- LatentSync 1.0
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can LatentSync be used in a commercial project?
With conditions. License: Weights are OpenRAIL++ (commercial use with usage restrictions), code is Apache 2.0; the built-in InsightFace face detector is non-commercial and must be replaced for business use. Restrictions vary — region, company revenue, attribution requirements. Have a lawyer check the terms before a commercial launch.
What hardware does LatentSync need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does LatentSync support Russian?
Language does not matter for this model: it does not work with text.
Where can I download LatentSync and what does it cost?
The LatentSync weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
DetailsAvatarsWav2LipIIIT Hyderabad · IndiaNon-commercial onlyThe classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.
DetailsAvatarsMultiTalk / InfiniteTalkMeituan · ChinaCommercial use allowedDubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
DetailsSource: github.com/bytedance/LatentSync. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


