daVinci-MagiHuman
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
- Developer
- SII-GAIR and Sand.ai, China
- First release
- Mar 2026
- Latest release
- Mar 2026
- Sizes
- 15B
- License
- Commercial use allowedApache 2.0
- Russian
- Not supported
- Running
- On your own serverNeeds a GPU
- Industries
- Marketing and content, Media and production, Education
What it does
- Presenter video from a script
- Ad videos with a talking character
- Training videos with a narrator
Where it is used
Hardware requirements
Versions
- daVinci-MagiHuman (Base, Distill, 256p–1080p)
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can daVinci-MagiHuman be used in a commercial project?
Yes. License: Apache 2.0. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does daVinci-MagiHuman need?
At minimum: One GPU with 16–80 GB — mid-size versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does daVinci-MagiHuman support Russian?
No. The model card lists its languages and Russian is not among them.
Where can I download daVinci-MagiHuman and what does it cost?
The daVinci-MagiHuman weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
DetailsVideoOviCharacter.AI · USACommercial use allowedGenerates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
DetailsAvatarsMultiTalk / InfiniteTalkMeituan · ChinaCommercial use allowedDubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
DetailsSource: huggingface.co/GAIR/daVinci-MagiHuman. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


