MMAudio
Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.
- Developer
- University of Illinois and Sony AI, USA / Japan
- First release
- Dec 2024
- Latest release
- Dec 2024
- Sizes
- size not stated on the model card
- License
- Non-commercial onlyWeights — CC-BY-NC 4.0 (non-commercial), code — MIT
- Running
- On your own serverNeeds a GPU
- Industries
- Media and production
What it does
- Sound effects for silent video
- Sound for clips from AI generators
- Draft sound design for editing
Where it is used
Hardware requirements
Versions
- MMAudio large 44k v2
- MMAudio (small, medium, large)
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can MMAudio be used in a commercial project?
No. License: Weights — CC-BY-NC 4.0 (non-commercial), code — MIT. A commercial product needs a different model or a separate agreement with the rights holder.
What hardware does MMAudio need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does MMAudio support Russian?
Language does not matter for this model: it does not work with text.
Where can I download MMAudio and what does it cost?
The MMAudio weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.
DetailsMusic and soundThinkSound / PrismAudioAlibaba Tongyi (FunAudioLLM) · ChinaCommercial use allowedGenerates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
DetailsMusic and soundStable AudioStability AI · UKCommercial use with conditionsGenerates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.
DetailsSource: github.com/hkchengrex/MMAudio. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


