Open-source AI models for music and sound

Audio models create music, jingles and sound effects from a text prompt, and some can split a track into stems. That gives you background audio for videos, podcasts and ads without stock libraries. The license matters most here: check whether commercial use of the output is allowed.

22 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Music and soundGGUF2025–2026

YuE

M-A-P and HKUST · China

Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.

  • Songs and jingles from lyrics
  • Demo versions of tracks
  • Music for videos
Sizes
0.5B – 7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Music and sound2026

HeartMuLa

HeartMuLa Team · not disclosed

An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.

  • Songs and jingles from lyrics
  • Music for videos
  • Transcribing song lyrics
Sizes
3B
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and soundGGUF2022–2026

MERT

m-a-p (Multimodal Art Projection) · UK / China

A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.

  • Automatic tagging of a music catalog
  • Finding similar tracks
  • Detecting genre, mood and tempo
Sizes
95M – 632M
Hardware
from: Laptop
Non-commercial onlyDetails
VideoGGUF2025–2026

MAGI

Sand AI · China

Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.

  • Long videos with continuation
  • Video with sound
  • Animating images
Sizes
4.5B – 114B-A6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice: speakers and sound2022–2026

UVR / MDX-Net / RoFormer (разделение звука)

Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community

A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.

  • Clean vocals from a recording with music
  • Backing tracks and stems for karaoke
  • Removing background music and noise from videos
Sizes
from tens to hundreds of millions of parameters
Hardware
from: Laptop
Commercial use with conditionsDetails
Video2024–2026

LTX-Video / LTX-2

Lightricks · Israel

A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.

  • Ad videos with sound
  • Video from a product photo
  • Voiced scenes for social media
Sizes
2B – 22B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Voice: speakers and sound2022–2026

Demucs

Meta AI, then Alexandre Défossez · France

A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.

  • Separating vocals from music in a recording
  • Backing tracks and karaoke stems
  • Cleaning speech in videos with background music
Sizes
tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundRUGGUF2025–2026

ACE-Step

ACE Studio and StepFun · China

Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.

  • Songs and jingles for ads
  • Background music for videos
  • Demo versions of tracks
Sizes
about 2B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundGGUF2024–2026

Stable Audio

Stability AI · UK

Generates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.

  • Sound effects for videos and games
  • Background music and jingles
  • Interface sounds
Sizes
about 0.5B – 2.3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Music and sound2025–2026

ThinkSound / PrismAudio

Alibaba Tongyi (FunAudioLLM) · China

Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.

  • Audio for video based on the scene
  • Editing individual sounds in a track
  • Sound effects from a description
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2026

daVinci-MagiHuman

SII-GAIR and Sand.ai · China

Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.

  • Presenter video from a script
  • Ad videos with a talking character
  • Training videos with a narrator
Sizes
15B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2026

MOVA

OpenMOSS / MOSI · China

Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.

  • Short clips with speech and sound from a description
  • Ad scenes with dialogue
  • Video prototypes for storyboards
Sizes
32B-A18B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice: speakers and sound2025

SAM Audio

Meta · USA

A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.

  • Isolating one person's voice from a noisy recording
  • Removing unwanted sound from a video
  • Splitting a recording into separate sound sources
Sizes
small, base, large (5 to 15 GB of weights)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2025

Perception Encoder (PE)

Meta · USA

Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.

  • Search photos and videos by description
  • Catalog labeling and tagging
  • Search across audio and video (PE-AV)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Video2025

Ovi

Character.AI · USA

Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.

  • Short clips with talking characters
  • Animating an image with voice-over
  • Ad scene prototypes
Sizes
11B
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and sound2025

DiffRhythm

ASLP-lab (Northwestern Polytechnical University) · China

Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.

  • Songs and jingles from lyrics
  • Music for videos
  • Demo versions of tracks
Sizes
about 1.1B
Hardware
from: Laptop
Commercial use allowedDetails
Music and sound2025

HunyuanVideo-Foley

Tencent Hunyuan · China

Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.

  • Foley and sound effects for video
  • Sound for AI-generated ads
  • Sound design for short videos
Sizes
not stated on the model card (weights about 10 GB; XL version with memory offloading)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Music and sound2024

MMAudio

University of Illinois and Sony AI · USA / Japan

Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.

  • Sound effects for silent video
  • Sound for clips from AI generators
  • Draft sound design for editing
Sizes
size not stated on the model card
Hardware
from: Laptop
Non-commercial onlyDetails
Music and sound2024

MuQ / MuQ-MuLan

Tencent AI Lab · China

The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.

  • Searching music by text description
  • Tagging tracks by genre and mood
  • Finding similar music
Sizes
300M – 700M
Hardware
from: Laptop
Non-commercial onlyDetails
Music and sound2023–2024

MusicGen (AudioCraft)

Meta · USA

Generates instrumental music from a text description or a sample melody. One of the first open models of its kind.

  • Draft music sketches
  • Music for video prototypes
  • Research
Sizes
300M – 3.3B
Hardware
from: Laptop
Non-commercial onlyDetails
Music and soundNot maintained2023

CLAP (LAION)

LAION · Germany

CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.

  • Search sounds and music by description
  • Automatic tags for an audio library
  • Recognizing sound types (siren, breaking glass, voice)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundNot maintained2022

AST (Audio Spectrogram Transformer)

MIT · USA

A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.

  • Sound event recognition
  • Tagging an audio archive
  • Detecting alarm sounds
Sizes
about 87M
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment