YuE
Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.
- Songs and jingles from lyrics
- Demo versions of tracks
- Music for videos
- Sizes
- 0.5B – 7B
- Hardware
- from: 1 GPU
Audio models create music, jingles and sound effects from a text prompt, and some can split a track into stems. That gives you background audio for videos, podcasts and ads without stock libraries. The license matters most here: check whether commercial use of the output is allowed.
Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.
An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.
A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.
A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.
A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.
Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.
Generates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.
Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.
A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.
Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.
Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.
The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.
Generates instrumental music from a text description or a sample melody. One of the first open models of its kind.
CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.
A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.