Demucs or UVR: which tool splits vocals from music

Both do the same job: split a recording into vocals and everything else. Demucs is one clear model: MIT, tens of millions of parameters, splits a track into vocals, drums, bass and the rest, runs on a CPU, with version 4.1.0 released in 2026-07. UVR is not a model but a large community collection: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet, DTTNet, listed in the catalog as the quality leader for vocals among open solutions, with updates arriving constantly, the latest in 2026-08. The price is licensing: the code is MIT, but each set of weights carries its own terms and has to be checked individually. For a production pipeline Demucs is more predictable; for the best result on a specific recording it pays to try weights from UVR.

Comparison based on catalog data

ParameterDemucsUVR / MDX-Net / RoFormer (разделение звука)
CategoryVoice: speakers and sound, Music and soundVoice: speakers and sound, Music and sound
DeveloperMeta AI, then Alexandre Défossez, FranceCommunity: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep, International community
ReleasesDec 2022 – Jul 2026Dec 2022 – Aug 2026
Sizestens of millions of parametersfrom tens to hundreds of millions of parameters
HardwareLaptop, 1 GPULaptop, 1 GPU
Commercial useCommercial use allowedCommercial use with conditions
LicenseMITCode MIT; individual weights have different terms, check each model
RussianNot applicableNot applicable
OllamaNoNo
Without GPUYesYes
Tasks
  • Separating vocals from music in a recording
  • Backing tracks and karaoke stems
  • Cleaning speech in videos with background music
  • Clean vocals from a recording with music
  • Backing tracks and stems for karaoke
  • Removing background music and noise from videos

Choose Demucs if

  • You want one model under a clean MIT license, with no terms to untangle
  • Separation is built into a service and must behave the same on every file
  • You need all four stems: vocals, drums, bass and the rest
Demucs

Choose UVR / MDX-Net / RoFormer (разделение звука) if

  • You want the best vocal quality and will tune weights per recording
  • You want a choice: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet, DTTNet
  • Recency matters: the newest weights in the collection are dated 2026-08
UVR / MDX-Net / RoFormer (разделение звука)

Other comparisons

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment