Voice: speakers and sound2022–2026
Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community
A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.
- Clean vocals from a recording with music
- Backing tracks and stems for karaoke
- Removing background music and noise from videos
- Sizes
- from tens to hundreds of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and soundGGUF2022–2026
WeNet community · China
A set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.
- Voice verification of a customer during a call
- Finding repeat calls from the same person
- Splitting a recording by speaker
- Sizes
- from a few to tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2022–2026
Meta AI, then Alexandre Défossez · France
A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.
- Separating vocals from music in a recording
- Backing tracks and karaoke stems
- Cleaning speech in videos with background music
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2023–2026
RVC-Project community · China
The most widely used open voice conversion tool: a model for a specific voice trains on 10–30 minutes of recording and works in real time. Use only with the voice owner's consent.
- Voicing content with one brand voice
- Covers and vocal work
- Real-time voice changing
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and soundRU2026
FireRedTeam (Xiaohongshu) · China
A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.
- Cutting recordings before speech recognition
- Separating speech from music and singing in broadcasts and videos
- Speech detection in voice bots
- Sizes
- compact, exact size not stated
- Hardware
- from: Laptop
Voice: speakers and soundRU2025–2026
Daily (Pipecat) · USA
Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.
- Voice bot does not interrupt the customer during pauses
- Fast reply when the customer has really finished
- An add-on to a standard speech detector in voice assistants
- Sizes
- 8M (v3) – 580M (v1)
- Hardware
- from: Laptop
Voice: speakers and soundRU2020–2025
Silero · Russia
The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.
- Cutting calls and recordings before speech recognition
- Detecting when the customer is speaking in a voice bot
- Filtering out silence and noise to save on transcription
- Sizes
- about 2 MB
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
NVIDIA · USA
NVIDIA models for "who is speaking": TitaNet recognizes a specific person's voice, Sortformer splits a recording into up to 4 speakers, including live during a call.
- Real-time speaker tagging in conversations
- Checking that the same person is calling (voiceprint)
- Preparing meeting transcripts
- Sizes
- 23M (TitaNet) – 117M (Sortformer)
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
pyannoteAI (Hervé Bredin) · France
The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.
- Tagging calls: which part is the agent, which is the customer
- Meeting minutes with speaker labels
- Preparing recordings for transcription and analysis
- Sizes
- a few million parameters
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
Alibaba (Tongyi Lab) · China
Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.
- Noise suppression in conversation recordings
- Separating two voices speaking at once
- Improving old and phone recordings
- Sizes
- under 1B
- Hardware
- from: Laptop
Voice: speakers and sound2025
Agora (TEN project) · USA / China
A lightweight speech detector for real-time voice assistants: it notices the start and end of a phrase faster than Silero VAD. Runs on servers, phones and in the browser.
- Zero-lag speech detection in a voice bot
- Fast assistant response at the end of a phrase
- Use in mobile apps and the browser
- Sizes
- very small, the library is smaller than Silero VAD
- Hardware
- from: Laptop
Voice: speakers and soundRU2023–2025
Community (xbgoose and others), Dusha dataset from SberDevices · Russia
Models that detect emotion from voice in Russian speech: neutral, anger, positive, sadness. Trained on the open Dusha dataset from SberDevices.
- Finding calls with irritated customers
- Assessing the tone of operator conversations
- Prioritizing complaints in a call center
- Sizes
- 21M – 316M
- Hardware
- from: Laptop