Open-source models for voice and audio analysis

These models identify who spoke when in a recording, clean up noise, separate voice from music and recognize speakers. They are a key part of call analytics and meeting transcription. Check quality on your own recordings, processing speed and the license.

22 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Voice: speakers and sound2022–2026

UVR / MDX-Net / RoFormer (разделение звука)

Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community

A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.

  • Clean vocals from a recording with music
  • Backing tracks and stems for karaoke
  • Removing background music and noise from videos
Sizes
from tens to hundreds of millions of parameters
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and soundGGUF2022–2026

WeSpeaker

WeNet community · China

A set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.

  • Voice verification of a customer during a call
  • Finding repeat calls from the same person
  • Splitting a recording by speaker
Sizes
from a few to tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2022–2026

Demucs

Meta AI, then Alexandre Défossez · France

A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.

  • Separating vocals from music in a recording
  • Backing tracks and karaoke stems
  • Cleaning speech in videos with background music
Sizes
tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2023–2026

RVC (Retrieval-based Voice Conversion)

RVC-Project community · China

The most widely used open voice conversion tool: a model for a specific voice trains on 10–30 minutes of recording and works in real time. Use only with the voice owner's consent.

  • Voicing content with one brand voice
  • Covers and vocal work
  • Real-time voice changing
Sizes
tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundRU2026

FireRedVAD

FireRedTeam (Xiaohongshu) · China

A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.

  • Cutting recordings before speech recognition
  • Separating speech from music and singing in broadcasts and videos
  • Speech detection in voice bots
Sizes
compact, exact size not stated
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundRU2025–2026

Smart Turn

Daily (Pipecat) · USA

Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.

  • Voice bot does not interrupt the customer during pauses
  • Fast reply when the customer has really finished
  • An add-on to a standard speech detector in voice assistants
Sizes
8M (v3) – 580M (v1)
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2025

SAM Audio

Meta · USA

A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.

  • Isolating one person's voice from a noisy recording
  • Removing unwanted sound from a video
  • Splitting a recording into separate sound sources
Sizes
small, base, large (5 to 15 GB of weights)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Voice: speakers and soundRU2020–2025

Silero VAD

Silero · Russia

The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.

  • Cutting calls and recordings before speech recognition
  • Detecting when the customer is speaking in a voice bot
  • Filtering out silence and noise to save on transcription
Sizes
about 2 MB
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2022–2025

NVIDIA Sortformer / TitaNet

NVIDIA · USA

NVIDIA models for "who is speaking": TitaNet recognizes a specific person's voice, Sortformer splits a recording into up to 4 speakers, including live during a call.

  • Real-time speaker tagging in conversations
  • Checking that the same person is calling (voiceprint)
  • Preparing meeting transcripts
Sizes
23M (TitaNet) – 117M (Sortformer)
Hardware
from: Laptop
Commercial use with conditionsDetails
Deepfake detection2025

AntiDeepfake (NII)

National Institute of Informatics, Yamagishi Lab · Japan

Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.

  • Checking audio recordings for synthesis
  • Fine-tuning for your own language and recording channel
  • Comparing several encoders on your own data
Sizes
0,3B – 2B
Hardware
from: Laptop
Non-commercial onlyDetails
Voice: speakers and sound2022–2025

pyannote (диаризация)

pyannoteAI (Hervé Bredin) · France

The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.

  • Tagging calls: which part is the agent, which is the customer
  • Meeting minutes with speaker labels
  • Preparing recordings for transcription and analysis
Sizes
a few million parameters
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2024–2025

ClearerVoice (MossFormer)

Alibaba (Tongyi Lab) · China

Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.

  • Noise suppression in conversation recordings
  • Separating two voices speaking at once
  • Improving old and phone recordings
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2025

TEN VAD

Agora (TEN project) · USA / China

A lightweight speech detector for real-time voice assistants: it notices the start and end of a phrase faster than Silero VAD. Runs on servers, phones and in the browser.

  • Zero-lag speech detection in a voice bot
  • Fast assistant response at the end of a phrase
  • Use in mobile apps and the browser
Sizes
very small, the library is smaller than Silero VAD
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and soundRU2023–2025

Эмоции в русской речи (модели на датасете Dusha)

Community (xbgoose and others), Dusha dataset from SberDevices · Russia

Models that detect emotion from voice in Russian speech: neutral, anger, positive, sadness. Trained on the open Dusha dataset from SberDevices.

  • Finding calls with irritated customers
  • Assessing the tone of operator conversations
  • Prioritizing complaints in a call center
Sizes
21M – 316M
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and sound2024–2025

GPT-SoVITS

RVC-Boss and community · China

Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.

  • Voicing texts with a specific narrator's voice
  • Voice for a bot or assistant
  • Dubbing training videos
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2024–2025

Seed-VC

Songting Liu (Plachtaa) · China

Voice conversion without training: transfers timbre from a 1–30 second sample, can sing and work in real time; V2 also changes accent. Use only with the voice owner's consent.

  • Re-voicing a video with a different voice
  • Voice anonymization in recordings
  • Real-time voice for streams
Sizes
about 70M – 200M
Hardware
from: Laptop
Commercial use with conditionsDetails
Deepfake detection2024

AudioSeal

Meta · USA

An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.

  • Marking speech synthesized by your service
  • Finding your own mark in third-party publications
  • Checking whether synthesis was mixed into a call recording
Sizes
a watermark generator and detector, 16-bit message
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundNot maintained2024

OpenVoice

MyShell and MIT · USA

Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.

  • Voicing videos with the company narrator's voice
  • Voice bot with a recognizable brand voice
  • Transferring timbre onto existing speech synthesis
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundNot maintained2023

Resemble Enhance

Resemble AI · USA

A speech enhancement model: removes noise and restores lost frequencies so a muffled recording sounds studio-quality. Good for preparing a voice for voiceover.

  • Restoring old and phone recordings
  • Cleaning a voice before voiceover and cloning
  • Improving audio in videos and podcasts
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detectionNot maintained2022–2023

SSL Anti-spoofing (wav2vec 2.0 + AASIST)

EURECOM · France

A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.

  • Spotting synthetic speech in calls
  • Checking voice messages and recordings
  • Fine-tuning for your own data and codecs
Sizes
about 0.3B (wav2vec 2.0 XLS-R encoder)
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundNot maintained2022–2023

DeepFilterNet

Hendrik Schröter (University of Erlangen) · Germany

Lightweight real-time speech noise suppression that works even on a regular CPU and low-power devices. Removes hum, street and office noise while keeping the voice.

  • Cleaning calls and voice messages of noise
  • Preparing recordings before speech recognition
  • Noise suppression for video calls
Sizes
about 2M
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detectionNot maintained2021

AASIST

NAVER Clova AI Research and EURECOM · South Korea

The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.

  • Voice check during phone authentication
  • Filtering replays and synthesis in a voice menu
  • A baseline when comparing voice detectors
Sizes
weight files of 0.4 and 1.3 MB
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment