Open-source speech recognition models

Speech recognition models turn calls, meetings and voice messages into searchable, analyzable text. Local deployment matters when call recordings must stay in-house. Check accuracy in your languages and with background noise, speaker separation, the license and speed on your hardware.

20 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Speech to textRUGGUF2025–2026

VibeVoice

Microsoft · USA

Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.

  • Transcribing long meetings with speaker labels
  • Voicing podcasts and dialogues
  • Real-time voice for assistants
Sizes
0.5B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

MiMo

Xiaomi · China

Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.

  • Logic and calculation tasks
  • Agents with tools
  • Help for developers
Sizes
7B – 1,02T-A42B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textGGUF2024–2026

Moonshine

Moonshine AI (Useful Sensors) · USA

Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.

  • Voice control of devices
  • Offline recognition on a phone
  • Live subtitles
Sizes
27M – 245M
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textGGUF2025–2026

Granite Speech

IBM · USA

IBM speech models for recognizing and translating speech in English, several European languages and Japanese. Designed for enterprise use.

  • Transcribing business meetings
  • Translating speech into text in another language
  • Voice assistants
Sizes
470M – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRUGGUF2024–2026

GigaAM

Sber · Russia

Sber's models for Russian speech recognition, among the most accurate for Russian. Includes emotion recognition, v3 with punctuation, and a multilingual version (Russian, Kazakh, Kyrgyz, Uzbek).

  • Transcribing calls in Russian
  • Meeting minutes
  • Voice control of services
Sizes
220M – 600M
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRUGGUF2026

Qwen3-ASR

Alibaba (Qwen) · China

Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.

  • Transcribing calls and meetings
  • Video subtitles
  • Multilingual recognition
Sizes
0.6B – 1.7B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textGGUF2026

Cohere Transcribe

Cohere · Canada

Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.

  • Transcribing meetings and interviews
  • Subtitles
  • Searching an audio archive
Sizes
2B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRUGGUF2025–2026

Higgs Audio

Boson AI · USA

Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.

  • Expressive video voiceover
  • Voicing dialogues
  • Voice cloning
Sizes
about 3B – 8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Speech to textRU2023–2026

NVIDIA Parakeet / Canary / Nemotron Speech

NVIDIA · USA

Fast NVIDIA speech recognition models, including streaming ones for real-time use. Parakeet TDT v3 and Nemotron 3.5 ASR understand Russian.

  • Transcribing calls and meetings
  • Video subtitles
  • Real-time voice input
Sizes
110M – 2.5B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speech2025–2026

Kyutai STT, TTS и Pocket TTS

Kyutai · France

Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.

  • Streaming speech transcription for voice bots
  • Voicing replies with minimal delay
  • Speech synthesis on a server without a GPU (Pocket TTS)
Sizes
100M (Pocket TTS) – 2.6B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRUGGUF2025–2026

Voxtral

Mistral AI · France

Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.

  • Transcribing and summarizing recordings
  • Asking questions about audio
  • Real-time recognition
Sizes
3B – 24B
Hardware
from: Laptop
Commercial use with conditionsDetails
Speech to textRU2025

Omnilingual ASR

Meta · USA

Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.

  • Transcription in rare and local languages
  • Digitizing oral archives
  • Subtitles in many languages
Sizes
300M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to text2024–2025

SenseVoice / Paraformer / Fun-ASR

Alibaba (Tongyi, FunAudioLLM) · China

Alibaba's set of fast speech recognition models, primarily for Chinese and Asian languages. SenseVoice also detects emotions and sound events.

  • Transcribing calls
  • Detecting emotions in the voice
  • Recognizing laughter, music and other sounds
Sizes
about 230M – 800M
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUGGUF2024–2025

Vikhr

Vikhr Models · Russia

Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.

  • Russian-language assistant on your own PC or server
  • Knowledge-base answers (RAG)
  • Texts and emails in Russian
Sizes
0.5B – 24B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRU2023–2025

Vosk (русские модели)

Alpha Cephei · Russia

Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.

  • Transcribing Russian calls and recordings without the cloud
  • Voice control in apps and kiosks
  • Low-latency streaming speech recognition
Sizes
about 45 MB – 1.8 GB
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2022–2025

pyannote (диаризация)

pyannoteAI (Hervé Bredin) · France

The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.

  • Tagging calls: which part is the agent, which is the customer
  • Meeting minutes with speaker labels
  • Preparing recordings for transcription and analysis
Sizes
a few million parameters
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRU2025

T-one

T-Bank · Russia

A compact T-Bank streaming model for recognizing Russian speech in phone calls. Works in real time even without a GPU.

  • Transcribing phone calls
  • Voice robots on the line
  • Call quality control
Sizes
72M
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2025

Kimi-Audio

Moonshot AI · China

A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.

  • Speech recognition
  • Detecting emotions and sound events
  • Speech-to-speech voice dialogue
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Speech to textRUGGUF2022–2025

Whisper

OpenAI · USA

Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.

  • Transcription of calls and video meetings
  • Video subtitles
  • Voice messages to text
Sizes
39M – 1,5B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistantsRUNot maintained2023

SeamlessM4T / Seamless

Meta · USA

Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.

  • Speech-to-speech translation
  • Translating and transcribing recordings
  • Streaming translation
Sizes
281M – 2.3B
Hardware
from: Laptop
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment