Speech recognition models turn calls, meetings and voice messages into searchable, analyzable text. Local deployment matters when call recordings must stay in-house. Check accuracy in your languages and with background noise, speaker separation, the license and speed on your hardware.
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.
Sber's models for Russian speech recognition, among the most accurate for Russian. Includes emotion recognition, v3 with punctuation, and a multilingual version (Russian, Kazakh, Kyrgyz, Uzbek).
Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.
Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
Streaming speech transcription for voice bots
Voicing replies with minimal delay
Speech synthesis on a server without a GPU (Pocket TTS)
Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.
Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
Russian-language assistant on your own PC or server
Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.
Transcribing Russian calls and recordings without the cloud
The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.
Tagging calls: which part is the agent, which is the customer
Meeting minutes with speaker labels
Preparing recordings for transcription and analysis
A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.
Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.
Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.