Speech to textRUGGUF2025–2026
Microsoft · USA
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
- Transcribing long meetings with speaker labels
- Voicing podcasts and dialogues
- Real-time voice for assistants
- Sizes
- 0.5B – 9B
- Hardware
- from: Laptop
Text to speechGGUF2025–2026
bilibili · China
Speech synthesis with voice cloning and precise duration control, handy for video dubbing. Controls emotion separately from timbre.
- Video dubbing matched to timing
- Voice cloning
- Emotional voiceover
- Sizes
- about 1B – 2B
- Hardware
- from: Laptop
TextGGUF2025–2026
Meituan · China
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
- Agents for orders and service processes
- Corporate assistant
- Analysis of long documents
- Sizes
- 1B – 1.6T-A48B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Boson AI · USA
Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.
- Expressive video voiceover
- Voicing dialogues
- Voice cloning
- Sizes
- about 3B – 8B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
Zyphra · USA
Speech synthesis with voice cloning and fine control over emotion, speed and pitch.
- Voice cloning
- Emotional voiceover
- Voicing videos
- Sizes
- about 1.6B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Resemble AI · USA
Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- about 350M – 500M
- Hardware
- from: Laptop
Text to speechRU2025–2026
OpenMOSS (Fudan University) · China
A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.
- Voicing podcasts and dialogues
- Voice for an assistant
- Voice cloning
- Sizes
- 100M – 8.5B
- Hardware
- from: Laptop
Text to speechRU2025–2026
Supertone · South Korea
Very fast, lightweight speech synthesis that runs directly on the device, without a GPU or the cloud. Supertonic 3 speaks 31 languages, including Russian.
- Voicing voice bot replies on an ordinary server
- Voiceover in offline and mobile apps
- Reading texts and notifications aloud
- Sizes
- about 99M
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.
- Voice cloning
- Voicing videos and audiobooks
- Voice for an assistant
- Sizes
- 0.5B – 2.3B
- Hardware
- from: Laptop
Text to speech2025–2026
Kyutai · France
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
- Streaming speech transcription for voice bots
- Voicing replies with minimal delay
- Speech synthesis on a server without a GPU (Pocket TTS)
- Sizes
- 100M (Pocket TTS) – 2.6B
- Hardware
- from: Laptop
Speech to textRUGGUF2025–2026
Mistral AI · France
Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.
- Transcribing and summarizing recordings
- Asking questions about audio
- Real-time recognition
- Sizes
- 3B – 24B
- Hardware
- from: Laptop
Text to speechRU2024–2026
Fish Audio · USA / China
Speech synthesis with voice cloning and emotion control in 80+ languages, including Russian. Quality is close to paid services, but the weights are for research only.
- Voice cloning
- Emotional voiceover
- Multilingual voiceover
- Sizes
- 0.5B – about 4.5B
- Hardware
- from: Laptop
Text to speechRU2026
k2-fsa (Next-gen Kaldi) · China
Speech synthesis with voice cloning from a short sample in 646 languages, including Russian and languages of Russia's peoples. A voice can be described in words. Weights are for non-commercial use only.
- Voiceover in rare languages
- Voice cloning from a sample
- Research and prototypes of multilingual voiceover
- Sizes
- 0.6B
- Hardware
- from: Laptop
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
Text to speechRU2022–2025
Silero · Russia
Lightweight Russian speech synthesis that runs on a regular CPU. Version v5 added CIS languages and languages of Russia's peoples: Tatar, Bashkir, Yakut, Kazakh and others.
- Voicing voice bot replies
- Reading texts in Russian
- Voices in the languages of Russia's peoples
- Sizes
- tens of megabytes
- Hardware
- from: Laptop
Text to speech2025
Nari Labs · South Korea
A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.
- Voicing dialogues and podcasts
- Ads with natural speech
- Training role-plays
- Sizes
- 1B – 2B
- Hardware
- from: Laptop
Speech to textRU2023–2025
Alpha Cephei · Russia
Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.
- Transcribing Russian calls and recordings without the cloud
- Voice control in apps and kiosks
- Low-latency streaming speech recognition
- Sizes
- about 45 MB – 1.8 GB
- Hardware
- from: Laptop
Text to speechRU2025
ESpeech (independent group of Russian-speaking developers) · Russia
Russian speech synthesis with voice cloning based on the F5-TTS architecture, trained on Russian speech datasets collected by the authors. Stress is placed automatically. Several variants, including a "podcaster" one.
- Voicing videos and audiobooks in Russian
- Cloning a narrator's voice from a sample
- Voice for a bot or assistant in Russian
- Sizes
- about 340M
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
RVC-Boss and community · China
Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.
- Voicing texts with a specific narrator's voice
- Voice for a bot or assistant
- Dubbing training videos
- Sizes
- under 1B
- Hardware
- from: Laptop
Text to speechRUGGUF2023–2025
Rhasspy / Open Home Foundation · USA
Very fast speech synthesis that runs even on a Raspberry Pi. Ready-made voices in 35+ languages, including several Russian ones.
- Voicing notifications and bot replies
- Voice for offline devices
- Voice menus
- Sizes
- about 5M – 30M
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
Shanghai Jiao Tong University and partners · China
A voice cloning model that needs only a few seconds of a sample, in English and Chinese. The community has released many fine-tuned versions for other languages, including Russian.
- Voice cloning
- Voicing audiobooks and videos
- Research and prototypes
- Sizes
- about 340M
- Hardware
- from: Laptop
Text to speechGGUF2025
Canopy Labs · USA
Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.
- Real-time voice for an assistant
- Emotional voiceover
- Voice cloning
- Sizes
- 3B
- Hardware
- from: Laptop
Text to speechGGUF2025
Sesame · USA
A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.
- Voice for a conversational assistant
- Voicing dialogues
- Voice product prototypes
- Sizes
- 1B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
hexgrad (independent developer) · not disclosed
A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.
- Voicing articles and notifications
- Voice for apps without a GPU
- Bulk text voiceover
- Sizes
- 82M
- Hardware
- from: Laptop
Text to speechGGUF2024
Hugging Face · USA
Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.
- Choosing a voice by description
- Voicing videos
- Voice service prototypes
- Sizes
- 880M – 2.2B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
Text to speechNot maintained2024
MyShell and MIT · USA
Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.
- Voicing bot replies in foreign languages
- Voicing training materials
- Reading texts aloud on a server without a GPU
- Sizes
- small, runs in real time on a CPU
- Hardware
- from: Laptop
Text to speechNot maintained2023
Columbia University · USA
A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.
- Voicing texts in English
- A base for fine-tuning your own voice
- Voice service prototypes
- Sizes
- about 150M
- Hardware
- from: Laptop
Text to speechRUNot maintained2023
Coqui · Germany
A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.
- Voice cloning from a sample
- Multilingual voiceover
- Research and prototypes
- Sizes
- about 470M
- Hardware
- from: Laptop
Text to speechRUGGUFNot maintained2023
Suno · USA
One of the first open models to voice text with intonation, laughter and pauses. Supports about ten languages, including Russian. Now outdated.
- Draft voiceovers for videos
- Voice service prototypes
- Sound effects in speech
- Sizes
- about 300M – 1B
- Hardware
- from: Laptop