GigaAM or Vosk: which speech recognition to pick
Both models come from Russia, handle Russian, run without a GPU and allow commercial use: GigaAM under MIT, Vosk under Apache 2.0. Sber's GigaAM covers more conversation analytics: emotion recognition, a v3 version with punctuation, and a multilingual version adding Kazakh, Kyrgyz and Uzbek, last updated 2026-07. Vosk is stronger where offline work on weak hardware matters: models from 45 MB run on a Raspberry Pi or a phone, there are streaming models for live audio, and Vosk TTS provides simple Russian speech synthesis. If you use Vosk TTS for voiceover, remember that a voice may only be cloned with the consent of the person it belongs to. The choice comes down to conversation analytics versus on-device autonomy.
Comparison based on catalog data
| Parameter | GigaAM | Vosk (русские модели) |
|---|---|---|
| Category | Speech to text | Speech to text, Text to speech |
| Developer | Sber, Russia | Alpha Cephei, Russia |
| Releases | Apr 2024 – Jul 2026 | Feb 2023 – Sep 2025 |
| Sizes | 220M – 600M | about 45 MB – 1.8 GB |
| Hardware | Laptop | Laptop |
| Commercial use | Commercial use allowed | Commercial use allowed |
| License | MIT | Apache 2.0 |
| Russian | Supported | Supported |
| Ollama | No | No |
| Without GPU | Yes | Yes |
| Tasks |
|
|
Choose GigaAM if
- You need punctuation in transcripts and emotion labels in conversations
- Banking, public sector or call center work with Russian call analysis
- You need neighboring languages: Kazakh, Kyrgyz, Uzbek in the multilingual version
Choose Vosk (русские модели) if
- Recognition has to run offline on a weak device or a phone
- You need low-latency streaming for live audio
- One project needs both recognition and simple Russian speech synthesis
Other comparisons
- Whisper or GigaAM: which to choose for business
- Whisper or Parakeet: which to choose for business
- Silero or Piper: which to choose for business
- Saiga or Vikhr: which Russian fine-tune to choose
- XTTS or F5-TTS: which voice cloning model to pick
- Piper or Kokoro: which text-to-speech model to pick
- F5-TTS or Fish Speech: which voiceover model to pick
- Parakeet or Qwen3-ASR: which transcription model to pick
- pyannote or WeSpeaker: which speaker tool to pick
- Qwen or GigaChat: which to choose for business
- Qwen or Llama: which to choose for business
- GigaChat or YandexGPT: which to choose for business
Need a model for your task?
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.
- SelectThe model and size for your task and hardware budget
- DeployOn your server or in a closed network, with an API
- Fine-tuneOn your data, or connect a knowledge base
- IntegrateInto your CRM, ERP, bot, website or team chat


