Voxtral
Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.
- Developer
- Mistral AI, France
- First release
- Jul 2025
- Latest release
- Mar 2026
- Sizes
- 3B – 24B
- License
- Commercial use with conditionsRecognition models (Mini 3B, Small 24B, Mini 4B Realtime) are Apache 2.0; Voxtral-4B-TTS is CC BY-NC 4.0, non-commercial only
- Russian
- Supported
- Ready-made builds
- GGUF, MLX (Apple)
- Running
- On your own serverAlso runs without a GPU
- Industries
- Customer support, Media and production, Legal
What it does
- Transcribing and summarizing recordings
- Asking questions about audio
- Real-time recognition
- Text-to-speech (non-commercial)
Where it is used
Hardware requirements
Versions
- Voxtral 4B TTS
- Voxtral Mini 4B Realtime
- Voxtral Mini 3B и Small 24B
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems. Quantization compresses a model so it takes less video memory and runs on more modest hardware. Answers change slightly, so quality is checked on your own examples.
Frequently asked questions
Can Voxtral be used in a commercial project?
With conditions. License: Recognition models (Mini 3B, Small 24B, Mini 4B Realtime) are Apache 2.0; Voxtral-4B-TTS is CC BY-NC 4.0, non-commercial only. Restrictions vary — region, company revenue, attribution requirements. Have a lawyer check the terms before a commercial launch.
What hardware does Voxtral need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Some versions also run on an ordinary CPU, without a GPU. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does Voxtral support Russian?
Yes, Russian is listed on the model card.
Where can I download Voxtral and what does it cost?
The Voxtral weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.
DetailsSpeech to textQwen3-ASRAlibaba (Qwen) · ChinaCommercial use allowedSpeech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.
DetailsTextMistralMistral AI · FranceCommercial use with conditionsEuropean models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
DetailsSource: huggingface.co/mistralai/Voxtral-4B-TTS-2603. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


