F5-TTS or Fish Speech: which voiceover model to pick
Both models clone a voice from a short sample, and both close off commercial use, so a product needs a separate agreement or a different model. A voice may only be cloned with the consent of the person it belongs to. F5-TTS is about 340M, runs on a CPU, officially covers English and Chinese, and gets Russian through community fine-tunes. Fish Speech is larger (0.5B to 4.5B) and needs a GPU, but the catalog lists 80+ languages including Russian plus emotion control. Fish Speech moves faster: the latest S2-Pro shipped in 2026-03 against an F5-TTS update in 2025-03, though the S2-Pro weights use the Fish Audio Research License and commercial use requires a separate agreement.
Comparison based on catalog data
| Parameter | F5-TTS | Fish Speech / OpenAudio |
|---|---|---|
| Category | Text to speech | Text to speech |
| Developer | Shanghai Jiao Tong University and partners, China | Fish Audio, USA / China |
| Releases | Oct 2024 – Mar 2025 | Apr 2024 – Mar 2026 |
| Sizes | about 340M | 0.5B – about 4.5B |
| Hardware | Laptop | Laptop, 1 GPU |
| Commercial use | Non-commercial only | Non-commercial only |
| License | Code MIT, official weights CC BY-NC 4.0, non-commercial | Early versions CC BY-NC-SA 4.0, S2-Pro under the Fish Audio Research License; commercial use only under a separate agreement |
| Russian | Not supported | Supported |
| Ollama | No | No |
| Without GPU | No | No |
| Tasks |
|
|
Choose F5-TTS if
- No GPU available: F5-TTS runs on a regular CPU
- You want a smaller model, about 340M instead of several billion
- English and Chinese are enough, or community fine-tunes work for you
Choose Fish Speech / OpenAudio if
- You need Russian and broad language coverage: the catalog lists 80+
- Emotional delivery matters for your voiceover
- You want the newest release: S2-Pro shipped in 2026-03
Other comparisons
- XTTS or F5-TTS: which voice cloning model to pick
- Silero or Piper: which to choose for business
- Piper or Kokoro: which text-to-speech model to pick
- GigaAM or Vosk: which speech recognition to pick
- Qwen or GigaChat: which to choose for business
- Qwen or Llama: which to choose for business
- GigaChat or YandexGPT: which to choose for business
- DeepSeek or Qwen: which to choose for business
- Gemma or Llama: which to choose for business
- Mistral or Llama: which to choose for business
- Whisper or GigaAM: which to choose for business
- Whisper or Parakeet: which to choose for business
Need a model for your task?
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.
- SelectThe model and size for your task and hardware budget
- DeployOn your server or in a closed network, with an API
- Fine-tuneOn your data, or connect a knowledge base
- IntegrateInto your CRM, ERP, bot, website or team chat


