Fish Speech / OpenAudio
Speech synthesis with voice cloning and emotion control in 80+ languages, including Russian. Quality is close to paid services, but the weights are for research only.
- Developer
- Fish Audio, USA / China
- First release
- Apr 2024
- Latest release
- Mar 2026
- Sizes
- 0.5B – about 4.5B
- License
- Non-commercial onlyEarly versions CC BY-NC-SA 4.0, S2-Pro under the Fish Audio Research License; commercial use only under a separate agreement
- Russian
- Supported
- Running
- On your own serverNeeds a GPU
- Industries
- Media and production, Science and research
What it does
- Voice cloning
- Emotional voiceover
- Multilingual voiceover
Where it is used
Hardware requirements
Versions
- Fish Audio S2-Pro
- OpenAudio S1-mini
- Fish Speech 1.5
- Fish Speech 1.0
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can Fish Speech / OpenAudio be used in a commercial project?
No. License: Early versions CC BY-NC-SA 4.0, S2-Pro under the Fish Audio Research License; commercial use only under a separate agreement. A commercial product needs a different model or a separate agreement with the rights holder.
What hardware does Fish Speech / OpenAudio need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does Fish Speech / OpenAudio support Russian?
Yes, Russian is listed on the model card.
Where can I download Fish Speech / OpenAudio and what does it cost?
The Fish Speech / OpenAudio weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Comparisons
Similar models
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
DetailsText to speechIndexTTSbilibili · ChinaCommercial use with conditionsSpeech synthesis with voice cloning and precise duration control, handy for video dubbing. Controls emotion separately from timbre.
DetailsText to speechQwen3-TTSAlibaba (Qwen) · ChinaCommercial use allowedSpeech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
DetailsSource: huggingface.co/fishaudio/s2-pro. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


