Video stack: footage, voice, subtitles, translation

A clip is assembled in layers, and models work the same way. Picture, voice, on-screen text and translation are four independent jobs, and asking for all of them in one shot gives you something you cannot revise in parts. With the layers separated you can redo only the voice-over or only the translation without rebuilding the footage. For a business that matters more than the quality of any single frame, because revisions always come.

Updated 22 Sep 2026Find a model in 4 questions

Step 1. Build the footage

A video model creates short scenes from a text description or animates a still image. It usually holds up for a few seconds, so clips are assembled from short pieces rather than requested as a whole minute. This part drives both cost and time: video takes far longer to compute than images and needs a serious graphics card. Without it you are back to stock footage and filming.

What does the work

Video2025–2026

Wan

Alibaba · China

Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).

  • Short promo videos
  • Animating product photos
  • Videos for social media
Sizes
1,3B – 14B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2024–2025

HunyuanVideo

Tencent · China

Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.

  • Video from a text script
  • Animating images
  • Base for fine-tuning your own video models
Sizes
8.3B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2024–2026

LTX-Video / LTX-2

Lightricks · Israel

A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.

  • Ad videos with sound
  • Video from a product photo
  • Voiced scenes for social media
Sizes
2B – 22B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Step 2. Add the voice-over

Speech synthesis reads the script in a chosen voice at the pace and pauses you set. This is where recognisability is decided: the same voice across all your clips works as part of the brand. Cloning someone else without their permission is not on, and that is a legal question rather than a technical one, described here in general terms with a lawyer reviewing your case. Without this part every text edit needs a voice actor.

What does the work

Text to speechGGUF2024–2025

F5-TTS

Shanghai Jiao Tong University and partners · China

A voice cloning model that needs only a few seconds of a sample, in English and Chinese. The community has released many fine-tuned versions for other languages, including Russian.

  • Voice cloning
  • Voicing audiobooks and videos
  • Research and prototypes
Sizes
about 340M
Hardware
from: Laptop
Non-commercial onlyDetails
Text to speechRU2024–2025

CosyVoice / Fun-CosyVoice

Alibaba (Tongyi, FunAudioLLM) · China

Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.

  • Voice for a bot or assistant
  • Cloning a brand voice
  • Voicing videos
Sizes
300M – 0.5B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRUNot maintained2023

Coqui XTTS

Coqui · Germany

A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.

  • Voice cloning from a sample
  • Multilingual voiceover
  • Research and prototypes
Sizes
about 470M
Hardware
from: Laptop
Non-commercial onlyDetails

Step 3. Generate subtitles

Speech recognition turns the finished voice-over back into text with timestamps, which gives subtitles that match the audio frame for frame. It looks redundant since you already have the script, but narration has pauses and feed viewers watch with the sound off. Without subtitles half of a social audience never reaches the point of the clip.

What does the work

Speech to textRUGGUF2022–2025

Whisper

OpenAI · USA

Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.

  • Transcription of calls and video meetings
  • Video subtitles
  • Voice messages to text
Sizes
39M – 1,5B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRUGGUF2026

Qwen3-ASR

Alibaba (Qwen) · China

Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.

  • Transcribing calls and meetings
  • Video subtitles
  • Multilingual recognition
Sizes
0.6B – 1.7B
Hardware
from: Laptop
Commercial use allowedDetails

Step 4. Localise for other markets

A translation model carries the script and subtitles into other languages, and the same clip is then re-voiced from step two. The result is a variant rather than a new shoot, and the cost of entering a neighbouring market drops to the cost of voice-over. Without this part every language becomes its own production cycle with its own contractor and schedule.

What does the work

TranslationRUGGUF2025

Seed-X

ByteDance Seed · China

A compact ByteDance translator for 28 languages, close in quality to large closed systems. Russian is supported. Ready-made compressed versions are available.

  • Translating business correspondence and documents
  • Translating product cards
  • Translating technical and legal texts
Sizes
7B
Hardware
from: Laptop
Commercial use allowedDetails
TranslationRUGGUFNot maintained2022–2023

NLLB-200

Meta · USA

A translator for 200 languages, including rare and minor ones. Russian is supported. Strong language coverage, but the license prohibits commercial use.

  • Translating texts between 200 languages
  • Translating into rare languages where no other models exist
  • Comparing quality when choosing a translator
Sizes
600M to 3.3B (plus 54B MoE)
Hardware
from: Laptop
Non-commercial onlyDetails
TranslationRUGGUF2024–2025

Tower

Unbabel · Portugal

Language models tailored for translation and multilingual text work: translating, editing, and assessing translation quality. Russian is supported. Non-commercial license.

  • Translation that respects context and terminology
  • Post-editing machine translation
  • Assessing the quality of a finished translation
Sizes
2B – 72B
Hardware
from: Laptop
Non-commercial onlyDetails

What to check before you start

Expectations are usually too high on two points: scene length and control. A model does not hand you the exact shot on the first try, so budget for iterations rather than a single request. Agree in advance where your line is: a real person in frame, someone else brand, an employee voice. Rights to generated material, labelling of synthetic content and use of someone likeness are covered here in general terms only, with a lawyer reviewing the specific case.

Other stacks

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment