AI models for media and production

For media and production, open models generate video, music and sound, produce subtitles, voiceovers and translations, and process footage. Your own infrastructure removes cloud service limits on volume and turnaround. Check the commercial license terms and GPU requirements carefully.

181 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Image generationGGUF2025–2026

Qwen-Image

Alibaba · China

Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.

  • Infographics for product cards
  • Photo editing by text command
  • Ad creatives
Sizes
7B – 20B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Speech to textRUGGUF2025–2026

VibeVoice

Microsoft · USA

Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.

  • Transcribing long meetings with speaker labels
  • Voicing podcasts and dialogues
  • Real-time voice for assistants
Sizes
0.5B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundGGUF2025–2026

YuE

M-A-P and HKUST · China

Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.

  • Songs and jingles from lyrics
  • Demo versions of tracks
  • Music for videos
Sizes
0.5B – 7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Music and sound2026

HeartMuLa

HeartMuLa Team · not disclosed

An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.

  • Songs and jingles from lyrics
  • Music for videos
  • Transcribing song lyrics
Sizes
3B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextRU2022–2026

YandexGPT / AliceAI

Yandex · Russia

Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.

  • Russian-language assistant and chatbot
  • Answers based on the company knowledge base
  • Base for industry-specific fine-tuning
Sizes
8B – 100B
Hardware
from: Laptop
Commercial use with conditionsDetails
3D2024–2026

DUSt3R / MASt3R

NAVER LABS Europe · France (NAVER, South Korea)

The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.

  • 3D scene from several photos without calibration
  • Point matching between images
  • Mapping from video (SLAM)
Sizes
0.57B – 0.69B
Hardware
from: Laptop
Non-commercial onlyDetails
Music and soundGGUF2022–2026

MERT

m-a-p (Multimodal Art Projection) · UK / China

A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.

  • Automatic tagging of a music catalog
  • Finding similar tracks
  • Detecting genre, mood and tempo
Sizes
95M – 632M
Hardware
from: Laptop
Non-commercial onlyDetails
Voice assistants2026

Samsone

Samsung · South Korea

Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.

  • Describing an audio recording in words
  • Answering questions about a sound
  • Identifying the type of sound and the setting
Sizes
99M – 356M
Hardware
from: Laptop
Non-commercial onlyDetails
Deepfake detection2023–2026

TrustMark

Adobe Research and University of Surrey · USA

An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.

  • Marking images on the way out of your own pipeline
  • Checking the provenance of a submitted image
  • Linking with content provenance metadata
Sizes
model types Q and P with different mark capacity
Hardware
from: Laptop
Commercial use allowedDetails
VideoGGUF2025–2026

MAGI

Sand AI · China

Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.

  • Long videos with continuation
  • Video with sound
  • Animating images
Sizes
4.5B – 114B-A6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer vision2024–2026

MoGe

Microsoft Research · USA

Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.

  • Measuring rooms and objects from photos
  • 3D point cloud from a single shot
  • Preparing data for robots and AR
Sizes
ViT-S – ViT-G
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textGGUF2024–2026

Moonshine

Moonshine AI (Useful Sensors) · USA

Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.

  • Voice control of devices
  • Offline recognition on a phone
  • Live subtitles
Sizes
27M – 245M
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechGGUF2025–2026

IndexTTS

bilibili · China

Speech synthesis with voice cloning and precise duration control, handy for video dubbing. Controls emotion separately from timbre.

  • Video dubbing matched to timing
  • Voice cloning
  • Emotional voiceover
Sizes
about 1B – 2B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text2024–2026

Hunyuan / Hy

Tencent · China

Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.

  • Corporate assistant
  • Translation and multilingual texts
  • Agents with tools
Sizes
0.5B – 770B-A49B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and sound2022–2026

UVR / MDX-Net / RoFormer (разделение звука)

Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community

A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.

  • Clean vocals from a recording with music
  • Backing tracks and stems for karaoke
  • Removing background music and noise from videos
Sizes
from tens to hundreds of millions of parameters
Hardware
from: Laptop
Commercial use with conditionsDetails
Deepfake detection2025–2026

Community Forensics

University of Michigan · USA

A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.

  • Checking submitted photos and illustrations
  • Filtering AI images in a content flow
  • Flagging suspicious images for manual review
Sizes
22M
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2023–2026

UniversalFakeDetect

University of Wisconsin-Madison · USA

An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.

  • Checking images from new, unfamiliar generators
  • A baseline when comparing detectors
  • Fast rollout of a check without training a large model
Sizes
a linear classifier on top of CLIP ViT-L/14
Hardware
from: Laptop
Commercial use allowedDetails
Video2025–2026

Wan

Alibaba · China

Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).

  • Short promo videos
  • Animating product photos
  • Videos for social media
Sizes
1,3B – 14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationRU2022–2026

Kandinsky

Sber (Kandinsky Lab) · Russia

Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.

  • Images from Russian-language descriptions
  • Short promo videos from text or a photo
  • Instruction-based image editing
Sizes
2B – 19B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025–2026

Chroma

lodestones (independent developer) · not disclosed

A community model retrained from FLUX.1-schnell with a simplified architecture. No style censorship; popular as a base for fine-tuning.

  • Base for fine-tuning your own styles
  • Artistic illustrations
  • Images for games
Sizes
4B – 8.9B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2024–2026

LTX-Video / LTX-2

Lightricks · Israel

A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.

  • Ad videos with sound
  • Video from a product photo
  • Voiced scenes for social media
Sizes
2B – 22B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2026

MiniMax H3 (Hailuo)

MiniMax · China

Open weights of MiniMax's Hailuo video model. A large 33B model that makes video from text and images, but needs several server GPUs.

  • Cinematic ad videos
  • Video from text and images
  • Complex scenes with motion
Sizes
33B + 32B encoder
Hardware
from: Cluster
Commercial use with conditionsDetails
Robotics2025–2026

NVIDIA Cosmos

NVIDIA · USA

"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.

  • Synthetic video for training robots and self-driving vehicles
  • Testing scenarios in simulation
  • Robot control (Policy versions)
Sizes
2B – 65B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice: speakers and sound2022–2026

Demucs

Meta AI, then Alexandre Défossez · France

A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.

  • Separating vocals from music in a recording
  • Backing tracks and karaoke stems
  • Cleaning speech in videos with background music
Sizes
tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2023–2026

RVC (Retrieval-based Voice Conversion)

RVC-Project community · China

The most widely used open voice conversion tool: a model for a specific voice trains on 10–30 minutes of recording and works in real time. Use only with the voice owner's consent.

  • Voicing content with one brand voice
  • Covers and vocal work
  • Real-time voice changing
Sizes
tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2026

MOSS-VL

OpenMOSS (Fudan University) · China

An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.

  • Analyzing long videos and finding events by time
  • Real-time streaming video analysis
  • Understanding photos and documents
Sizes
about 11B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2026

Krea 2

Krea · USA

A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.

  • Realistic photos for advertising
  • High-resolution images
  • Style fine-tuning
Sizes
12B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2026

Ideogram 4

Ideogram · Canada

Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.

  • Testing text-in-image generation
  • Research and prototypes
Sizes
about 9B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Speech to textRUGGUF2026

Qwen3-ASR

Alibaba (Qwen) · China

Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.

  • Transcribing calls and meetings
  • Video subtitles
  • Multilingual recognition
Sizes
0.6B – 1.7B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textGGUF2026

Cohere Transcribe

Cohere · Canada

Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.

  • Transcribing meetings and interviews
  • Subtitles
  • Searching an audio archive
Sizes
2B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRUGGUF2025–2026

Higgs Audio

Boson AI · USA

Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.

  • Expressive video voiceover
  • Voicing dialogues
  • Voice cloning
Sizes
about 3B – 8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Text to speechRUGGUF2025–2026

Zonos

Zyphra · USA

Speech synthesis with voice cloning and fine control over emotion, speed and pitch.

  • Voice cloning
  • Emotional voiceover
  • Voicing videos
Sizes
about 1.6B
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundRUGGUF2025–2026

ACE-Step

ACE Studio and StepFun · China

Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.

  • Songs and jingles for ads
  • Background music for videos
  • Demo versions of tracks
Sizes
about 2B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUGGUF2025–2026

MiniMax

MiniMax · China

Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.

  • Analysis of large document archives in a single request
  • Agents with tools
  • Help for developers
Sizes
230B-A10B – 456B-A46B
Hardware
from: Cluster
Commercial use with conditionsDetails
Image + text2023–2026

InternVideo

Shanghai AI Lab (OpenGVLab) · China

A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.

  • Searching a video archive with a text query
  • Action recognition in video
  • Answering questions about a long recording
Sizes
small encoders – 9B
Hardware
from: Laptop
Commercial use allowedDetails
VideoGGUF2025–2026

SCAIL

Zhipu AI (Z.ai) and Tsinghua University · China

Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.

  • Transferring an actor's motion to a character
  • Replacing a character in a finished video
  • Animating mascots and illustrations
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025–2026

HiDream

HiDream.ai · China

Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.

  • Image generation from descriptions
  • Editing images with words
  • Variations of product photos
Sizes
about 9B – 17B
Hardware
from: 1 GPU
Commercial use allowedDetails
AvatarsGGUF2025–2026

LongCat-Video-Avatar

Meituan · China

Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.

  • News or course presenter videos
  • Promo videos with a talking character
  • Singing and voice-over
Sizes
based on LongCat-Video 13.6B
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2024–2026

TripoSR / TripoSG

VAST (TripoSR together with Stability AI) · China

VAST family: a 3D model from a single photo. TripoSR runs in under a second, TripoSG gives cleaner geometry, TripoSplat builds a scene from Gaussian points.

  • 3D product model from a photo
  • Object assets for games and AR
  • Quick 3D prototype for printing
Sizes
up to 1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textRU2023–2026

NVIDIA Parakeet / Canary / Nemotron Speech

NVIDIA · USA

Fast NVIDIA speech recognition models, including streaming ones for real-time use. Parakeet TDT v3 and Nemotron 3.5 ASR understand Russian.

  • Transcribing calls and meetings
  • Video subtitles
  • Real-time voice input
Sizes
110M – 2.5B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speechRUGGUF2025–2026

Chatterbox

Resemble AI · USA

Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.

  • Voice for a bot or assistant
  • Cloning a brand voice
  • Voicing videos
Sizes
about 350M – 500M
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRU2025–2026

MOSS-TTS / MOSS-TTSD

OpenMOSS (Fudan University) · China

A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.

  • Voicing podcasts and dialogues
  • Voice for an assistant
  • Voice cloning
Sizes
100M – 8.5B
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundGGUF2024–2026

Stable Audio

Stability AI · UK

Generates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.

  • Sound effects for videos and games
  • Background music and jingles
  • Interface sounds
Sizes
about 0.5B – 2.3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Photo editing2023–2026

BRIA RMBG

BRIA AI · Israel

BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.

  • Cutting products out onto a white background
  • Staff and expert photos without background
  • Background removal in video
Sizes
44M – 220M
Hardware
from: Laptop
Non-commercial onlyDetails
Image + textGGUF2025–2026

Keye-VL

Kuaishou · China

Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.

  • Analysing and describing short videos
  • Reviewing clips and content
  • Finding the right moment in a video
Sizes
8B – 671B-A37B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRU2025–2026

Supertonic

Supertone · South Korea

Very fast, lightweight speech synthesis that runs directly on the device, without a GPU or the cloud. Supertonic 3 speaks 31 languages, including Russian.

  • Voicing voice bot replies on an ordinary server
  • Voiceover in offline and mobile apps
  • Reading texts and notifications aloud
Sizes
about 99M
Hardware
from: Laptop
Commercial use with conditionsDetails
3D2025–2026

VGGT

Meta and the University of Oxford (VGG) · USA / UK

Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.

  • 3D model of a room or object from a photo series
  • Camera pose estimation for photogrammetry
  • Point cloud for measurements and comparison with the plan
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2024–2026

Sapiens

Meta · USA

Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.

  • Pose and body keypoint detection
  • Segmentation of body parts and clothing
  • Separating a person from the background
Sizes
0.1B – 5B
Hardware
from: Laptop
Commercial use with conditionsDetails
Avatars2024–2026

Hallo

Fudan University · China

A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.

  • Presenter video from a photo and audio
  • Long training videos
  • Live avatar
Sizes
about 1B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2025–2026

HY-World (HunyuanWorld)

Tencent · China

Generates whole 3D worlds and scenes from text or an image that you can walk through. The second version builds a scene from video and photos.

  • 3D scenes for games and virtual tours
  • Backgrounds and environments for video production
  • Draft locations for simulations
Sizes
set of several models
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Text to speechRUGGUF2025–2026

VoxCPM

OpenBMB (ModelBest, Tsinghua University) · China

Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.

  • Voice cloning
  • Voicing videos and audiobooks
  • Voice for an assistant
Sizes
0.5B – 2.3B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speech2025–2026

Kyutai STT, TTS и Pocket TTS

Kyutai · France

Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.

  • Streaming speech transcription for voice bots
  • Voicing replies with minimal delay
  • Speech synthesis on a server without a GPU (Pocket TTS)
Sizes
100M (Pocket TTS) – 2.6B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2024–2026

Audio Flamingo

NVIDIA · USA

Models that listen to speech, sounds and music and answer questions about them. Audio Flamingo Next handles recordings up to 30 minutes. Research use only.

  • Detailed descriptions of audio recordings
  • Questions and answers about a long recording
  • Tagging music and sounds
Sizes
0.5B – 8B
Hardware
from: Laptop
Non-commercial onlyDetails
Computer vision2023–2026

Segment Anything (SAM)

Meta · USA

Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.

  • Background removal from product photos
  • Counting objects in photos
  • Data labeling for training
Sizes
91M – ~0,85B
Hardware
from: Laptop
Commercial use with conditionsDetails
Speech to textRUGGUF2025–2026

Voxtral

Mistral AI · France

Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.

  • Transcribing and summarizing recordings
  • Asking questions about audio
  • Real-time recognition
Sizes
3B – 24B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speechRU2024–2026

Fish Speech / OpenAudio

Fish Audio · USA / China

Speech synthesis with voice cloning and emotion control in 80+ languages, including Russian. Quality is close to paid services, but the weights are for research only.

  • Voice cloning
  • Emotional voiceover
  • Multilingual voiceover
Sizes
0.5B – about 4.5B
Hardware
from: Laptop
Non-commercial onlyDetails
Text2025–2026

Reka Flash и Reka Edge

Reka AI · USA

Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.

  • Photo and video analysis (Edge)
  • Object detection in images
  • Reasoning tasks (Flash)
Sizes
7B – 21B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and soundRU2026

FireRedVAD

FireRedTeam (Xiaohongshu) · China

A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.

  • Cutting recordings before speech recognition
  • Separating speech from music and singing in broadcasts and videos
  • Speech detection in voice bots
Sizes
compact, exact size not stated
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRU2026

OmniVoice

k2-fsa (Next-gen Kaldi) · China

Speech synthesis with voice cloning from a short sample in 646 languages, including Russian and languages of Russia's peoples. A voice can be described in words. Weights are for non-commercial use only.

  • Voiceover in rare languages
  • Voice cloning from a sample
  • Research and prototypes of multilingual voiceover
Sizes
0.6B
Hardware
from: Laptop
Non-commercial onlyDetails
Video2025–2026

Matrix-Game

Skywork AI (Kunlun) · China

An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.

  • Game world prototypes without an engine
  • Interactive demos and simulations
  • Generating data to train agents
Sizes
1.8B – 17B
Hardware
from: 1 GPU
Commercial use allowedDetails
Photo editing2025–2026

MatAnyone

S-Lab, Nanyang Technological University · Singapore

Cuts a person out of video with a precise alpha mask, including hair and edges, without a green screen. Needs a first-frame mask, for example from SAM.

  • Background replacement in video without chroma key
  • Cutting out a person for editing and effects
  • Preparing videos for advertising and social media
Sizes
about 35M
Hardware
from: Laptop
Non-commercial onlyDetails
Music and sound2025–2026

ThinkSound / PrismAudio

Alibaba Tongyi (FunAudioLLM) · China

Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.

  • Audio for video based on the scene
  • Editing individual sounds in a track
  • Sound effects from a description
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2026

daVinci-MagiHuman

SII-GAIR and Sand.ai · China

Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.

  • Presenter video from a script
  • Ad videos with a talking character
  • Training videos with a narrator
Sizes
15B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2024–2026

FLUX

Black Forest Labs · Germany

Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.

  • Images for product cards
  • Banners and covers
  • Photo editing by description (Kontext)
Sizes
4B – 32B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2024–2026

Hunyuan Image

Tencent · China

Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.

  • Complex scenes from long descriptions
  • Images with Chinese and English text
  • Instruction-based image editing
Sizes
1.5B – 80B-A13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2025–2026

Z-Image

Alibaba (Tongyi-MAI) · China

A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.

  • Photorealistic ad images
  • Images with English and Chinese text
  • Bulk visual generation
Sizes
6B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2025–2026

SkyReels

Skywork AI (Kunlun Tech) · China

Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.

  • Long videos with continuation
  • Video with one character from a reference
  • Talking avatar from a voice
Sizes
1.3B – 19B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Text to speechRUGGUF2026

Qwen3-TTS

Alibaba (Qwen) · China

Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.

  • Voice for a bot or assistant
  • Cloning a brand voice
  • Choosing a voice by description
Sizes
0.6B – 1.7B
Hardware
from: Laptop
Commercial use allowedDetails
VideoGGUF2026

MOVA

OpenMOSS / MOSI · China

Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.

  • Short clips with speech and sound from a description
  • Ad scenes with dialogue
  • Video prototypes for storyboards
Sizes
32B-A18B
Hardware
from: 1 GPU
Commercial use allowedDetails
3DGGUF2024–2025

TRELLIS

Microsoft · USA

One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.

  • 3D models of products and interiors from photos
  • Assets for games and AR/VR
  • Prototypes for 3D printing
Sizes
up to 4B (TRELLIS.2)
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer visionGGUF2024–2025

Depth Anything

ByteDance and the University of Hong Kong · China

Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.

  • Estimating distances and volumes from a camera
  • Depth effects for photo and video
  • Navigation for robots and drones
Sizes
25M – 1.4B
Hardware
from: Laptop
Commercial use with conditionsDetails
Speech to textRU2025

Omnilingual ASR

Meta · USA

Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.

  • Transcription in rare and local languages
  • Digitizing oral archives
  • Subtitles in many languages
Sizes
300M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to text2024–2025

SenseVoice / Paraformer / Fun-ASR

Alibaba (Tongyi, FunAudioLLM) · China

Alibaba's set of fast speech recognition models, primarily for Chinese and Asian languages. SenseVoice also detects emotions and sound events.

  • Transcribing calls
  • Detecting emotions in the voice
  • Recognizing laughter, music and other sounds
Sizes
about 230M – 800M
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speechRU2024–2025

CosyVoice / Fun-CosyVoice

Alibaba (Tongyi, FunAudioLLM) · China

Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.

  • Voice for a bot or assistant
  • Cloning a brand voice
  • Voicing videos
Sizes
300M – 0.5B
Hardware
from: Laptop
Commercial use allowedDetails
Text2025

HyperCLOVA X SEED

Naver · South Korea

Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.

  • Lightweight Korean-English assistant
  • Analysis of images and documents
  • Text classification
Sizes
0.5B – 32B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and sound2025

SAM Audio

Meta · USA

A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.

  • Isolating one person's voice from a noisy recording
  • Removing unwanted sound from a video
  • Splitting a recording into separate sound sources
Sizes
small, base, large (5 to 15 GB of weights)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3D2025

MapAnything

Meta and Carnegie Mellon University · USA

A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.

  • 3D reconstruction of an object or room from photos
  • Exporting the scene to COLMAP format for further processing
  • Depth and camera pose estimation
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2025

Pi3 (π³)

Shanghai AI Lab · China

Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.

  • 3D scene reconstruction from video
  • Camera pose estimation from frames
  • Point clouds for research and prototypes
Sizes
0.96B – 1.4B
Hardware
from: 1 GPU
Non-commercial onlyDetails
3D2025

HY-Motion

Tencent Hunyuan · China

Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.

  • Character animation from a text description
  • Draft animation for games and videos
  • Motion library for avatars
Sizes
0.46B – 1B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2025

Perception Encoder (PE)

Meta · USA

Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.

  • Search photos and videos by description
  • Catalog labeling and tagging
  • Search across audio and video (PE-AV)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2024–2025

VideoSeal

Meta · USA

A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.

  • Marking video created or processed by AI
  • Finding your own mark in re-uploaded clips
  • Protecting ad materials from being reused as someone else's
Sizes
a mark of 96 to 1024 bits
Hardware
from: Laptop
Commercial use allowedDetails
VideoGGUF2024–2025

HunyuanVideo

Tencent · China

Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.

  • Video from a text script
  • Animating images
  • Base for fine-tuning your own video models
Sizes
8.3B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3DGGUF2025

SAM 3D

Meta · USA

Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.

  • 3D model of an item from a catalog photo
  • Estimating body pose and shape from a photo
  • Try-on and AR scenarios
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Text to speech2025

Dia

Nari Labs · South Korea

A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.

  • Voicing dialogues and podcasts
  • Ads with natural speech
  • Training role-plays
Sizes
1B – 2B
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundRU2020–2025

Silero VAD

Silero · Russia

The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.

  • Cutting calls and recordings before speech recognition
  • Detecting when the customer is speaking in a voice bot
  • Filtering out silence and noise to save on transcription
Sizes
about 2 MB
Hardware
from: Laptop
Commercial use allowedDetails
Video2025

Ovi

Character.AI · USA

Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.

  • Short clips with talking characters
  • Animating an image with voice-over
  • Ad scene prototypes
Sizes
11B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer vision2023–2025

MetaCLIP / MetaCLIP 2

Meta · USA

Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.

  • Image search by text
  • Image classification without training
  • Search research and prototypes
Sizes
0.15B – 3.6B
Hardware
from: Laptop
Non-commercial onlyDetails
VideoGGUF2025

LongCat-Video

Meituan · China

A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.

  • Long videos
  • Video from a photo
  • Continuing an existing video
Sizes
13.6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and sound2025

DiffRhythm

ASLP-lab (Northwestern Polytechnical University) · China

Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.

  • Songs and jingles from lyrics
  • Music for videos
  • Demo versions of tracks
Sizes
about 1.1B
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2025

AntiDeepfake (NII)

National Institute of Informatics, Yamagishi Lab · Japan

Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.

  • Checking audio recordings for synthesis
  • Fine-tuning for your own language and recording channel
  • Comparing several encoders on your own data
Sizes
0,3B – 2B
Hardware
from: Laptop
Non-commercial onlyDetails
3D2024–2025

Hunyuan3D

Tencent · China

Tencent's open 3D line: shape and texture from an image, at the level of paid services. Omni adds control of pose and shape, Part splits a model into parts.

  • Textured 3D product models
  • Characters and objects for games
  • Splitting a model into parts for printing
Sizes
set of models: shape and textures
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Speech to textRU2023–2025

Vosk (русские модели)

Alpha Cephei · Russia

Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.

  • Transcribing Russian calls and recordings without the cloud
  • Voice control in apps and kiosks
  • Low-latency streaming speech recognition
Sizes
about 45 MB – 1.8 GB
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistantsRUGGUF2025

Qwen Omni

Alibaba (Qwen) · China

Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.

  • Voice assistant for customers
  • Analyzing calls and videos
  • Voice answers about documents and images
Sizes
3B – 30B-A3B
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and sound2022–2025

pyannote (диаризация)

pyannoteAI (Hervé Bredin) · France

The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.

  • Tagging calls: which part is the agent, which is the customer
  • Meeting minutes with speaker labels
  • Preparing recordings for transcription and analysis
Sizes
a few million parameters
Hardware
from: Laptop
Commercial use allowedDetails
Music and sound2025

HunyuanVideo-Foley

Tencent Hunyuan · China

Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.

  • Foley and sound effects for video
  • Sound for AI-generated ads
  • Sound design for short videos
Sizes
not stated on the model card (weights about 10 GB; XL version with memory offloading)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Deepfake detection2023–2025

RADAR

IBM Research and The Chinese University of Hong Kong · USA

An AI-text detector trained together with a paraphraser: it was deliberately taught not to give up when the text has been rewritten. It errs in both directions; a human reviews the output.

  • Checking texts that may have been rewritten after generation
  • First-pass filtering in a newsroom or admissions office
  • Comparison against simpler detectors
Sizes
about 355M (RoBERTa-large)
Hardware
from: Laptop
Commercial use with conditionsDetails
Avatars2025

MultiTalk / InfiniteTalk

Meituan · China

Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.

  • Video dubbing with matched facial expressions
  • Dialogue between two characters from audio
  • Long videos with a presenter
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice: speakers and sound2024–2025

ClearerVoice (MossFormer)

Alibaba (Tongyi Lab) · China

Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.

  • Noise suppression in conversation recordings
  • Separating two voices speaking at once
  • Improving old and phone recordings
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRU2025

ESpeech-TTS

ESpeech (independent group of Russian-speaking developers) · Russia

Russian speech synthesis with voice cloning based on the F5-TTS architecture, trained on Russian speech datasets collected by the authors. Stress is placed automatically. Several variants, including a "podcaster" one.

  • Voicing videos and audiobooks in Russian
  • Cloning a narrator's voice from a sample
  • Voice for a bot or assistant in Russian
Sizes
about 340M
Hardware
from: Laptop
Commercial use with conditionsDetails
Video2025

Hunyuan-GameCraft

Tencent Hunyuan · China

Turns a single image into a controllable game-scene video: the camera moves on keyboard commands. Minimum 24 GB of GPU memory, 80 GB recommended.

  • Interactive video prototypes of game locations
  • Camera walkthrough videos of a scene
  • Level demos for pitches
Sizes
based on HunyuanVideo
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Deepfake detection2023–2025

DeepfakeBench

The Chinese University of Hong Kong, Shenzhen (SCLBD) · China

Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.

  • First-pass check of a submitted video or selfie
  • Comparing several detectors on your own data
  • Fine-tuning a detector for your own flow of applications
Sizes
Xception- and EfficientNet-class detectors, tens of millions of parameters
Hardware
from: Laptop
Non-commercial onlyDetails
TranslationRUGGUF2025

Seed-X

ByteDance Seed · China

A compact ByteDance translator for 28 languages, close in quality to large closed systems. Russian is supported. Ready-made compressed versions are available.

  • Translating business correspondence and documents
  • Translating product cards
  • Translating technical and legal texts
Sizes
7B
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingGGUF2024–2025

BiRefNet

Nankai University · China

An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.

  • Bulk background removal from product photos
  • Precise masks for design and print
  • Cutting out people with hair for advertising
Sizes
about 220M (lightweight lite versions available)
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2024–2025

Watermark Anything (WAM)

Meta · USA

An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.

  • Marking generated and edited images
  • Finding a marked fragment inside a collage
  • Tracking which parts of a picture were made by AI
Sizes
a mark encoder and decoder for images
Hardware
from: Laptop
Commercial use allowedDetails
Image generationGGUF2024–2025

OmniGen

BAAI (Beijing Academy of Artificial Intelligence) · China

An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.

  • Placing a product or person into a new scene
  • Instruction-based photo editing
  • Generation from multiple references
Sizes
about 4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Photo editing2025

SeedVR / SeedVR2

ByteDance Seed · China

ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.

  • Upscaling photos and video to 2K–4K
  • Restoring old videos and photos
  • Enhancing user photos before publishing
Sizes
3B – 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice: speakers and sound2024–2025

GPT-SoVITS

RVC-Boss and community · China

Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.

  • Voicing texts with a specific narrator's voice
  • Voice for a bot or assistant
  • Dubbing training videos
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Avatars2025

LatentSync

ByteDance · China

Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.

  • Dubbing videos into another language with lip sync
  • Editing lines in finished video without reshooting
  • Talking avatars for training courses
Sizes
requires 8–18 GB of VRAM
Hardware
from: Laptop
Commercial use with conditionsDetails
Deepfake detection2024–2025

AIDE

Xiaohongshu, USTC and Shanghai Jiao Tong University · China

An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.

  • Checking realistic AI images without obvious artifacts
  • Comparing detectors on hard examples
  • Fine-tuning for your own type of content
Sizes
several experts based on ConvNeXt and CLIP
Hardware
from: 1 GPU
Commercial use with conditionsDetails
AvatarsGGUF2025

HunyuanVideo-Avatar

Tencent · China

Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.

  • Presenter video from a photo and audio
  • Scenes with several speakers
  • Cartoon characters
Sizes
about 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Voice: speakers and sound2024–2025

Seed-VC

Songting Liu (Plachtaa) · China

Voice conversion without training: transfers timbre from a 1–30 second sample, can sing and work in real time; V2 also changes accent. Use only with the voice owner's consent.

  • Re-voicing a video with a different voice
  • Voice anonymization in recordings
  • Real-time voice for streams
Sizes
about 70M – 200M
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safety2023–2025

NSFW-классификаторы (Falconsai, Freepik)

Falconsai and Freepik · USA and Spain

Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.

  • Filtering user photos and avatars
  • Checking generated images before publishing
  • Labeling a media library
Sizes
86M
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2025

Kimi-Audio

Moonshot AI · China

A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.

  • Speech recognition
  • Detecting emotions and sound events
  • Speech-to-speech voice dialogue
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Speech to textRUGGUF2022–2025

Whisper

OpenAI · USA

Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.

  • Transcription of calls and video meetings
  • Video subtitles
  • Voice messages to text
Sizes
39M – 1,5B
Hardware
from: Laptop
Commercial use allowedDetails
Video2024–2025

Open-Sora

HPC-AI Tech · Singapore

A fully open video generation project: weights, code and training recipe. Version 2.0 at 11B makes video from text and from an image.

  • Video from a text description
  • Animating images
  • Training your own video model
Sizes
up to 11B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025

Step-Video

StepFun · China

A large 30B video model producing clips of up to 204 frames. Needs server hardware, but is open under MIT.

  • Video from a description
  • Animating images
Sizes
30B
Hardware
from: Cluster
Commercial use allowedDetails
Avatars2024–2025

MuseTalk

Tencent Music (Lyra Lab) · China

Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.

  • Dubbing video into another language
  • Live avatar in a video chat
  • Editing speech in a finished video
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechGGUF2024–2025

F5-TTS

Shanghai Jiao Tong University and partners · China

A voice cloning model that needs only a few seconds of a sample, in English and Chinese. The community has released many fine-tuned versions for other languages, including Russian.

  • Voice cloning
  • Voicing audiobooks and videos
  • Research and prototypes
Sizes
about 340M
Hardware
from: Laptop
Non-commercial onlyDetails
Text to speechGGUF2025

Orpheus TTS

Canopy Labs · USA

Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.

  • Real-time voice for an assistant
  • Emotional voiceover
  • Voice cloning
Sizes
3B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechGGUF2025

Sesame CSM

Sesame · USA

A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.

  • Voice for a conversational assistant
  • Voicing dialogues
  • Voice product prototypes
Sizes
1B
Hardware
from: Laptop
Commercial use allowedDetails
Faces2025

InfiniteYou

ByteDance · China

FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.

  • Portraits from one photo with a precise scene description
  • Testing characters for advertising
  • Comparing face-preservation methods
Sizes
adapter for FLUX.1-dev
Hardware
from: 1 GPU
Non-commercial onlyDetails
Deepfake detection2025

Desklib AI Text Detector

Desklib · India

A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.

  • Checking submitted articles and reports
  • Filtering templated reviews and applications
  • First-pass check of student work
Sizes
0.4B (DeBERTa-v3-large)
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2025

SigLIP (наследник CLIP)

Google · USA

Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.

  • Image search by text query
  • Automatic catalog labeling and tagging
  • Filtering prohibited content
Sizes
about 0.2B to 2B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechGGUF2024–2025

Kokoro

hexgrad (independent developer) · not disclosed

A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.

  • Voicing articles and notifications
  • Voice for apps without a GPU
  • Bulk text voiceover
Sizes
82M
Hardware
from: Laptop
Commercial use allowedDetails
3D2024–2025

Stable Fast 3D / SPAR3D

Stability AI · UK

Stability AI models that turn a single photo into a textured 3D model in about a second. SPAR3D lets you adjust the shape through a point cloud.

  • 3D product cards from photos
  • Assets for games and AR
  • Quick mock-ups for design
Sizes
1B – 2B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Photo editing2024–2025

BEN2

Prama LLC · USA

A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.

  • Cutting out products and people from photos
  • Background removal in video
  • Preparing photos for a catalog
Sizes
about 95M
Hardware
from: Laptop
Commercial use allowedDetails
Image + textGGUF2024–2025

Janus

DeepSeek · China

A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.

  • Answering questions about images
  • Draft illustrations from a description
  • Experiments with a unified vision and generation model
Sizes
1B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer vision2022–2025

ViTPose

University of Sydney and JD Explore Academy · Australia / China

A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.

  • Body keypoints in photos and video
  • Motion analysis in sports and rehabilitation
  • Monitoring work postures and safety practices
Sizes
33M – about 1B
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2023–2025

VideoLLaMA

Alibaba DAMO Academy · China

Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.

  • Video description and short summary
  • Finding a moment in a recording by question
  • Tagging a video archive
Sizes
2B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Virtual try-on2024

Leffa

Meta AI (with King's College London) · USA

Meta's model for virtual try-on and changing a person's pose in a photo. Carefully transfers fine fabric details and lettering. MIT license, but the training data is non-commercial.

  • Trying clothes on a model photo
  • Changing the model pose in an existing shot
  • Adding extra angles for a product card
Sizes
based on Stable Diffusion
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Music and sound2024

MMAudio

University of Illinois and Sony AI · USA / Japan

Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.

  • Sound effects for silent video
  • Sound for clips from AI generators
  • Draft sound design for editing
Sizes
size not stated on the model card
Hardware
from: Laptop
Non-commercial onlyDetails
Music and sound2024

MuQ / MuQ-MuLan

Tencent AI Lab · China

The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.

  • Searching music by text description
  • Tagging tracks by genre and mood
  • Finding similar music
Sizes
300M – 700M
Hardware
from: Laptop
Non-commercial onlyDetails
Deepfake detection2024

AudioSeal

Meta · USA

An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.

  • Marking speech synthesized by your service
  • Finding your own mark in third-party publications
  • Checking whether synthesis was mixed into a call recording
Sizes
a watermark generator and detector, 16-bit message
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2024

GME (General Multimodal Embedding)

Alibaba (Tongyi Lab) · China

One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.

  • Finding a product by photo
  • Search across a catalogue of images and cards
  • Search across document pages as images
Sizes
2B and 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text to speechGGUF2024

Parler-TTS

Hugging Face · USA

Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.

  • Choosing a voice by description
  • Voicing videos
  • Voice service prototypes
Sizes
880M – 2.2B
Hardware
from: Laptop
Commercial use allowedDetails
Music and sound2023–2024

MusicGen (AudioCraft)

Meta · USA

Generates instrumental music from a text description or a sample melody. One of the first open models of its kind.

  • Draft music sketches
  • Music for video prototypes
  • Research
Sizes
300M – 3.3B
Hardware
from: Laptop
Non-commercial onlyDetails
Image generationGGUF2022–2024

Stable Diffusion

Stability AI · UK

The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.

  • Illustrations and banners for advertising
  • Backgrounds and scenes for product cards
  • Fine-tuning to a brand style
Sizes
0.9B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
VideoGGUF2024

Mochi

Genmo · USA

An open 10B video model with realistic motion. At release it was among the strongest open models; no updates now.

  • Video from a description
  • Short ad scenes
Sizes
10B
Hardware
from: 1 GPU
Commercial use allowedDetails
Faces2024

PuLID

ByteDance · China

Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.

  • Portraits and avatars from one photo
  • Ad characters with a recognizable face
  • Photoshoot prototypes
Sizes
adapters for SDXL and FLUX.1-dev
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2022–2024

RIFE (Practical-RIFE)

hzwer (Zhewei Huang) and co-authors · China

Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.

  • Increasing video frame rate
  • Smooth slow-motion video
  • Smoothing clips from AI generators
Sizes
lightweight model (size not stated on the model card)
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2024

VLM2Vec

TIGER-Lab · Canada

Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.

  • Search across a mixed archive of texts and images
  • Search across document pages as images
  • Finding similar cards and illustrations
Sizes
about 4B (based on Phi-3.5-V)
Hardware
from: 1 GPU
Commercial use allowedDetails
Photo editing2024

Flux.1-dev ControlNet Upscaler

Jasper AI · USA

A FLUX add-on for upscaling small and blurry images with detail reconstruction. Popular, but under the non-commercial FLUX dev license.

  • Upscaling small images with detail reconstruction
  • Enhancing generated images
  • Upscaling pilots for a catalog
Sizes
add-on for FLUX.1 dev 12B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Image + textNot maintained2023–2024

IDEFICS

Hugging Face · France / USA

Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.

  • Answering questions about images
  • Analysing documents and screenshots
  • A base for fine-tuning
Sizes
8B – 80B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationNot maintained2023–2024

ControlNet

Lvmin Zhang (Stanford) and the community · USA

An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.

  • Image from a sketch or outline
  • Keeping pose and composition
  • Interior visualization from a floor plan
Sizes
0.4B – 1.3B
Hardware
from: Laptop
Commercial use allowedDetails
VideoNot maintained2023–2024

AnimateDiff

Shanghai AI Lab and CUHK · China

A module that brings Stable Diffusion image models to life, turning them into short animations. One of the first open video technologies.

  • Short animations in brand style
  • Animated covers and banners
  • Animated stickers
Sizes
motion module on top of SD 1.5 / SDXL
Hardware
from: Laptop
Commercial use allowedDetails
AvatarsNot maintained2024

LivePortrait

Kuaishou (Kling) · China

Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.

  • Animating portraits
  • Transferring an actor's expressions to a character
  • Mascot animation
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textNot maintained2024

Florence-2

Microsoft · USA

A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.

  • Reading text in photos
  • Finding and highlighting objects
  • Automatic photo captions
Sizes
0.23B – 0.77B
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2023–2024

PowerPaint

Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China

All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.

  • Removing and replacing objects in photos
  • Extending the frame to a required format
  • Inserting a product or detail from a text description
Sizes
based on SD 1.5
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2024

IC-Light

Lvmin Zhang (author of ControlNet) · USA

Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.

  • Matching product lighting to a new background
  • Studio lighting for portraits without a reshoot
  • Consistent lighting style across a catalog
Sizes
based on SD 1.5
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detectionNot maintained2024

MAGE

UC Santa Barbara and co-authors · USA

A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.

  • Checking long articles and reports as a whole
  • Filtering machine text in a publication flow
  • Comparing detectors on your own data
Sizes
about 150M (Longformer-base)
Hardware
from: Laptop
Commercial use allowedDetails
Image generationNot maintained2023–2024

PixArt

Huawei Noah's Ark Lab and partners · China

A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.

  • Illustrations for articles and social media
  • Backgrounds for product cards
  • Quick visual drafts
Sizes
0.6B
Hardware
from: Laptop
Commercial use allowedDetails
3DNot maintained2024

InstantMesh

Tencent ARC · China

Builds a 3D mesh from a single image in about 10 seconds: first it draws the object from several angles, then assembles the model from them.

  • 3D model of an object from a photo
  • Assets for games and visualizations
  • Prototypes for 3D printing
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice: speakers and soundNot maintained2024

OpenVoice

MyShell and MIT · USA

Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.

  • Voicing videos with the company narrator's voice
  • Voice bot with a recognizable brand voice
  • Transferring timbre onto existing speech synthesis
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
FacesNot maintained2023–2024

IP-Adapter-FaceID

Tencent AI Lab (h94) · China

One of the first adapters that transfer a face from a photo into a generated image. The SD 1.5 versions run on low-end cards, but the weights are non-commercial.

  • Portraits from a photo in different styles
  • Image series with one character
  • Avatar experiments
Sizes
adapters for SD 1.5 and SDXL
Hardware
from: Laptop
Non-commercial onlyDetails
Text to speechNot maintained2024

MeloTTS

MyShell and MIT · USA

Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.

  • Voicing bot replies in foreign languages
  • Voicing training materials
  • Reading texts aloud on a server without a GPU
Sizes
small, runs in real time on a CPU
Hardware
from: Laptop
Commercial use allowedDetails
Image generationNot maintained2023–2024

Playground

Playground AI · USA

An SDXL-based model focused on aesthetics: vivid colors, contrast, portraits. Compatible with SDXL ecosystem tools.

  • Aesthetic ad visuals
  • Portraits and lifestyle images
  • Post covers
Sizes
about 2.6B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Photo editingNot maintained2024

SUPIR

XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shanghai AI Lab and others) · China

Powerful SDXL-based restoration of badly damaged photos: it recreates details rather than just upscaling. Hardware-hungry; non-commercial license.

  • Restoring old and blurry photos
  • Upscaling with detail reconstruction
  • Archive restoration pilots
Sizes
based on SDXL, plus LLaVA 13B for captions
Hardware
from: 1 GPU
Non-commercial onlyDetails
FacesNot maintained2024

InstantID

InstantX (Xiaohongshu) · China

Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.

  • Portraits in different styles from one photo
  • Avatar and character sketches
  • Photoshoot prototypes
Sizes
adapter for SDXL
Hardware
from: 1 GPU
Non-commercial onlyDetails
Photo editingNot maintained2023

DDColor

Alibaba DAMO Academy · China

Colorizes black-and-white photos in natural colors. A lightweight model with a commercial-friendly license; a compact tiny version is available.

  • Colorizing archival photos
  • Color versions of historical photos for publications
  • Family photo restoration service
Sizes
DDColor-T (tiny) and DDColor-L
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundNot maintained2023

Resemble Enhance

Resemble AI · USA

A speech enhancement model: removes noise and restores lost frequencies so a muffled recording sounds studio-quality. Good for preparing a voice for voiceover.

  • Restoring old and phone recordings
  • Cleaning a voice before voiceover and cloning
  • Improving audio in videos and podcasts
Sizes
under 1B
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detectionNot maintained2023

HC3 ChatGPT Detector

Hello-SimpleAI · China

One of the first open AI-text classifiers, trained on the HC3 corpus of paired human and ChatGPT answers. It errs in both directions: its output is a reason to talk to the author, not proof.

  • First-pass check of student work
  • Filtering templated applications and reviews
  • Flagging suspicious texts for manual review
Sizes
about 125M (RoBERTa-base)
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speechNot maintained2023

StyleTTS 2

Columbia University · USA

A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.

  • Voicing texts in English
  • A base for fine-tuning your own voice
  • Voice service prototypes
Sizes
about 150M
Hardware
from: Laptop
Commercial use with conditionsDetails
TranslationRUGGUFNot maintained2023

MADLAD-400

Google · USA

Google's translator for more than 400 languages under a permissive license. Russian is supported. A good substitute for NLLB when commercial use is needed.

  • Translating documents and emails
  • Translating catalogs and product descriptions
  • Translating into CIS and Asian languages
Sizes
3B – 10B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRUNot maintained2023

Coqui XTTS

Coqui · Germany

A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.

  • Voice cloning from a sample
  • Multilingual voiceover
  • Research and prototypes
Sizes
about 470M
Hardware
from: Laptop
Non-commercial onlyDetails
Music and soundNot maintained2023

CLAP (LAION)

LAION · Germany

CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.

  • Search sounds and music by description
  • Automatic tags for an audio library
  • Recognizing sound types (siren, breaking glass, voice)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2022–2023

InSPyReNet / transparent-background

Taehoon Kim (POSTECH) · South Korea

A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.

  • Batch background removal from photos
  • Replacing the background with a color or blur
  • Background removal in video
Sizes
small (based on Swin-B)
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2022–2023

HAT

XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and others) · China

Transformer-based photo upscaling, more accurate than SwinIR on fine details. Versions for real noisy photos and a lightweight HAT-S.

  • Upscaling product and interior photos
  • Preparing images for print
  • Sharpening archival photos
Sizes
9M – 40M
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundNot maintained2022–2023

DeepFilterNet

Hendrik Schröter (University of Erlangen) · Germany

Lightweight real-time speech noise suppression that works even on a regular CPU and low-power devices. Removes hum, street and office noise while keeping the voice.

  • Cleaning calls and voice messages of noise
  • Preparing recordings before speech recognition
  • Noise suppression for video calls
Sizes
about 2M
Hardware
from: Laptop
Commercial use allowedDetails
TextRUGGUFNot maintained2023

ruGPT-3.5

Sber (ai-forever) · Russia

Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.

  • Base for fine-tuning on a narrow Russian-language task
  • Generating template Russian texts
  • Experiments with Russian-language models without license restrictions
Sizes
13B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text to speechRUGGUFNot maintained2023

Bark

Suno · USA

One of the first open models to voice text with intonation, laughter and pauses. Supports about ten languages, including Russian. Now outdated.

  • Draft voiceovers for videos
  • Voice service prototypes
  • Sound effects in speech
Sizes
about 300M – 1B
Hardware
from: Laptop
Commercial use allowedDetails
3DNot maintained2022–2023

Point-E / Shap-E

OpenAI · USA

Early open OpenAI models that create a 3D object from text or an image in seconds. Quality is basic, but they are fast and easy to run.

  • Rough 3D mock-ups from a description
  • Quick object prototypes for games and AR
  • Training and research pilots in 3D
Sizes
40M – 1B
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2023

StableSR

S-Lab, Nanyang Technological University · Singapore

One of the first Stable Diffusion-based photo upscalers: restores realistic details. Non-commercial license.

  • Upscaling photos with detail reconstruction
  • Research restoration pilots
  • Comparison with classic upscalers
Sizes
based on SD 2.1
Hardware
from: 1 GPU
Non-commercial onlyDetails
Photo editingNot maintained2022–2023

CodeFormer

S-Lab, Nanyang Technological University · Singapore

Popular face restoration for old and blurry photos; also works on video. Has face inpainting and colorization modes. Non-commercial license.

  • Restoring faces in old photos
  • Enhancing faces in low-quality video
  • Family archive restoration pilots
Sizes
small, up to 0.1B
Hardware
from: Laptop
Non-commercial onlyDetails
Music and soundNot maintained2022

AST (Audio Spectrogram Transformer)

MIT · USA

A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.

  • Sound event recognition
  • Tagging an audio archive
  • Detecting alarm sounds
Sizes
about 87M
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingGGUFNot maintained2021–2022

Real-ESRGAN

Tencent ARC Lab · China

The classic for upscaling photos 2–4x while cleaning noise and compression artifacts. Lightweight, runs even on a CPU. Versions for drawings and anime.

  • Upscaling old and small product photos
  • Cleaning images of compression artifacts
  • Preparing images for print
Sizes
about 17M
Hardware
from: Laptop
Commercial use allowedDetails
Computer visionNot maintained2021–2022

CLIP (OpenAI)

OpenAI · USA

The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.

  • Image search by text query
  • Automatic tags for a catalog
  • Finding similar images
Sizes
about 0.15B – 0.6B
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2021–2022

GFPGAN

Tencent ARC Lab · China

Proven face restoration for old and compressed photos, with a commercial-friendly license. Often paired with Real-ESRGAN; the most used versions are 1.3 and 1.4.

  • Restoring faces in old photos
  • Enhancing avatars and profile photos
  • Restoration in a photo shop or online service
Sizes
small, up to 0.1B
Hardware
from: Laptop
Commercial use allowedDetails
Computer visionRUNot maintained2022

ruCLIP (ai-forever)

Sber AI and SberDevices (ai-forever) · Russia

A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.

  • Product search by photo and by Russian description
  • Sorting images into categories without labeling
  • Checking that a photo matches its caption
Sizes
150M – 430M
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingGGUFNot maintained2021

LaMa

Samsung AI Center Moscow (with Skoltech) · Russia

Removes unwanted objects, text and watermarks from photos with clean background fill. Lightweight and fast; still the standard for this task.

  • Removing price tags, people and clutter from photos
  • Cleaning interior and real estate photos
  • Removing text and dates from archival photos
Sizes
about 51M
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2021

SwinIR

ETH Zurich · Switzerland

A transformer model for upscaling, denoising and removing JPEG artifacts from photos. Lightweight and proven; often embedded in other systems.

  • Photo upscaling
  • Image denoising
  • Removing compression artifacts
Sizes
about 12M
Hardware
from: Laptop
Commercial use allowedDetails
AvatarsNot maintained2020

Wav2Lip

IIIT Hyderabad · India

The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.

  • Quick dubbing tests
  • Comparison with newer lip-sync models
  • Educational and research projects
Sizes
small model, 96-pixel face
Hardware
from: Laptop
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment