Image generationGGUF2025–2026
Alibaba · China
Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.
- Infographics for product cards
- Photo editing by text command
- Ad creatives
- Sizes
- 7B – 20B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2025–2026
Microsoft · USA
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
- Transcribing long meetings with speaker labels
- Voicing podcasts and dialogues
- Real-time voice for assistants
- Sizes
- 0.5B – 9B
- Hardware
- from: Laptop
Music and soundGGUF2025–2026
M-A-P and HKUST · China
Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.
- Songs and jingles from lyrics
- Demo versions of tracks
- Music for videos
- Sizes
- 0.5B – 7B
- Hardware
- from: 1 GPU
Music and sound2026
HeartMuLa Team · not disclosed
An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.
- Songs and jingles from lyrics
- Music for videos
- Transcribing song lyrics
- Sizes
- 3B
- Hardware
- from: 1 GPU
TextRU2022–2026
Yandex · Russia
Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.
- Russian-language assistant and chatbot
- Answers based on the company knowledge base
- Base for industry-specific fine-tuning
- Sizes
- 8B – 100B
- Hardware
- from: Laptop
3D2024–2026
NAVER LABS Europe · France (NAVER, South Korea)
The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.
- 3D scene from several photos without calibration
- Point matching between images
- Mapping from video (SLAM)
- Sizes
- 0.57B – 0.69B
- Hardware
- from: Laptop
Music and soundGGUF2022–2026
m-a-p (Multimodal Art Projection) · UK / China
A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.
- Automatic tagging of a music catalog
- Finding similar tracks
- Detecting genre, mood and tempo
- Sizes
- 95M – 632M
- Hardware
- from: Laptop
Voice assistants2026
Samsung · South Korea
Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.
- Describing an audio recording in words
- Answering questions about a sound
- Identifying the type of sound and the setting
- Sizes
- 99M – 356M
- Hardware
- from: Laptop
Deepfake detection2023–2026
Adobe Research and University of Surrey · USA
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
- Marking images on the way out of your own pipeline
- Checking the provenance of a submitted image
- Linking with content provenance metadata
- Sizes
- model types Q and P with different mark capacity
- Hardware
- from: Laptop
VideoGGUF2025–2026
Sand AI · China
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
- Long videos with continuation
- Video with sound
- Animating images
- Sizes
- 4.5B – 114B-A6B
- Hardware
- from: 1 GPU
Computer vision2024–2026
Microsoft Research · USA
Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.
- Measuring rooms and objects from photos
- 3D point cloud from a single shot
- Preparing data for robots and AR
- Sizes
- ViT-S – ViT-G
- Hardware
- from: Laptop
Speech to textGGUF2024–2026
Moonshine AI (Useful Sensors) · USA
Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.
- Voice control of devices
- Offline recognition on a phone
- Live subtitles
- Sizes
- 27M – 245M
- Hardware
- from: Laptop
Text to speechGGUF2025–2026
bilibili · China
Speech synthesis with voice cloning and precise duration control, handy for video dubbing. Controls emotion separately from timbre.
- Video dubbing matched to timing
- Voice cloning
- Emotional voiceover
- Sizes
- about 1B – 2B
- Hardware
- from: Laptop
Text2024–2026
Tencent · China
Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.
- Corporate assistant
- Translation and multilingual texts
- Agents with tools
- Sizes
- 0.5B – 770B-A49B
- Hardware
- from: Laptop
Voice: speakers and sound2022–2026
Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community
A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.
- Clean vocals from a recording with music
- Backing tracks and stems for karaoke
- Removing background music and noise from videos
- Sizes
- from tens to hundreds of millions of parameters
- Hardware
- from: Laptop
Deepfake detection2025–2026
University of Michigan · USA
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
- Checking submitted photos and illustrations
- Filtering AI images in a content flow
- Flagging suspicious images for manual review
- Sizes
- 22M
- Hardware
- from: Laptop
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
Video2025–2026
Alibaba · China
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
- Short promo videos
- Animating product photos
- Videos for social media
- Sizes
- 1,3B – 14B
- Hardware
- from: 1 GPU
Image generationRU2022–2026
Sber (Kandinsky Lab) · Russia
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
- Images from Russian-language descriptions
- Short promo videos from text or a photo
- Instruction-based image editing
- Sizes
- 2B – 19B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
lodestones (independent developer) · not disclosed
A community model retrained from FLUX.1-schnell with a simplified architecture. No style censorship; popular as a base for fine-tuning.
- Base for fine-tuning your own styles
- Artistic illustrations
- Images for games
- Sizes
- 4B – 8.9B
- Hardware
- from: 1 GPU
Video2024–2026
Lightricks · Israel
A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.
- Ad videos with sound
- Video from a product photo
- Voiced scenes for social media
- Sizes
- 2B – 22B
- Hardware
- from: 1 GPU
Video2026
MiniMax · China
Open weights of MiniMax's Hailuo video model. A large 33B model that makes video from text and images, but needs several server GPUs.
- Cinematic ad videos
- Video from text and images
- Complex scenes with motion
- Sizes
- 33B + 32B encoder
- Hardware
- from: Cluster
Robotics2025–2026
NVIDIA · USA
"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.
- Synthetic video for training robots and self-driving vehicles
- Testing scenarios in simulation
- Robot control (Policy versions)
- Sizes
- 2B – 65B
- Hardware
- from: 1 GPU
Voice: speakers and sound2022–2026
Meta AI, then Alexandre Défossez · France
A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.
- Separating vocals from music in a recording
- Backing tracks and karaoke stems
- Cleaning speech in videos with background music
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2023–2026
RVC-Project community · China
The most widely used open voice conversion tool: a model for a specific voice trains on 10–30 minutes of recording and works in real time. Use only with the voice owner's consent.
- Voicing content with one brand voice
- Covers and vocal work
- Real-time voice changing
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Image + text2026
OpenMOSS (Fudan University) · China
An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.
- Analyzing long videos and finding events by time
- Real-time streaming video analysis
- Understanding photos and documents
- Sizes
- about 11B
- Hardware
- from: 1 GPU
Image generationGGUF2026
Krea · USA
A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.
- Realistic photos for advertising
- High-resolution images
- Style fine-tuning
- Sizes
- 12B
- Hardware
- from: 1 GPU
Image generationGGUF2026
Ideogram · Canada
Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.
- Testing text-in-image generation
- Research and prototypes
- Sizes
- about 9B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2026
Alibaba (Qwen) · China
Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.
- Transcribing calls and meetings
- Video subtitles
- Multilingual recognition
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
Speech to textGGUF2026
Cohere · Canada
Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.
- Transcribing meetings and interviews
- Subtitles
- Searching an audio archive
- Sizes
- 2B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Boson AI · USA
Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.
- Expressive video voiceover
- Voicing dialogues
- Voice cloning
- Sizes
- about 3B – 8B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
Zyphra · USA
Speech synthesis with voice cloning and fine control over emotion, speed and pitch.
- Voice cloning
- Emotional voiceover
- Voicing videos
- Sizes
- about 1.6B
- Hardware
- from: Laptop
Music and soundRUGGUF2025–2026
ACE Studio and StepFun · China
Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.
- Songs and jingles for ads
- Background music for videos
- Demo versions of tracks
- Sizes
- about 2B – 4B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
MiniMax · China
Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.
- Analysis of large document archives in a single request
- Agents with tools
- Help for developers
- Sizes
- 230B-A10B – 456B-A46B
- Hardware
- from: Cluster
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
VideoGGUF2025–2026
Zhipu AI (Z.ai) and Tsinghua University · China
Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.
- Transferring an actor's motion to a character
- Replacing a character in a finished video
- Animating mascots and illustrations
- Sizes
- 14B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
HiDream.ai · China
Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.
- Image generation from descriptions
- Editing images with words
- Variations of product photos
- Sizes
- about 9B – 17B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Meituan · China
Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.
- News or course presenter videos
- Promo videos with a talking character
- Singing and voice-over
- Sizes
- based on LongCat-Video 13.6B
- Hardware
- from: 1 GPU
3D2024–2026
VAST (TripoSR together with Stability AI) · China
VAST family: a 3D model from a single photo. TripoSR runs in under a second, TripoSG gives cleaner geometry, TripoSplat builds a scene from Gaussian points.
- 3D product model from a photo
- Object assets for games and AR
- Quick 3D prototype for printing
- Sizes
- up to 1.5B
- Hardware
- from: Laptop
Speech to textRU2023–2026
NVIDIA · USA
Fast NVIDIA speech recognition models, including streaming ones for real-time use. Parakeet TDT v3 and Nemotron 3.5 ASR understand Russian.
- Transcribing calls and meetings
- Video subtitles
- Real-time voice input
- Sizes
- 110M – 2.5B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Resemble AI · USA
Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- about 350M – 500M
- Hardware
- from: Laptop
Text to speechRU2025–2026
OpenMOSS (Fudan University) · China
A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.
- Voicing podcasts and dialogues
- Voice for an assistant
- Voice cloning
- Sizes
- 100M – 8.5B
- Hardware
- from: Laptop
Music and soundGGUF2024–2026
Stability AI · UK
Generates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.
- Sound effects for videos and games
- Background music and jingles
- Interface sounds
- Sizes
- about 0.5B – 2.3B
- Hardware
- from: Laptop
Photo editing2023–2026
BRIA AI · Israel
BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.
- Cutting products out onto a white background
- Staff and expert photos without background
- Background removal in video
- Sizes
- 44M – 220M
- Hardware
- from: Laptop
Image + textGGUF2025–2026
Kuaishou · China
Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.
- Analysing and describing short videos
- Reviewing clips and content
- Finding the right moment in a video
- Sizes
- 8B – 671B-A37B
- Hardware
- from: Laptop
Text to speechRU2025–2026
Supertone · South Korea
Very fast, lightweight speech synthesis that runs directly on the device, without a GPU or the cloud. Supertonic 3 speaks 31 languages, including Russian.
- Voicing voice bot replies on an ordinary server
- Voiceover in offline and mobile apps
- Reading texts and notifications aloud
- Sizes
- about 99M
- Hardware
- from: Laptop
3D2025–2026
Meta and the University of Oxford (VGG) · USA / UK
Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.
- 3D model of a room or object from a photo series
- Camera pose estimation for photogrammetry
- Point cloud for measurements and comparison with the plan
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
Computer vision2024–2026
Meta · USA
Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.
- Pose and body keypoint detection
- Segmentation of body parts and clothing
- Separating a person from the background
- Sizes
- 0.1B – 5B
- Hardware
- from: Laptop
Avatars2024–2026
Fudan University · China
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
- Presenter video from a photo and audio
- Long training videos
- Live avatar
- Sizes
- about 1B – 5B
- Hardware
- from: 1 GPU
3D2025–2026
Tencent · China
Generates whole 3D worlds and scenes from text or an image that you can walk through. The second version builds a scene from video and photos.
- 3D scenes for games and virtual tours
- Backgrounds and environments for video production
- Draft locations for simulations
- Sizes
- set of several models
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.
- Voice cloning
- Voicing videos and audiobooks
- Voice for an assistant
- Sizes
- 0.5B – 2.3B
- Hardware
- from: Laptop
Text to speech2025–2026
Kyutai · France
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
- Streaming speech transcription for voice bots
- Voicing replies with minimal delay
- Speech synthesis on a server without a GPU (Pocket TTS)
- Sizes
- 100M (Pocket TTS) – 2.6B
- Hardware
- from: Laptop
Voice assistants2024–2026
NVIDIA · USA
Models that listen to speech, sounds and music and answer questions about them. Audio Flamingo Next handles recordings up to 30 minutes. Research use only.
- Detailed descriptions of audio recordings
- Questions and answers about a long recording
- Tagging music and sounds
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Computer vision2023–2026
Meta · USA
Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.
- Background removal from product photos
- Counting objects in photos
- Data labeling for training
- Sizes
- 91M – ~0,85B
- Hardware
- from: Laptop
Speech to textRUGGUF2025–2026
Mistral AI · France
Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.
- Transcribing and summarizing recordings
- Asking questions about audio
- Real-time recognition
- Sizes
- 3B – 24B
- Hardware
- from: Laptop
Text to speechRU2024–2026
Fish Audio · USA / China
Speech synthesis with voice cloning and emotion control in 80+ languages, including Russian. Quality is close to paid services, but the weights are for research only.
- Voice cloning
- Emotional voiceover
- Multilingual voiceover
- Sizes
- 0.5B – about 4.5B
- Hardware
- from: Laptop
Text2025–2026
Reka AI · USA
Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.
- Photo and video analysis (Edge)
- Object detection in images
- Reasoning tasks (Flash)
- Sizes
- 7B – 21B
- Hardware
- from: Laptop
Voice: speakers and soundRU2026
FireRedTeam (Xiaohongshu) · China
A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.
- Cutting recordings before speech recognition
- Separating speech from music and singing in broadcasts and videos
- Speech detection in voice bots
- Sizes
- compact, exact size not stated
- Hardware
- from: Laptop
Text to speechRU2026
k2-fsa (Next-gen Kaldi) · China
Speech synthesis with voice cloning from a short sample in 646 languages, including Russian and languages of Russia's peoples. A voice can be described in words. Weights are for non-commercial use only.
- Voiceover in rare languages
- Voice cloning from a sample
- Research and prototypes of multilingual voiceover
- Sizes
- 0.6B
- Hardware
- from: Laptop
Video2025–2026
Skywork AI (Kunlun) · China
An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.
- Game world prototypes without an engine
- Interactive demos and simulations
- Generating data to train agents
- Sizes
- 1.8B – 17B
- Hardware
- from: 1 GPU
Photo editing2025–2026
S-Lab, Nanyang Technological University · Singapore
Cuts a person out of video with a precise alpha mask, including hair and edges, without a green screen. Needs a first-frame mask, for example from SAM.
- Background replacement in video without chroma key
- Cutting out a person for editing and effects
- Preparing videos for advertising and social media
- Sizes
- about 35M
- Hardware
- from: Laptop
Music and sound2025–2026
Alibaba Tongyi (FunAudioLLM) · China
Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
- Audio for video based on the scene
- Editing individual sounds in a track
- Sound effects from a description
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Avatars2026
SII-GAIR and Sand.ai · China
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
- Presenter video from a script
- Ad videos with a talking character
- Training videos with a narrator
- Sizes
- 15B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2026
Black Forest Labs · Germany
Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.
- Images for product cards
- Banners and covers
- Photo editing by description (Kontext)
- Sizes
- 4B – 32B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2026
Tencent · China
Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.
- Complex scenes from long descriptions
- Images with Chinese and English text
- Instruction-based image editing
- Sizes
- 1.5B – 80B-A13B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
Alibaba (Tongyi-MAI) · China
A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.
- Photorealistic ad images
- Images with English and Chinese text
- Bulk visual generation
- Sizes
- 6B
- Hardware
- from: 1 GPU
VideoGGUF2025–2026
Skywork AI (Kunlun Tech) · China
Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.
- Long videos with continuation
- Video with one character from a reference
- Talking avatar from a voice
- Sizes
- 1.3B – 19B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
VideoGGUF2026
OpenMOSS / MOSI · China
Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.
- Short clips with speech and sound from a description
- Ad scenes with dialogue
- Video prototypes for storyboards
- Sizes
- 32B-A18B
- Hardware
- from: 1 GPU
3DGGUF2024–2025
Microsoft · USA
One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.
- 3D models of products and interiors from photos
- Assets for games and AR/VR
- Prototypes for 3D printing
- Sizes
- up to 4B (TRELLIS.2)
- Hardware
- from: 1 GPU
Computer visionGGUF2024–2025
ByteDance and the University of Hong Kong · China
Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.
- Estimating distances and volumes from a camera
- Depth effects for photo and video
- Navigation for robots and drones
- Sizes
- 25M – 1.4B
- Hardware
- from: Laptop
Speech to textRU2025
Meta · USA
Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.
- Transcription in rare and local languages
- Digitizing oral archives
- Subtitles in many languages
- Sizes
- 300M – 7B
- Hardware
- from: Laptop
Speech to text2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Alibaba's set of fast speech recognition models, primarily for Chinese and Asian languages. SenseVoice also detects emotions and sound events.
- Transcribing calls
- Detecting emotions in the voice
- Recognizing laughter, music and other sounds
- Sizes
- about 230M – 800M
- Hardware
- from: Laptop
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
Text2025
Naver · South Korea
Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.
- Lightweight Korean-English assistant
- Analysis of images and documents
- Text classification
- Sizes
- 0.5B – 32B
- Hardware
- from: Laptop
Voice: speakers and sound2025
Meta · USA
A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.
- Isolating one person's voice from a noisy recording
- Removing unwanted sound from a video
- Splitting a recording into separate sound sources
- Sizes
- small, base, large (5 to 15 GB of weights)
- Hardware
- from: 1 GPU
3D2025
Meta and Carnegie Mellon University · USA
A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.
- 3D reconstruction of an object or room from photos
- Exporting the scene to COLMAP format for further processing
- Depth and camera pose estimation
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
3D2025
Shanghai AI Lab · China
Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.
- 3D scene reconstruction from video
- Camera pose estimation from frames
- Point clouds for research and prototypes
- Sizes
- 0.96B – 1.4B
- Hardware
- from: 1 GPU
3D2025
Tencent Hunyuan · China
Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.
- Character animation from a text description
- Draft animation for games and videos
- Motion library for avatars
- Sizes
- 0.46B – 1B
- Hardware
- from: 1 GPU
Computer vision2025
Meta · USA
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
- Search photos and videos by description
- Catalog labeling and tagging
- Search across audio and video (PE-AV)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
- Marking video created or processed by AI
- Finding your own mark in re-uploaded clips
- Protecting ad materials from being reused as someone else's
- Sizes
- a mark of 96 to 1024 bits
- Hardware
- from: Laptop
VideoGGUF2024–2025
Tencent · China
Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.
- Video from a text script
- Animating images
- Base for fine-tuning your own video models
- Sizes
- 8.3B – 13B
- Hardware
- from: 1 GPU
3DGGUF2025
Meta · USA
Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.
- 3D model of an item from a catalog photo
- Estimating body pose and shape from a photo
- Try-on and AR scenarios
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Text to speech2025
Nari Labs · South Korea
A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.
- Voicing dialogues and podcasts
- Ads with natural speech
- Training role-plays
- Sizes
- 1B – 2B
- Hardware
- from: Laptop
Voice: speakers and soundRU2020–2025
Silero · Russia
The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.
- Cutting calls and recordings before speech recognition
- Detecting when the customer is speaking in a voice bot
- Filtering out silence and noise to save on transcription
- Sizes
- about 2 MB
- Hardware
- from: Laptop
Video2025
Character.AI · USA
Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
- Short clips with talking characters
- Animating an image with voice-over
- Ad scene prototypes
- Sizes
- 11B
- Hardware
- from: 1 GPU
Computer vision2023–2025
Meta · USA
Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.
- Image search by text
- Image classification without training
- Search research and prototypes
- Sizes
- 0.15B – 3.6B
- Hardware
- from: Laptop
VideoGGUF2025
Meituan · China
A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.
- Long videos
- Video from a photo
- Continuing an existing video
- Sizes
- 13.6B
- Hardware
- from: 1 GPU
Music and sound2025
ASLP-lab (Northwestern Polytechnical University) · China
Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.
- Songs and jingles from lyrics
- Music for videos
- Demo versions of tracks
- Sizes
- about 1.1B
- Hardware
- from: Laptop
Deepfake detection2025
National Institute of Informatics, Yamagishi Lab · Japan
Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.
- Checking audio recordings for synthesis
- Fine-tuning for your own language and recording channel
- Comparing several encoders on your own data
- Sizes
- 0,3B – 2B
- Hardware
- from: Laptop
3D2024–2025
Tencent · China
Tencent's open 3D line: shape and texture from an image, at the level of paid services. Omni adds control of pose and shape, Part splits a model into parts.
- Textured 3D product models
- Characters and objects for games
- Splitting a model into parts for printing
- Sizes
- set of models: shape and textures
- Hardware
- from: 1 GPU
Speech to textRU2023–2025
Alpha Cephei · Russia
Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.
- Transcribing Russian calls and recordings without the cloud
- Voice control in apps and kiosks
- Low-latency streaming speech recognition
- Sizes
- about 45 MB – 1.8 GB
- Hardware
- from: Laptop
Voice assistantsRUGGUF2025
Alibaba (Qwen) · China
Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.
- Voice assistant for customers
- Analyzing calls and videos
- Voice answers about documents and images
- Sizes
- 3B – 30B-A3B
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
pyannoteAI (Hervé Bredin) · France
The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.
- Tagging calls: which part is the agent, which is the customer
- Meeting minutes with speaker labels
- Preparing recordings for transcription and analysis
- Sizes
- a few million parameters
- Hardware
- from: Laptop
Music and sound2025
Tencent Hunyuan · China
Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.
- Foley and sound effects for video
- Sound for AI-generated ads
- Sound design for short videos
- Sizes
- not stated on the model card (weights about 10 GB; XL version with memory offloading)
- Hardware
- from: 1 GPU
Deepfake detection2023–2025
IBM Research and The Chinese University of Hong Kong · USA
An AI-text detector trained together with a paraphraser: it was deliberately taught not to give up when the text has been rewritten. It errs in both directions; a human reviews the output.
- Checking texts that may have been rewritten after generation
- First-pass filtering in a newsroom or admissions office
- Comparison against simpler detectors
- Sizes
- about 355M (RoBERTa-large)
- Hardware
- from: Laptop
Avatars2025
Meituan · China
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
- Video dubbing with matched facial expressions
- Dialogue between two characters from audio
- Long videos with a presenter
- Sizes
- 14B
- Hardware
- from: 1 GPU
Voice: speakers and sound2024–2025
Alibaba (Tongyi Lab) · China
Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.
- Noise suppression in conversation recordings
- Separating two voices speaking at once
- Improving old and phone recordings
- Sizes
- under 1B
- Hardware
- from: Laptop
Text to speechRU2025
ESpeech (independent group of Russian-speaking developers) · Russia
Russian speech synthesis with voice cloning based on the F5-TTS architecture, trained on Russian speech datasets collected by the authors. Stress is placed automatically. Several variants, including a "podcaster" one.
- Voicing videos and audiobooks in Russian
- Cloning a narrator's voice from a sample
- Voice for a bot or assistant in Russian
- Sizes
- about 340M
- Hardware
- from: Laptop
Video2025
Tencent Hunyuan · China
Turns a single image into a controllable game-scene video: the camera moves on keyboard commands. Minimum 24 GB of GPU memory, 80 GB recommended.
- Interactive video prototypes of game locations
- Camera walkthrough videos of a scene
- Level demos for pitches
- Sizes
- based on HunyuanVideo
- Hardware
- from: 1 GPU
Deepfake detection2023–2025
The Chinese University of Hong Kong, Shenzhen (SCLBD) · China
Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.
- First-pass check of a submitted video or selfie
- Comparing several detectors on your own data
- Fine-tuning a detector for your own flow of applications
- Sizes
- Xception- and EfficientNet-class detectors, tens of millions of parameters
- Hardware
- from: Laptop
TranslationRUGGUF2025
ByteDance Seed · China
A compact ByteDance translator for 28 languages, close in quality to large closed systems. Russian is supported. Ready-made compressed versions are available.
- Translating business correspondence and documents
- Translating product cards
- Translating technical and legal texts
- Sizes
- 7B
- Hardware
- from: Laptop
Photo editingGGUF2024–2025
Nankai University · China
An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.
- Bulk background removal from product photos
- Precise masks for design and print
- Cutting out people with hair for advertising
- Sizes
- about 220M (lightweight lite versions available)
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
- Marking generated and edited images
- Finding a marked fragment inside a collage
- Tracking which parts of a picture were made by AI
- Sizes
- a mark encoder and decoder for images
- Hardware
- from: Laptop
Image generationGGUF2024–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.
- Placing a product or person into a new scene
- Instruction-based photo editing
- Generation from multiple references
- Sizes
- about 4B
- Hardware
- from: 1 GPU
Photo editing2025
ByteDance Seed · China
ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.
- Upscaling photos and video to 2K–4K
- Restoring old videos and photos
- Enhancing user photos before publishing
- Sizes
- 3B – 7B
- Hardware
- from: 1 GPU
Voice: speakers and sound2024–2025
RVC-Boss and community · China
Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.
- Voicing texts with a specific narrator's voice
- Voice for a bot or assistant
- Dubbing training videos
- Sizes
- under 1B
- Hardware
- from: Laptop
Avatars2025
ByteDance · China
Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.
- Dubbing videos into another language with lip sync
- Editing lines in finished video without reshooting
- Talking avatars for training courses
- Sizes
- requires 8–18 GB of VRAM
- Hardware
- from: Laptop
Deepfake detection2024–2025
Xiaohongshu, USTC and Shanghai Jiao Tong University · China
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
- Checking realistic AI images without obvious artifacts
- Comparing detectors on hard examples
- Fine-tuning for your own type of content
- Sizes
- several experts based on ConvNeXt and CLIP
- Hardware
- from: 1 GPU
AvatarsGGUF2025
Tencent · China
Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.
- Presenter video from a photo and audio
- Scenes with several speakers
- Cartoon characters
- Sizes
- about 13B
- Hardware
- from: 1 GPU
Voice: speakers and sound2024–2025
Songting Liu (Plachtaa) · China
Voice conversion without training: transfers timbre from a 1–30 second sample, can sing and work in real time; V2 also changes accent. Use only with the voice owner's consent.
- Re-voicing a video with a different voice
- Voice anonymization in recordings
- Real-time voice for streams
- Sizes
- about 70M – 200M
- Hardware
- from: Laptop
Moderation and safety2023–2025
Falconsai and Freepik · USA and Spain
Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.
- Filtering user photos and avatars
- Checking generated images before publishing
- Labeling a media library
- Sizes
- 86M
- Hardware
- from: Laptop
Voice assistants2025
Moonshot AI · China
A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.
- Speech recognition
- Detecting emotions and sound events
- Speech-to-speech voice dialogue
- Sizes
- 7B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2022–2025
OpenAI · USA
Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.
- Transcription of calls and video meetings
- Video subtitles
- Voice messages to text
- Sizes
- 39M – 1,5B
- Hardware
- from: Laptop
Video2024–2025
HPC-AI Tech · Singapore
A fully open video generation project: weights, code and training recipe. Version 2.0 at 11B makes video from text and from an image.
- Video from a text description
- Animating images
- Training your own video model
- Sizes
- up to 11B
- Hardware
- from: 1 GPU
Video2025
StepFun · China
A large 30B video model producing clips of up to 204 frames. Needs server hardware, but is open under MIT.
- Video from a description
- Animating images
- Sizes
- 30B
- Hardware
- from: Cluster
Avatars2024–2025
Tencent Music (Lyra Lab) · China
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
- Dubbing video into another language
- Live avatar in a video chat
- Editing speech in a finished video
- Sizes
- under 1B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
Shanghai Jiao Tong University and partners · China
A voice cloning model that needs only a few seconds of a sample, in English and Chinese. The community has released many fine-tuned versions for other languages, including Russian.
- Voice cloning
- Voicing audiobooks and videos
- Research and prototypes
- Sizes
- about 340M
- Hardware
- from: Laptop
Text to speechGGUF2025
Canopy Labs · USA
Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.
- Real-time voice for an assistant
- Emotional voiceover
- Voice cloning
- Sizes
- 3B
- Hardware
- from: Laptop
Text to speechGGUF2025
Sesame · USA
A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.
- Voice for a conversational assistant
- Voicing dialogues
- Voice product prototypes
- Sizes
- 1B
- Hardware
- from: Laptop
Faces2025
ByteDance · China
FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.
- Portraits from one photo with a precise scene description
- Testing characters for advertising
- Comparing face-preservation methods
- Sizes
- adapter for FLUX.1-dev
- Hardware
- from: 1 GPU
Deepfake detection2025
Desklib · India
A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.
- Checking submitted articles and reports
- Filtering templated reviews and applications
- First-pass check of student work
- Sizes
- 0.4B (DeBERTa-v3-large)
- Hardware
- from: Laptop
Computer vision2023–2025
Google · USA
Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.
- Image search by text query
- Automatic catalog labeling and tagging
- Filtering prohibited content
- Sizes
- about 0.2B to 2B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
hexgrad (independent developer) · not disclosed
A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.
- Voicing articles and notifications
- Voice for apps without a GPU
- Bulk text voiceover
- Sizes
- 82M
- Hardware
- from: Laptop
3D2024–2025
Stability AI · UK
Stability AI models that turn a single photo into a textured 3D model in about a second. SPAR3D lets you adjust the shape through a point cloud.
- 3D product cards from photos
- Assets for games and AR
- Quick mock-ups for design
- Sizes
- 1B – 2B
- Hardware
- from: 1 GPU
Photo editing2024–2025
Prama LLC · USA
A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.
- Cutting out products and people from photos
- Background removal in video
- Preparing photos for a catalog
- Sizes
- about 95M
- Hardware
- from: Laptop
Image + textGGUF2024–2025
DeepSeek · China
A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.
- Answering questions about images
- Draft illustrations from a description
- Experiments with a unified vision and generation model
- Sizes
- 1B – 7B
- Hardware
- from: Laptop
Computer vision2022–2025
University of Sydney and JD Explore Academy · Australia / China
A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.
- Body keypoints in photos and video
- Motion analysis in sports and rehabilitation
- Monitoring work postures and safety practices
- Sizes
- 33M – about 1B
- Hardware
- from: Laptop
Image + text2023–2025
Alibaba DAMO Academy · China
Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
- Video description and short summary
- Finding a moment in a recording by question
- Tagging a video archive
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Virtual try-on2024
Meta AI (with King's College London) · USA
Meta's model for virtual try-on and changing a person's pose in a photo. Carefully transfers fine fabric details and lettering. MIT license, but the training data is non-commercial.
- Trying clothes on a model photo
- Changing the model pose in an existing shot
- Adding extra angles for a product card
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
Music and sound2024
University of Illinois and Sony AI · USA / Japan
Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.
- Sound effects for silent video
- Sound for clips from AI generators
- Draft sound design for editing
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Music and sound2024
Tencent AI Lab · China
The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.
- Searching music by text description
- Tagging tracks by genre and mood
- Finding similar music
- Sizes
- 300M – 700M
- Hardware
- from: Laptop
Deepfake detection2024
Meta · USA
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
- Marking speech synthesized by your service
- Finding your own mark in third-party publications
- Checking whether synthesis was mixed into a call recording
- Sizes
- a watermark generator and detector, 16-bit message
- Hardware
- from: Laptop
Visual document search2024
Alibaba (Tongyi Lab) · China
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
- Finding a product by photo
- Search across a catalogue of images and cards
- Search across document pages as images
- Sizes
- 2B and 7B
- Hardware
- from: 1 GPU
Text to speechGGUF2024
Hugging Face · USA
Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.
- Choosing a voice by description
- Voicing videos
- Voice service prototypes
- Sizes
- 880M – 2.2B
- Hardware
- from: Laptop
Music and sound2023–2024
Meta · USA
Generates instrumental music from a text description or a sample melody. One of the first open models of its kind.
- Draft music sketches
- Music for video prototypes
- Research
- Sizes
- 300M – 3.3B
- Hardware
- from: Laptop
Image generationGGUF2022–2024
Stability AI · UK
The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.
- Illustrations and banners for advertising
- Backgrounds and scenes for product cards
- Fine-tuning to a brand style
- Sizes
- 0.9B – 8B
- Hardware
- from: Laptop
VideoGGUF2024
Genmo · USA
An open 10B video model with realistic motion. At release it was among the strongest open models; no updates now.
- Video from a description
- Short ad scenes
- Sizes
- 10B
- Hardware
- from: 1 GPU
Faces2024
ByteDance · China
Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.
- Portraits and avatars from one photo
- Ad characters with a recognizable face
- Photoshoot prototypes
- Sizes
- adapters for SDXL and FLUX.1-dev
- Hardware
- from: 1 GPU
Video2022–2024
hzwer (Zhewei Huang) and co-authors · China
Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.
- Increasing video frame rate
- Smooth slow-motion video
- Smoothing clips from AI generators
- Sizes
- lightweight model (size not stated on the model card)
- Hardware
- from: Laptop
Visual document search2024
TIGER-Lab · Canada
Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.
- Search across a mixed archive of texts and images
- Search across document pages as images
- Finding similar cards and illustrations
- Sizes
- about 4B (based on Phi-3.5-V)
- Hardware
- from: 1 GPU
Photo editing2024
Jasper AI · USA
A FLUX add-on for upscaling small and blurry images with detail reconstruction. Popular, but under the non-commercial FLUX dev license.
- Upscaling small images with detail reconstruction
- Enhancing generated images
- Upscaling pilots for a catalog
- Sizes
- add-on for FLUX.1 dev 12B
- Hardware
- from: 1 GPU
Image + textNot maintained2023–2024
Hugging Face · France / USA
Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.
- Answering questions about images
- Analysing documents and screenshots
- A base for fine-tuning
- Sizes
- 8B – 80B
- Hardware
- from: 1 GPU
Image generationNot maintained2023–2024
Lvmin Zhang (Stanford) and the community · USA
An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.
- Image from a sketch or outline
- Keeping pose and composition
- Interior visualization from a floor plan
- Sizes
- 0.4B – 1.3B
- Hardware
- from: Laptop
VideoNot maintained2023–2024
Shanghai AI Lab and CUHK · China
A module that brings Stable Diffusion image models to life, turning them into short animations. One of the first open video technologies.
- Short animations in brand style
- Animated covers and banners
- Animated stickers
- Sizes
- motion module on top of SD 1.5 / SDXL
- Hardware
- from: Laptop
AvatarsNot maintained2024
Kuaishou (Kling) · China
Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.
- Animating portraits
- Transferring an actor's expressions to a character
- Mascot animation
- Sizes
- under 1B
- Hardware
- from: Laptop
Image + textNot maintained2024
Microsoft · USA
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Sizes
- 0.23B – 0.77B
- Hardware
- from: Laptop
Photo editingNot maintained2023–2024
Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China
All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.
- Removing and replacing objects in photos
- Extending the frame to a required format
- Inserting a product or detail from a text description
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Photo editingNot maintained2024
Lvmin Zhang (author of ControlNet) · USA
Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.
- Matching product lighting to a new background
- Studio lighting for portraits without a reshoot
- Consistent lighting style across a catalog
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Deepfake detectionNot maintained2024
UC Santa Barbara and co-authors · USA
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
- Checking long articles and reports as a whole
- Filtering machine text in a publication flow
- Comparing detectors on your own data
- Sizes
- about 150M (Longformer-base)
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Huawei Noah's Ark Lab and partners · China
A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.
- Illustrations for articles and social media
- Backgrounds for product cards
- Quick visual drafts
- Sizes
- 0.6B
- Hardware
- from: Laptop
3DNot maintained2024
Tencent ARC · China
Builds a 3D mesh from a single image in about 10 seconds: first it draws the object from several angles, then assembles the model from them.
- 3D model of an object from a photo
- Assets for games and visualizations
- Prototypes for 3D printing
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
FacesNot maintained2023–2024
Tencent AI Lab (h94) · China
One of the first adapters that transfer a face from a photo into a generated image. The SD 1.5 versions run on low-end cards, but the weights are non-commercial.
- Portraits from a photo in different styles
- Image series with one character
- Avatar experiments
- Sizes
- adapters for SD 1.5 and SDXL
- Hardware
- from: Laptop
Text to speechNot maintained2024
MyShell and MIT · USA
Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.
- Voicing bot replies in foreign languages
- Voicing training materials
- Reading texts aloud on a server without a GPU
- Sizes
- small, runs in real time on a CPU
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Playground AI · USA
An SDXL-based model focused on aesthetics: vivid colors, contrast, portraits. Compatible with SDXL ecosystem tools.
- Aesthetic ad visuals
- Portraits and lifestyle images
- Post covers
- Sizes
- about 2.6B
- Hardware
- from: 1 GPU
Photo editingNot maintained2024
XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shanghai AI Lab and others) · China
Powerful SDXL-based restoration of badly damaged photos: it recreates details rather than just upscaling. Hardware-hungry; non-commercial license.
- Restoring old and blurry photos
- Upscaling with detail reconstruction
- Archive restoration pilots
- Sizes
- based on SDXL, plus LLaVA 13B for captions
- Hardware
- from: 1 GPU
FacesNot maintained2024
InstantX (Xiaohongshu) · China
Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.
- Portraits in different styles from one photo
- Avatar and character sketches
- Photoshoot prototypes
- Sizes
- adapter for SDXL
- Hardware
- from: 1 GPU
Photo editingNot maintained2023
Alibaba DAMO Academy · China
Colorizes black-and-white photos in natural colors. A lightweight model with a commercial-friendly license; a compact tiny version is available.
- Colorizing archival photos
- Color versions of historical photos for publications
- Family photo restoration service
- Sizes
- DDColor-T (tiny) and DDColor-L
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2023
Resemble AI · USA
A speech enhancement model: removes noise and restores lost frequencies so a muffled recording sounds studio-quality. Good for preparing a voice for voiceover.
- Restoring old and phone recordings
- Cleaning a voice before voiceover and cloning
- Improving audio in videos and podcasts
- Sizes
- under 1B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023
Hello-SimpleAI · China
One of the first open AI-text classifiers, trained on the HC3 corpus of paired human and ChatGPT answers. It errs in both directions: its output is a reason to talk to the author, not proof.
- First-pass check of student work
- Filtering templated applications and reviews
- Flagging suspicious texts for manual review
- Sizes
- about 125M (RoBERTa-base)
- Hardware
- from: Laptop
Text to speechNot maintained2023
Columbia University · USA
A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.
- Voicing texts in English
- A base for fine-tuning your own voice
- Voice service prototypes
- Sizes
- about 150M
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2023
Google · USA
Google's translator for more than 400 languages under a permissive license. Russian is supported. A good substitute for NLLB when commercial use is needed.
- Translating documents and emails
- Translating catalogs and product descriptions
- Translating into CIS and Asian languages
- Sizes
- 3B – 10B
- Hardware
- from: Laptop
Text to speechRUNot maintained2023
Coqui · Germany
A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.
- Voice cloning from a sample
- Multilingual voiceover
- Research and prototypes
- Sizes
- about 470M
- Hardware
- from: Laptop
Music and soundNot maintained2023
LAION · Germany
CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.
- Search sounds and music by description
- Automatic tags for an audio library
- Recognizing sound types (siren, breaking glass, voice)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
Taehoon Kim (POSTECH) · South Korea
A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.
- Batch background removal from photos
- Replacing the background with a color or blur
- Background removal in video
- Sizes
- small (based on Swin-B)
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and others) · China
Transformer-based photo upscaling, more accurate than SwinIR on fine details. Versions for real noisy photos and a lightweight HAT-S.
- Upscaling product and interior photos
- Preparing images for print
- Sharpening archival photos
- Sizes
- 9M – 40M
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2022–2023
Hendrik Schröter (University of Erlangen) · Germany
Lightweight real-time speech noise suppression that works even on a regular CPU and low-power devices. Removes hum, street and office noise while keeping the voice.
- Cleaning calls and voice messages of noise
- Preparing recordings before speech recognition
- Noise suppression for video calls
- Sizes
- about 2M
- Hardware
- from: Laptop
TextRUGGUFNot maintained2023
Sber (ai-forever) · Russia
Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.
- Base for fine-tuning on a narrow Russian-language task
- Generating template Russian texts
- Experiments with Russian-language models without license restrictions
- Sizes
- 13B
- Hardware
- from: 1 GPU
Text to speechRUGGUFNot maintained2023
Suno · USA
One of the first open models to voice text with intonation, laughter and pauses. Supports about ten languages, including Russian. Now outdated.
- Draft voiceovers for videos
- Voice service prototypes
- Sound effects in speech
- Sizes
- about 300M – 1B
- Hardware
- from: Laptop
3DNot maintained2022–2023
OpenAI · USA
Early open OpenAI models that create a 3D object from text or an image in seconds. Quality is basic, but they are fast and easy to run.
- Rough 3D mock-ups from a description
- Quick object prototypes for games and AR
- Training and research pilots in 3D
- Sizes
- 40M – 1B
- Hardware
- from: Laptop
Photo editingNot maintained2023
S-Lab, Nanyang Technological University · Singapore
One of the first Stable Diffusion-based photo upscalers: restores realistic details. Non-commercial license.
- Upscaling photos with detail reconstruction
- Research restoration pilots
- Comparison with classic upscalers
- Sizes
- based on SD 2.1
- Hardware
- from: 1 GPU
Photo editingNot maintained2022–2023
S-Lab, Nanyang Technological University · Singapore
Popular face restoration for old and blurry photos; also works on video. Has face inpainting and colorization modes. Non-commercial license.
- Restoring faces in old photos
- Enhancing faces in low-quality video
- Family archive restoration pilots
- Sizes
- small, up to 0.1B
- Hardware
- from: Laptop
Music and soundNot maintained2022
MIT · USA
A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.
- Sound event recognition
- Tagging an audio archive
- Detecting alarm sounds
- Sizes
- about 87M
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021–2022
Tencent ARC Lab · China
The classic for upscaling photos 2–4x while cleaning noise and compression artifacts. Lightweight, runs even on a CPU. Versions for drawings and anime.
- Upscaling old and small product photos
- Cleaning images of compression artifacts
- Preparing images for print
- Sizes
- about 17M
- Hardware
- from: Laptop
Computer visionNot maintained2021–2022
OpenAI · USA
The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.
- Image search by text query
- Automatic tags for a catalog
- Finding similar images
- Sizes
- about 0.15B – 0.6B
- Hardware
- from: Laptop
Photo editingNot maintained2021–2022
Tencent ARC Lab · China
Proven face restoration for old and compressed photos, with a commercial-friendly license. Often paired with Real-ESRGAN; the most used versions are 1.3 and 1.4.
- Restoring faces in old photos
- Enhancing avatars and profile photos
- Restoration in a photo shop or online service
- Sizes
- small, up to 0.1B
- Hardware
- from: Laptop
Computer visionRUNot maintained2022
Sber AI and SberDevices (ai-forever) · Russia
A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.
- Product search by photo and by Russian description
- Sorting images into categories without labeling
- Checking that a photo matches its caption
- Sizes
- 150M – 430M
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021
Samsung AI Center Moscow (with Skoltech) · Russia
Removes unwanted objects, text and watermarks from photos with clean background fill. Lightweight and fast; still the standard for this task.
- Removing price tags, people and clutter from photos
- Cleaning interior and real estate photos
- Removing text and dates from archival photos
- Sizes
- about 51M
- Hardware
- from: Laptop
Photo editingNot maintained2021
ETH Zurich · Switzerland
A transformer model for upscaling, denoising and removing JPEG artifacts from photos. Lightweight and proven; often embedded in other systems.
- Photo upscaling
- Image denoising
- Removing compression artifacts
- Sizes
- about 12M
- Hardware
- from: Laptop
AvatarsNot maintained2020
IIIT Hyderabad · India
The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.
- Quick dubbing tests
- Comparison with newer lip-sync models
- Educational and research projects
- Sizes
- small model, 96-pixel face
- Hardware
- from: Laptop