Image generationGGUF2025–2026
Alibaba · China
Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.
- Infographics for product cards
- Photo editing by text command
- Ad creatives
- Sizes
- 7B – 20B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2025–2026
Microsoft · USA
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
- Transcribing long meetings with speaker labels
- Voicing podcasts and dialogues
- Real-time voice for assistants
- Sizes
- 0.5B – 9B
- Hardware
- from: Laptop
Music and soundGGUF2025–2026
M-A-P and HKUST · China
Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.
- Songs and jingles from lyrics
- Demo versions of tracks
- Music for videos
- Sizes
- 0.5B – 7B
- Hardware
- from: 1 GPU
Music and sound2026
HeartMuLa Team · not disclosed
An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.
- Songs and jingles from lyrics
- Music for videos
- Transcribing song lyrics
- Sizes
- 3B
- Hardware
- from: 1 GPU
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
Deepfake detection2023–2026
Adobe Research and University of Surrey · USA
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
- Marking images on the way out of your own pipeline
- Checking the provenance of a submitted image
- Linking with content provenance metadata
- Sizes
- model types Q and P with different mark capacity
- Hardware
- from: Laptop
VideoGGUF2025–2026
Sand AI · China
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
- Long videos with continuation
- Video with sound
- Animating images
- Sizes
- 4.5B – 114B-A6B
- Hardware
- from: 1 GPU
Video2025–2026
NVIDIA · USA
NVIDIA's lightweight, fast video model. Produces 720p clips on a single GPU; a 4-step version enables quick generation.
- Quick clips for social media
- Bulk video generation
- Video from an image
- Sizes
- 2B – 5B
- Hardware
- from: 1 GPU
Text to speechGGUF2025–2026
bilibili · China
Speech synthesis with voice cloning and precise duration control, handy for video dubbing. Controls emotion separately from timbre.
- Video dubbing matched to timing
- Voice cloning
- Emotional voiceover
- Sizes
- about 1B – 2B
- Hardware
- from: Laptop
Deepfake detection2025–2026
University of Michigan · USA
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
- Checking submitted photos and illustrations
- Filtering AI images in a content flow
- Flagging suspicious images for manual review
- Sizes
- 22M
- Hardware
- from: Laptop
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
Video2025–2026
Alibaba · China
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
- Short promo videos
- Animating product photos
- Videos for social media
- Sizes
- 1,3B – 14B
- Hardware
- from: 1 GPU
Image generationRU2022–2026
Sber (Kandinsky Lab) · Russia
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
- Images from Russian-language descriptions
- Short promo videos from text or a photo
- Instruction-based image editing
- Sizes
- 2B – 19B
- Hardware
- from: 1 GPU
Video2024–2026
Lightricks · Israel
A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.
- Ad videos with sound
- Video from a product photo
- Voiced scenes for social media
- Sizes
- 2B – 22B
- Hardware
- from: 1 GPU
Video2026
MiniMax · China
Open weights of MiniMax's Hailuo video model. A large 33B model that makes video from text and images, but needs several server GPUs.
- Cinematic ad videos
- Video from text and images
- Complex scenes with motion
- Sizes
- 33B + 32B encoder
- Hardware
- from: Cluster
Image generationGGUF2026
Krea · USA
A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.
- Realistic photos for advertising
- High-resolution images
- Style fine-tuning
- Sizes
- 12B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
Boson AI · USA
Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.
- Expressive video voiceover
- Voicing dialogues
- Voice cloning
- Sizes
- about 3B – 8B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
Zyphra · USA
Speech synthesis with voice cloning and fine control over emotion, speed and pitch.
- Voice cloning
- Emotional voiceover
- Voicing videos
- Sizes
- about 1.6B
- Hardware
- from: Laptop
Music and soundRUGGUF2025–2026
ACE Studio and StepFun · China
Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.
- Songs and jingles for ads
- Background music for videos
- Demo versions of tracks
- Sizes
- about 2B – 4B
- Hardware
- from: Laptop
VideoGGUF2025–2026
Zhipu AI (Z.ai) and Tsinghua University · China
Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.
- Transferring an actor's motion to a character
- Replacing a character in a finished video
- Animating mascots and illustrations
- Sizes
- 14B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
HiDream.ai · China
Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.
- Image generation from descriptions
- Editing images with words
- Variations of product photos
- Sizes
- about 9B – 17B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Meituan · China
Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.
- News or course presenter videos
- Promo videos with a talking character
- Singing and voice-over
- Sizes
- based on LongCat-Video 13.6B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
Resemble AI · USA
Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- about 350M – 500M
- Hardware
- from: Laptop
Music and soundGGUF2024–2026
Stability AI · UK
Generates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.
- Sound effects for videos and games
- Background music and jingles
- Interface sounds
- Sizes
- about 0.5B – 2.3B
- Hardware
- from: Laptop
Photo editing2023–2026
BRIA AI · Israel
BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.
- Cutting products out onto a white background
- Staff and expert photos without background
- Background removal in video
- Sizes
- 44M – 220M
- Hardware
- from: Laptop
Image + textGGUF2025–2026
Kuaishou · China
Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.
- Analysing and describing short videos
- Reviewing clips and content
- Finding the right moment in a video
- Sizes
- 8B – 671B-A37B
- Hardware
- from: Laptop
Photo editing2025–2026
S-Lab, Nanyang Technological University · Singapore
Cuts a person out of video with a precise alpha mask, including hair and edges, without a green screen. Needs a first-frame mask, for example from SAM.
- Background replacement in video without chroma key
- Cutting out a person for editing and effects
- Preparing videos for advertising and social media
- Sizes
- about 35M
- Hardware
- from: Laptop
Music and sound2025–2026
Alibaba Tongyi (FunAudioLLM) · China
Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
- Audio for video based on the scene
- Editing individual sounds in a track
- Sound effects from a description
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Avatars2026
SII-GAIR and Sand.ai · China
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
- Presenter video from a script
- Ad videos with a talking character
- Training videos with a narrator
- Sizes
- 15B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
Meituan · China
Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.
- Image generation from descriptions
- Instruction-based photo editing
- Visuals for product cards
- Sizes
- 6B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2026
Black Forest Labs · Germany
Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.
- Images for product cards
- Banners and covers
- Photo editing by description (Kontext)
- Sizes
- 4B – 32B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2026
Tencent · China
Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.
- Complex scenes from long descriptions
- Images with Chinese and English text
- Instruction-based image editing
- Sizes
- 1.5B – 80B-A13B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
Alibaba (Tongyi-MAI) · China
A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.
- Photorealistic ad images
- Images with English and Chinese text
- Bulk visual generation
- Sizes
- 6B
- Hardware
- from: 1 GPU
Image generation2026
Zhipu AI (Z.ai) · China
A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.
- Posters and banners with text
- Infographics
- Illustrations for presentations
- Sizes
- 9B + 7B
- Hardware
- from: 1 GPU
VideoGGUF2025–2026
Skywork AI (Kunlun Tech) · China
Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.
- Long videos with continuation
- Video with one character from a reference
- Talking avatar from a voice
- Sizes
- 1.3B – 19B
- Hardware
- from: 1 GPU
Avatars2024–2026
Ant Group · China
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
- Presenter video from a photo and voice
- Avatar with gestures for presentations
- Voiced characters
- Sizes
- up to 1.3B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
VideoGGUF2026
OpenMOSS / MOSI · China
Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.
- Short clips with speech and sound from a description
- Ad scenes with dialogue
- Video prototypes for storyboards
- Sizes
- 32B-A18B
- Hardware
- from: 1 GPU
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
TextRUGGUF2024–2025
Vikhr Models · Russia
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
- Russian-language assistant on your own PC or server
- Knowledge-base answers (RAG)
- Texts and emails in Russian
- Sizes
- 0.5B – 24B
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
- Marking video created or processed by AI
- Finding your own mark in re-uploaded clips
- Protecting ad materials from being reused as someone else's
- Sizes
- a mark of 96 to 1024 bits
- Hardware
- from: Laptop
VideoGGUF2024–2025
Tencent · China
Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.
- Video from a text script
- Animating images
- Base for fine-tuning your own video models
- Sizes
- 8.3B – 13B
- Hardware
- from: 1 GPU
Text to speech2025
Nari Labs · South Korea
A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.
- Voicing dialogues and podcasts
- Ads with natural speech
- Training role-plays
- Sizes
- 1B – 2B
- Hardware
- from: Laptop
Video2025
Character.AI · USA
Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
- Short clips with talking characters
- Animating an image with voice-over
- Ad scene prototypes
- Sizes
- 11B
- Hardware
- from: 1 GPU
VideoGGUF2025
Meituan · China
A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.
- Long videos
- Video from a photo
- Continuing an existing video
- Sizes
- 13.6B
- Hardware
- from: 1 GPU
Music and sound2025
ASLP-lab (Northwestern Polytechnical University) · China
Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.
- Songs and jingles from lyrics
- Music for videos
- Demo versions of tracks
- Sizes
- about 1.1B
- Hardware
- from: Laptop
TextRU2025
Avito Tech · Russia
Avito's model based on Qwen3-8B, retrained for Russian: its own tokenizer makes Russian text 15–25% faster. Supports function calling.
- Product and listing descriptions in Russian
- Chatbot that calls internal services
- Request analysis and classification
- Sizes
- 7.9B
- Hardware
- from: Laptop
Image + textRU2025
Avito Tech · Russia
Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.
- Product descriptions from photos in Russian
- Checking that a photo matches its description
- Reading brands and text in images
- Sizes
- 7.4B
- Hardware
- from: 1 GPU
Music and sound2025
Tencent Hunyuan · China
Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.
- Foley and sound effects for video
- Sound for AI-generated ads
- Sound design for short videos
- Sizes
- not stated on the model card (weights about 10 GB; XL version with memory offloading)
- Hardware
- from: 1 GPU
Avatars2025
Meituan · China
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
- Video dubbing with matched facial expressions
- Dialogue between two characters from audio
- Long videos with a presenter
- Sizes
- 14B
- Hardware
- from: 1 GPU
Photo editingGGUF2024–2025
Nankai University · China
An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.
- Bulk background removal from product photos
- Precise masks for design and print
- Cutting out people with hair for advertising
- Sizes
- about 220M (lightweight lite versions available)
- Hardware
- from: Laptop
Voice: speakers and soundRU2023–2025
Community (xbgoose and others), Dusha dataset from SberDevices · Russia
Models that detect emotion from voice in Russian speech: neutral, anger, positive, sadness. Trained on the open Dusha dataset from SberDevices.
- Finding calls with irritated customers
- Assessing the tone of operator conversations
- Prioritizing complaints in a call center
- Sizes
- 21M – 316M
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
- Marking generated and edited images
- Finding a marked fragment inside a collage
- Tracking which parts of a picture were made by AI
- Sizes
- a mark encoder and decoder for images
- Hardware
- from: Laptop
Image generationGGUF2024–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.
- Placing a product or person into a new scene
- Instruction-based photo editing
- Generation from multiple references
- Sizes
- about 4B
- Hardware
- from: 1 GPU
TranslationRUGGUF2024–2025
Unbabel · Portugal
Language models tailored for translation and multilingual text work: translating, editing, and assessing translation quality. Russian is supported. Non-commercial license.
- Translation that respects context and terminology
- Post-editing machine translation
- Assessing the quality of a finished translation
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Avatars2025
ByteDance · China
Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.
- Dubbing videos into another language with lip sync
- Editing lines in finished video without reshooting
- Talking avatars for training courses
- Sizes
- requires 8–18 GB of VRAM
- Hardware
- from: Laptop
Deepfake detection2024–2025
Xiaohongshu, USTC and Shanghai Jiao Tong University · China
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
- Checking realistic AI images without obvious artifacts
- Comparing detectors on hard examples
- Fine-tuning for your own type of content
- Sizes
- several experts based on ConvNeXt and CLIP
- Hardware
- from: 1 GPU
AvatarsGGUF2025
Tencent · China
Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.
- Presenter video from a photo and audio
- Scenes with several speakers
- Cartoon characters
- Sizes
- about 13B
- Hardware
- from: 1 GPU
TextRUGGUF2023–2025
Ilya Gusev (IlyaGusev) · Russia
The best-known Russian community fine-tune: open models (Llama, Mistral, Gemma, YandexGPT) trained to act as a Russian-speaking assistant. A convenient starting point for a Russian chatbot on your own server.
- Russian-language chat assistant
- Answers based on the company knowledge base
- Drafts of emails, descriptions and posts in Russian
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Image generationGGUF2024–2025
NVIDIA · USA
NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.
- Bulk image generation
- High-resolution visuals
- Real-time generation inside apps
- Sizes
- 0.6B – 4.8B
- Hardware
- from: Laptop
Video2024–2025
HPC-AI Tech · Singapore
A fully open video generation project: weights, code and training recipe. Version 2.0 at 11B makes video from text and from an image.
- Video from a text description
- Animating images
- Training your own video model
- Sizes
- up to 11B
- Hardware
- from: 1 GPU
Faces2025
ByteDance · China
FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.
- Portraits from one photo with a precise scene description
- Testing characters for advertising
- Comparing face-preservation methods
- Sizes
- adapter for FLUX.1-dev
- Hardware
- from: 1 GPU
Virtual try-on2025
Beijing University of Posts and Telecommunications and others · China
An all-round FLUX-based apparel toolkit: try-on, generating a model wearing a given item, and "taking off" an item from a person into a separate product photo.
- Try-on from a product photo
- Photo of a model wearing an item from a text description
- Clean product photo extracted from a shot of a person
- Sizes
- add-ons (LoRA) for FLUX.1 dev 12B
- Hardware
- from: 1 GPU
Photo editing2024–2025
Prama LLC · USA
A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.
- Cutting out products and people from photos
- Background removal in video
- Preparing photos for a catalog
- Sizes
- about 95M
- Hardware
- from: Laptop
Image + textGGUF2024–2025
DeepSeek · China
A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.
- Answering questions about images
- Draft illustrations from a description
- Experiments with a unified vision and generation model
- Sizes
- 1B – 7B
- Hardware
- from: Laptop
Text analysisRUOllama2024–2025
Jina AI · Germany
Small models that turn raw web page HTML into clean Markdown or JSON. Handy for preparing websites for a knowledge base. Non-commercial license only.
- Cleaning website pages for a knowledge base
- Extracting data from pages into JSON
- Preparing texts for RAG
- Sizes
- 0.5B – 1.5B
- Hardware
- from: Laptop
Fact-checking and judges2025
Atla · UK
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring chatbot answers against your own criteria
- Comparing two versions of a prompt or model
- Filtering out weak answers before they reach a person
- Sizes
- 8B
- Hardware
- from: 1 GPU
Deepfake detection2024
Meta · USA
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
- Marking speech synthesized by your service
- Finding your own mark in third-party publications
- Checking whether synthesis was mixed into a call recording
- Sizes
- a watermark generator and detector, 16-bit message
- Hardware
- from: Laptop
Video2024
Zhipu AI (Z.ai) and Tsinghua University · China
A 2–5B video model that runs on a single gaming GPU. A popular base for research and add-ons.
- Short clips from text
- Animating images
- Video fine-tuning experiments
- Sizes
- 2B – 5B
- Hardware
- from: 1 GPU
Text to speechGGUF2024
Hugging Face · USA
Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.
- Choosing a voice by description
- Voicing videos
- Voice service prototypes
- Sizes
- 880M – 2.2B
- Hardware
- from: Laptop
TextRU2024
MTS AI (MWS AI) · Russia
A lightweight Russian-language model from MTS AI for Russian texts: answers, summaries, drafts. A ready version for CPU without a GPU is available. The larger Cotype Pro is not released openly.
- Drafts of emails and descriptions in Russian
- Short document summaries
- Answers to common customer questions
- Sizes
- 1.5B
- Hardware
- from: Laptop
Image generationGGUF2022–2024
Stability AI · UK
The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.
- Illustrations and banners for advertising
- Backgrounds and scenes for product cards
- Fine-tuning to a brand style
- Sizes
- 0.9B – 8B
- Hardware
- from: Laptop
VideoGGUF2024
Genmo · USA
An open 10B video model with realistic motion. At release it was among the strongest open models; no updates now.
- Video from a description
- Short ad scenes
- Sizes
- 10B
- Hardware
- from: 1 GPU
Faces2024
ByteDance · China
Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.
- Portraits and avatars from one photo
- Ad characters with a recognizable face
- Photoshoot prototypes
- Sizes
- adapters for SDXL and FLUX.1-dev
- Hardware
- from: 1 GPU
Video2022–2024
hzwer (Zhewei Huang) and co-authors · China
Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.
- Increasing video frame rate
- Smooth slow-motion video
- Smoothing clips from AI generators
- Sizes
- lightweight model (size not stated on the model card)
- Hardware
- from: Laptop
Finance2023–2024
AI4Finance Foundation · USA
An open set of lightweight add-ons for ordinary language models that work with financial texts and news. Not investment advice: decisions are made by a specialist.
- Assessing the tone of financial news and reports
- Tagging mentions of companies and instruments in text
- Preparing digests from a stream of business news
- Sizes
- adapters for 6B - 20B base models
- Hardware
- from: 1 GPU
Photo editing2024
Jasper AI · USA
A FLUX add-on for upscaling small and blurry images with detail reconstruction. Popular, but under the non-commercial FLUX dev license.
- Upscaling small images with detail reconstruction
- Enhancing generated images
- Upscaling pilots for a catalog
- Sizes
- add-on for FLUX.1 dev 12B
- Hardware
- from: 1 GPU
Image generationNot maintained2023–2024
Lvmin Zhang (Stanford) and the community · USA
An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.
- Image from a sketch or outline
- Keeping pose and composition
- Interior visualization from a floor plan
- Sizes
- 0.4B – 1.3B
- Hardware
- from: Laptop
VideoNot maintained2023–2024
Shanghai AI Lab and CUHK · China
A module that brings Stable Diffusion image models to life, turning them into short animations. One of the first open video technologies.
- Short animations in brand style
- Animated covers and banners
- Animated stickers
- Sizes
- motion module on top of SD 1.5 / SDXL
- Hardware
- from: Laptop
AvatarsNot maintained2024
Kuaishou (Kling) · China
Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.
- Animating portraits
- Transferring an actor's expressions to a character
- Mascot animation
- Sizes
- under 1B
- Hardware
- from: Laptop
Photo editingNot maintained2023–2024
Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China
All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.
- Removing and replacing objects in photos
- Extending the frame to a required format
- Inserting a product or detail from a text description
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Photo editingNot maintained2024
Lvmin Zhang (author of ControlNet) · USA
Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.
- Matching product lighting to a new background
- Studio lighting for portraits without a reshoot
- Consistent lighting style across a catalog
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Huawei Noah's Ark Lab and partners · China
A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.
- Illustrations for articles and social media
- Backgrounds for product cards
- Quick visual drafts
- Sizes
- 0.6B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
FacesNot maintained2023–2024
Tencent AI Lab (h94) · China
One of the first adapters that transfer a face from a photo into a generated image. The SD 1.5 versions run on low-end cards, but the weights are non-commercial.
- Portraits from a photo in different styles
- Image series with one character
- Avatar experiments
- Sizes
- adapters for SD 1.5 and SDXL
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Playground AI · USA
An SDXL-based model focused on aesthetics: vivid colors, contrast, portraits. Compatible with SDXL ecosystem tools.
- Aesthetic ad visuals
- Portraits and lifestyle images
- Post covers
- Sizes
- about 2.6B
- Hardware
- from: 1 GPU
FacesNot maintained2024
InstantX (Xiaohongshu) · China
Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.
- Portraits in different styles from one photo
- Avatar and character sketches
- Photoshoot prototypes
- Sizes
- adapter for SDXL
- Hardware
- from: 1 GPU
VideoNot maintained2023
Stability AI · UK
Stability AI's first open video model: turns a photo into a 2–4 second clip. Now behind newer models in quality.
- Animating product photos
- Short video intros
- Animating illustrations
- Sizes
- about 1.5B
- Hardware
- from: 1 GPU
Music and soundNot maintained2023
LAION · Germany
CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.
- Search sounds and music by description
- Automatic tags for an audio library
- Recognizing sound types (siren, breaking glass, voice)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
AvatarsNot maintained2023
Xi'an Jiaotong University and Tencent AI Lab · China
An older lightweight talking-head model: one photo plus audio becomes a video. Runs on weak hardware, but quality is noticeably below newer models.
- Talking photo for greetings
- Simple voiced avatars
- Sizes
- under 1B
- Hardware
- from: Laptop
TextRUGGUFNot maintained2023
Sber (ai-forever) · Russia
Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.
- Base for fine-tuning on a narrow Russian-language task
- Generating template Russian texts
- Experiments with Russian-language models without license restrictions
- Sizes
- 13B
- Hardware
- from: 1 GPU
Text to speechRUGGUFNot maintained2023
Suno · USA
One of the first open models to voice text with intonation, laughter and pauses. Supports about ten languages, including Russian. Now outdated.
- Draft voiceovers for videos
- Voice service prototypes
- Sound effects in speech
- Sizes
- about 300M – 1B
- Hardware
- from: Laptop
Text analysisRUNot maintained2022–2023
David Dale (cointegrated) and the community · Russia
Ready-made tiny rubert-tiny models for Russian text: detect rudeness and insults, sentiment and emotions. They run on a CPU in milliseconds.
- Filtering insults in Russian chats and comments
- Labeling reviews as positive, neutral or negative
- Spotting irritated customers in requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop