What runs on a single GPU

A single GPU is the most common business setup: a server with a 16, 24, 48 or 80 GB card covers almost everything except the largest models. This is where live chat, image generation and real-time document processing become practical.

146 familiesUpdated 22 Sep 2026Calculate video memory for your model size

What this means in practice

16 GB cardMid-sized language models in compressed form, document understanding, image generation. The cheapest way in.
24 GB cardThe workhorse: a bigger model, or several tasks on one card. Enough for a chatbot and background processing together.
48-80 GB cardModels at full precision, long context, video generation. More expensive, but one card replaces a small cluster.

What runs on it

Video 16

VideoGGUF2025–2026

MAGI

Sand AI · China

Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.

  • Long videos with continuation
  • Video with sound
  • Animating images
Sizes
4.5B – 114B-A6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025–2026

SANA-Video

NVIDIA · USA

NVIDIA's lightweight, fast video model. Produces 720p clips on a single GPU; a 4-step version enables quick generation.

  • Quick clips for social media
  • Bulk video generation
  • Video from an image
Sizes
2B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025–2026

Wan

Alibaba · China

Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).

  • Short promo videos
  • Animating product photos
  • Videos for social media
Sizes
1,3B – 14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2024–2026

LTX-Video / LTX-2

Lightricks · Israel

A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.

  • Ad videos with sound
  • Video from a product photo
  • Voiced scenes for social media
Sizes
2B – 22B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
VideoGGUF2025–2026

SCAIL

Zhipu AI (Z.ai) and Tsinghua University · China

Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.

  • Transferring an actor's motion to a character
  • Replacing a character in a finished video
  • Animating mascots and illustrations
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025–2026

Matrix-Game

Skywork AI (Kunlun) · China

An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.

  • Game world prototypes without an engine
  • Interactive demos and simulations
  • Generating data to train agents
Sizes
1.8B – 17B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2025–2026

SkyReels

Skywork AI (Kunlun Tech) · China

Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.

  • Long videos with continuation
  • Video with one character from a reference
  • Talking avatar from a voice
Sizes
1.3B – 19B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
VideoGGUF2026

MOVA

OpenMOSS / MOSI · China

Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.

  • Short clips with speech and sound from a description
  • Ad scenes with dialogue
  • Video prototypes for storyboards
Sizes
32B-A18B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2024–2025

HunyuanVideo

Tencent · China

Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.

  • Video from a text script
  • Animating images
  • Base for fine-tuning your own video models
Sizes
8.3B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Video2025

Ovi

Character.AI · USA

Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.

  • Short clips with talking characters
  • Animating an image with voice-over
  • Ad scene prototypes
Sizes
11B
Hardware
from: 1 GPU
Commercial use allowedDetails
VideoGGUF2025

LongCat-Video

Meituan · China

A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.

  • Long videos
  • Video from a photo
  • Continuing an existing video
Sizes
13.6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Video2025

Hunyuan-GameCraft

Tencent Hunyuan · China

Turns a single image into a controllable game-scene video: the camera moves on keyboard commands. Minimum 24 GB of GPU memory, 80 GB recommended.

  • Interactive video prototypes of game locations
  • Camera walkthrough videos of a scene
  • Level demos for pitches
Sizes
based on HunyuanVideo
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Image generation 15

Image generationGGUF2025–2026

Qwen-Image

Alibaba · China

Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.

  • Infographics for product cards
  • Photo editing by text command
  • Ad creatives
Sizes
7B – 20B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationRU2022–2026

Kandinsky

Sber (Kandinsky Lab) · Russia

Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.

  • Images from Russian-language descriptions
  • Short promo videos from text or a photo
  • Instruction-based image editing
Sizes
2B – 19B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025–2026

Chroma

lodestones (independent developer) · not disclosed

A community model retrained from FLUX.1-schnell with a simplified architecture. No style censorship; popular as a base for fine-tuning.

  • Base for fine-tuning your own styles
  • Artistic illustrations
  • Images for games
Sizes
4B – 8.9B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2026

Krea 2

Krea · USA

A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.

  • Realistic photos for advertising
  • High-resolution images
  • Style fine-tuning
Sizes
12B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2026

Ideogram 4

Ideogram · Canada

Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.

  • Testing text-in-image generation
  • Research and prototypes
Sizes
about 9B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Image generationGGUF2025–2026

HiDream

HiDream.ai · China

Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.

  • Image generation from descriptions
  • Editing images with words
  • Variations of product photos
Sizes
about 9B – 17B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025–2026

LongCat-Image

Meituan · China

Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.

  • Image generation from descriptions
  • Instruction-based photo editing
  • Visuals for product cards
Sizes
6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2024–2026

FLUX

Black Forest Labs · Germany

Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.

  • Images for product cards
  • Banners and covers
  • Photo editing by description (Kontext)
Sizes
4B – 32B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2024–2026

Hunyuan Image

Tencent · China

Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.

  • Complex scenes from long descriptions
  • Images with Chinese and English text
  • Instruction-based image editing
Sizes
1.5B – 80B-A13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2025–2026

Z-Image

Alibaba (Tongyi-MAI) · China

A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.

  • Photorealistic ad images
  • Images with English and Chinese text
  • Bulk visual generation
Sizes
6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generation2026

GLM-Image

Zhipu AI (Z.ai) · China

A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.

  • Posters and banners with text
  • Infographics
  • Illustrations for presentations
Sizes
9B + 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2024–2025

OmniGen

BAAI (Beijing Academy of Artificial Intelligence) · China

An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.

  • Placing a product or person into a new scene
  • Instruction-based photo editing
  • Generation from multiple references
Sizes
about 4B
Hardware
from: 1 GPU
Commercial use allowedDetails

Text 10

TextGGUF2025–2026

LLaDA

Renmin University of China (GSAI) and Ant Group (inclusionAI) · China

Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.

  • Fast generation of code and text
  • Agent scenarios with long context
  • Research into alternatives to standard LLMs
Sizes
8B – 100B (MoE)
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2026

Muse Glimmer

Meta Superintelligence Labs · USA

An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.

  • Agents with tool calling
  • Analysis of screenshots, charts and documents
  • Multilingual assistant
Sizes
30B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextGGUF2025–2026

Kimi

Moonshot AI · China

Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.

  • Multi-step agents: search, data collection, reports
  • In-depth document analysis
  • Help for developers
Sizes
16B-A3B – 2.8T-A104B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextOllama2023–2026

Solar

Upstage · South Korea

Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.

  • Working with office documents
  • Agents for routine tasks
  • Help for developers
Sizes
10.7B – 250B-A15B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextGGUF2025–2026

Step

StepFun · China

StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.

  • High-load agents
  • Analysis of documents with diagrams and screenshots
  • Help for developers
Sizes
8B – 321B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextGGUF2024–2026

INTELLECT

Prime Intellect · USA

Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.

  • Reasoning and math tasks
  • Programming help
  • Agents with tool calling
Sizes
10B – 106B-A12B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2025

gpt-oss

OpenAI · USA

OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.

  • AI agent that calls internal systems
  • Answers based on internal policies
  • Drafts of emails and reports
Sizes
20B, 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextGGUF2025

Seed-OSS

ByteDance · China

An open ByteDance 36B model with up to 512K tokens of context and an adjustable thinking budget. Fits on a single powerful GPU.

  • Analysis of long documents
  • Agents with tools
  • Corporate assistant
Sizes
36B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2024

Athene

Nexusflow · USA

Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.

  • Research on agents and function calling
  • Comparison with commercial models
  • Experiments with a chat assistant
Sizes
70B – 72B
Hardware
from: 1 GPU
Non-commercial onlyDetails
TextRUGGUFNot maintained2023

ruGPT-3.5

Sber (ai-forever) · Russia

Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.

  • Base for fine-tuning on a narrow Russian-language task
  • Generating template Russian texts
  • Experiments with Russian-language models without license restrictions
Sizes
13B
Hardware
from: 1 GPU
Commercial use allowedDetails

3D 10

3D2025–2026

VGGT

Meta and the University of Oxford (VGG) · USA / UK

Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.

  • 3D model of a room or object from a photo series
  • Camera pose estimation for photogrammetry
  • Point cloud for measurements and comparison with the plan
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3D2025–2026

HY-World (HunyuanWorld)

Tencent · China

Generates whole 3D worlds and scenes from text or an image that you can walk through. The second version builds a scene from video and photos.

  • 3D scenes for games and virtual tours
  • Backgrounds and environments for video production
  • Draft locations for simulations
Sizes
set of several models
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3DGGUF2024–2025

TRELLIS

Microsoft · USA

One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.

  • 3D models of products and interiors from photos
  • Assets for games and AR/VR
  • Prototypes for 3D printing
Sizes
up to 4B (TRELLIS.2)
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2025

MapAnything

Meta and Carnegie Mellon University · USA

A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.

  • 3D reconstruction of an object or room from photos
  • Exporting the scene to COLMAP format for further processing
  • Depth and camera pose estimation
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2025

Pi3 (π³)

Shanghai AI Lab · China

Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.

  • 3D scene reconstruction from video
  • Camera pose estimation from frames
  • Point clouds for research and prototypes
Sizes
0.96B – 1.4B
Hardware
from: 1 GPU
Non-commercial onlyDetails
3D2025

HY-Motion

Tencent Hunyuan · China

Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.

  • Character animation from a text description
  • Draft animation for games and videos
  • Motion library for avatars
Sizes
0.46B – 1B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3DGGUF2025

SAM 3D

Meta · USA

Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.

  • 3D model of an item from a catalog photo
  • Estimating body pose and shape from a photo
  • Try-on and AR scenarios
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3D2024–2025

Hunyuan3D

Tencent · China

Tencent's open 3D line: shape and texture from an image, at the level of paid services. Omni adds control of pose and shape, Part splits a model into parts.

  • Textured 3D product models
  • Characters and objects for games
  • Splitting a model into parts for printing
Sizes
set of models: shape and textures
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3D2024–2025

Stable Fast 3D / SPAR3D

Stability AI · UK

Stability AI models that turn a single photo into a textured 3D model in about a second. SPAR3D lets you adjust the shape through a point cloud.

  • 3D product cards from photos
  • Assets for games and AR
  • Quick mock-ups for design
Sizes
1B – 2B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
3DNot maintained2024

InstantMesh

Tencent ARC · China

Builds a 3D mesh from a single image in about 10 seconds: first it draws the object from several angles, then assembles the model from them.

  • 3D model of an object from a photo
  • Assets for games and visualizations
  • Prototypes for 3D printing
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use allowedDetails

Virtual try-on 9

Virtual try-on2026

FASHN VTON

FASHN AI · Israel

A rare open try-on model with a commercial license: mask-free, accepts a photo of the item on a model or a flat lay. Weights are about 2 GB.

  • Product cards on a model without a photo shoot
  • Fitting room on a store website
  • Catalog from flat-lay clothing photos
Sizes
972M
Hardware
from: 1 GPU
Commercial use allowedDetails
Virtual try-on2025

OmniTry

Kunbyte AI · China

Try-on beyond clothing: glasses, earrings, bags, hats, watches and other accessories. Works without a mask. Needs a GPU with 28 GB or more.

  • Trying accessories and jewelry on a photo
  • Product cards with an accessory on a model
  • Online fitting room for eyewear and jewelry
Sizes
add-on for FLUX.1 Fill dev 12B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Virtual try-on2025

FastFit

LavieAI and Sun Yat-sen University · China

Tries on several items at once: top, bottom, shoes, bag. Faster than earlier models thanks to caching. Non-commercial license.

  • Building an outfit from several products on one model
  • "Build a look" pilot on a website
  • Lookbook prototypes without a shoot
Sizes
based on SD 1.5 inpainting
Hardware
from: 1 GPU
Non-commercial onlyDetails
Virtual try-on2025

Any2AnyTryon

Beijing University of Posts and Telecommunications and others · China

An all-round FLUX-based apparel toolkit: try-on, generating a model wearing a given item, and "taking off" an item from a person into a separate product photo.

  • Try-on from a product photo
  • Photo of a model wearing an item from a text description
  • Clean product photo extracted from a shot of a person
Sizes
add-ons (LoRA) for FLUX.1 dev 12B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Virtual try-on2024

Leffa

Meta AI (with King's College London) · USA

Meta's model for virtual try-on and changing a person's pose in a photo. Carefully transfers fine fabric details and lettering. MIT license, but the training data is non-commercial.

  • Trying clothes on a model photo
  • Changing the model pose in an existing shot
  • Adding extra angles for a product card
Sizes
based on Stable Diffusion
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Virtual try-on2024

FitDiT

Tencent and Fudan University · China

Tencent's transformer-based try-on: more accurately reproduces fabric texture, fine prints and garment length. Non-commercial license.

  • Try-on of items with complex prints and textures
  • Test product cards on a model
  • Checking length and fit on a photo
Sizes
based on SD3
Hardware
from: 1 GPU
Non-commercial onlyDetails
Virtual try-onNot maintained2024

IDM-VTON

KAIST and OMNIOUS.AI · South Korea

One of the best-known open virtual try-on models: moves a garment from a product photo onto a photo of a person, keeping prints and logos well. Non-commercial license.

  • Pilot of a fitting room on a store website
  • Prototype product cards on a model without a photo shoot
  • Comparison with commercial try-on services
Sizes
based on SDXL
Hardware
from: 1 GPU
Non-commercial onlyDetails
Virtual try-onNot maintained2024

OOTDiffusion

Xiao-i Research · China

An early popular open try-on model: one version for upper-body garments, another for full-length outfits. Non-commercial license.

  • Trying tops on a model photo
  • Full-length try-on: tops, bottoms, dresses
  • Fitting room prototype for testing
Sizes
based on Stable Diffusion
Hardware
from: 1 GPU
Non-commercial onlyDetails
Virtual try-onNot maintained2024

StableVITON

KAIST · South Korea

A research try-on model from CVPR 2024, one of the first built on Stable Diffusion. Now mostly used as a comparison baseline.

  • Pilot of upper-body garment try-on
  • Comparing quality of different try-on models
  • Training your own try-on on the open code
Sizes
based on Stable Diffusion
Hardware
from: 1 GPU
Non-commercial onlyDetails

Robotics 8

Robotics2025–2026

GigaBrain

GigaAI · China

A robot control model trained mostly on synthetic data from a world model. It reduces spending on collecting data from real robots.

  • Controlling a robot arm
  • Fine-tuning on a small amount of your own data
  • Sorting and assembly pilots
Sizes
3.5B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2025–2026

NVIDIA Cosmos

NVIDIA · USA

"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.

  • Synthetic video for training robots and self-driving vehicles
  • Testing scenarios in simulation
  • Robot control (Policy versions)
Sizes
2B – 65B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2026

LingBot-VLA

Ant Group (Robbyant) · China

A robot control model from Ant Group trained on a large volume of data from real robots. Version 2.0 works with different types of robot arms.

  • Controlling a two-armed robot
  • Fine-tuning for your own operation
  • Assembly and sorting pilots
Sizes
4B – 6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2026

Xiaomi Robotics

Xiaomi · China

Open robot control models from Xiaomi. Robotics-1 is designed for household and kitchen tasks, U0 combines scene understanding and action.

  • Controlling a robot arm by command
  • Household and service scenarios
  • Base for fine-tuning
Sizes
4B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2025–2026

MolmoAct

Ai2 (Allen Institute for AI) · USA

A fully open robot control model that first "reasons" about space and trajectory, then acts. Its reasoning can be checked.

  • Controlling a robot arm with explainable steps
  • Fine-tuning for your own robot
  • Research pilots
Sizes
5B – 8B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2025–2026

NVIDIA Isaac GR00T

NVIDIA · USA

NVIDIA's foundation model for humanoid robots and robot arms: it sees, understands a command and outputs movements. Built into the Isaac ecosystem.

  • Controlling humanoid robots
  • Fine-tuning for your own robotic cell
  • Training in simulation with transfer to a real robot
Sizes
2B – 3B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Robotics2025

π0 / π0.5 (openpi)

Physical Intelligence · USA

Robot control models from Physical Intelligence: folding laundry, tidying up, handling objects. π0.5 copes better in unfamiliar settings.

  • Controlling robot arms and two-armed robots
  • Fine-tuning for your own operations
  • Pilots for automating manual work
Sizes
about 3B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Robotics2024–2025

OpenVLA

Stanford, Berkeley and partners · USA

The first large open vision-language-action model: a robot arm carries out commands like "put the apple in the bowl". OFT makes it several times faster.

  • Controlling a robot arm by text command
  • Pilots for robotizing simple operations
  • Base for fine-tuning to your own robot
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails

Avatars 7

AvatarsGGUF2025–2026

LongCat-Video-Avatar

Meituan · China

Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.

  • News or course presenter videos
  • Promo videos with a talking character
  • Singing and voice-over
Sizes
based on LongCat-Video 13.6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2024–2026

Hallo

Fudan University · China

A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.

  • Presenter video from a photo and audio
  • Long training videos
  • Live avatar
Sizes
about 1B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2026

daVinci-MagiHuman

SII-GAIR and Sand.ai · China

Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.

  • Presenter video from a script
  • Ad videos with a talking character
  • Training videos with a narrator
Sizes
15B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2024–2026

EchoMimic

Ant Group · China

Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.

  • Presenter video from a photo and voice
  • Avatar with gestures for presentations
  • Voiced characters
Sizes
up to 1.3B
Hardware
from: 1 GPU
Commercial use allowedDetails
AvatarsGGUF2025–2026

Live Avatar

Alibaba (Quark) · China

A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.

  • Live avatar for customer dialogue
  • Endless broadcasts with a presenter
  • Interactive characters
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
Avatars2025

MultiTalk / InfiniteTalk

Meituan · China

Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.

  • Video dubbing with matched facial expressions
  • Dialogue between two characters from audio
  • Long videos with a presenter
Sizes
14B
Hardware
from: 1 GPU
Commercial use allowedDetails
AvatarsGGUF2025

HunyuanVideo-Avatar

Tencent · China

Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.

  • Presenter video from a photo and audio
  • Scenes with several speakers
  • Cartoon characters
Sizes
about 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Visual document search 7

Visual document searchGGUF2026

EVIE

Tencent · China

Tencent models based on Qwen3.5 for searching scans and PDFs as images. According to the model card, among the top of the ViDoRe leaderboard at release.

  • Search across scans and PDFs without OCR
  • RAG over reports with tables and charts
  • Search across document archives
Sizes
4.5B – 8B
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2025

Nomic Embed Multimodal / ColNomic

Nomic AI · USA

Search across PDF pages and scans as images. The cards list English, Italian, French, German and Spanish — Russian is not among them.

  • Search across an archive of scans and PDFs
  • Search across tables and diagrams inside documents
  • Picking pages for an AI assistant answer
Sizes
3B and 7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Visual document search2025

LlamaIndex vdr

LlamaIndex · USA

A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.

  • Search across scans and PDFs without OCR
  • Search across invoices, acts and contracts
  • Picking pages for an AI assistant answer
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2024

GME (General Multimodal Embedding)

Alibaba (Tongyi Lab) · China

One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.

  • Finding a product by photo
  • Search across a catalogue of images and cards
  • Search across document pages as images
Sizes
2B and 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2024

VLM2Vec

TIGER-Lab · Canada

Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.

  • Search across a mixed archive of texts and images
  • Search across document pages as images
  • Finding similar cards and illustrations
Sizes
about 4B (based on Phi-3.5-V)
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2024

MonoQwen2-VL (LightOn)

LightOn · France

A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.

  • Refining search results over scans and PDFs
  • Selecting pages before an AI assistant answers
  • Sorting retrieved slides and reports
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2024

DSE (Document Screenshot Embedding)

University of Waterloo, Tevatron project · Canada

Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.

  • Search across scans and PDFs without OCR
  • Search across presentations and reports with complex layouts
  • Picking pages for an AI assistant answer
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails

Image + text 5

Image + text2026

MOSS-VL

OpenMOSS (Fudan University) · China

An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.

  • Analyzing long videos and finding events by time
  • Real-time streaming video analysis
  • Understanding photos and documents
Sizes
about 11B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + text2026

Qwen3-VL Resume Parser

Sukhrob Nurali · not disclosed

A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.

  • Moving a resume from PDF into a candidate record
  • Filling a candidate database without manual typing
  • Parsing resumes with different layouts and styling
Sizes
8B, a fine-tune of Qwen3-VL-8B-Instruct
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textRU2025

A-Vision (Авито)

Avito Tech · Russia

Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.

  • Product descriptions from photos in Russian
  • Checking that a photo matches its description
  • Reading brands and text in images
Sizes
7.4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textGGUF2025

Kimi-VL

Moonshot AI · China

An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.

  • Analysing long PDFs and presentations
  • Answering questions about video
  • Operating interfaces from screenshots
Sizes
16B-A3B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textNot maintained2023–2024

IDEFICS

Hugging Face · France / USA

Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.

  • Answering questions about images
  • Analysing documents and screenshots
  • A base for fine-tuning
Sizes
8B – 80B
Hardware
from: 1 GPU
Commercial use allowedDetails

Weather and climate 5

Weather and climate2024–2026

Ai2 ACE2 (климатический эмулятор)

Allen Institute for AI (Ai2) · USA

A fast climate model emulator: simulates the atmosphere years and decades ahead on a single GPU. Coupled with an ocean model (SamudrACE) for long-term scenarios.

  • Decades-long climate scenarios to assess long-term asset risks
  • Large-scale what-if runs on temperature and precipitation
  • Preparing data for crop yield and energy demand models
Sizes
checkpoint of about 1.8 GB
Hardware
from: 1 GPU
Commercial use allowedDetails
Weather and climate2023–2026

Google DeepMind GraphCast / GenCast / WeatherNext 2

Google DeepMind · UK

Google DeepMind's family of global weather models: GraphCast (10-day forecast), GenCast (probabilistic ensemble) and WeatherNext 2 with cyclone forecasting. Since August 2026 the weights are cleared for commercial use.

  • Medium-range weather forecasts for planning shifts, voyages and deliveries
  • Probabilistic assessment of extreme weather for insurance portfolios
  • Tropical cyclone track forecasts for marine and port operations
Sizes
from lightweight 1° versions to full 0.25°
Hardware
from: 1 GPU
Commercial use allowedDetails
Weather and climate2024–2026

Microsoft Aurora

Microsoft Research · USA

A foundation model of Earth's atmosphere: global weather forecasts, plus separate versions for air quality and ocean waves. Computes a forecast in seconds instead of hours on a supercomputer.

  • Your own forecast of temperature, wind and precipitation for company locations
  • Sea state estimates for planning voyages and port operations
  • Air pollution forecasts for industrial sites
Sizes
about 1.3B (a small test version is available)
Hardware
from: 1 GPU
Commercial use allowedDetails
Weather and climate2022–2026

NVIDIA FourCastNet и модели Earth-2

NVIDIA · USA

NVIDIA's set of weather and climate models: global FourCastNet forecasts, downscaling to kilometers (CorrDiff), regional storm forecasts (StormCast), climate generation (cBottle, Atlas). Run via Earth2Studio.

  • Global forecasts followed by downscaling to the region you need
  • Short-term forecasts of thunderstorms and heavy rain for dispatch services
  • Generating many weather scenarios for stress tests
Sizes
98M – 2.5B (FourCastNet 3 — about 711M)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Weather and climate2024–2026

ECMWF AIFS

European Centre for Medium-Range Weather Forecasts (ECMWF) · Europe (intergovernmental organization)

ECMWF's weather neural network running operationally: a 15-day forecast four times a day, an ensemble version with 51 scenarios, and since version 2, ocean waves.

  • Running your own forecast from open initial data
  • Ensemble forecasts to estimate the probability of frost, downpours and storms
  • Wave forecasts for marine operations
Sizes
checkpoint of about 1 GB
Hardware
from: 1 GPU
Commercial use allowedDetails

Code 4

CodeOllama2026

Ornith

DeepReinforce · not disclosed

Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.

  • A developer agent in the terminal
  • Fixing bugs from a task description
  • Understanding and extending a large repository
Sizes
9B – 397B
Hardware
from: 1 GPU
Commercial use allowedDetails
Code2025–2026

KAT-Coder / KAT-Dev

Kwaipilot (Kuaishou) · China

Kuaishou models for agentic development, trained to solve real tasks in repositories. KAT-Coder-V2.5-Dev (35B, 3B active) is the open version of their closed flagship.

  • An agent that fixes tasks in the repository
  • Code generation and refactoring
  • Automating routine development tasks
Sizes
32B – 72B, 35B-A3B
Hardware
from: 1 GPU
Commercial use allowedDetails
CodeGGUF2026

Laguna

Poolside · USA

Models for agentic programming: they edit code in a repository on their own. The small XS runs on a Mac with 36 GB of memory; S 2.1 has a 1M-token context.

  • Coding agent for in-house development
  • Bug fixing and code improvements
  • Working with large codebases
Sizes
33B-A3B – 225B-A23B
Hardware
from: 1 GPU
Commercial use allowedDetails
CodeRUOllama2025

Devstral

Mistral AI (with All Hands AI) · France

Mistral models for agentic development: they read the repository, edit files and run commands on their own. The 24B version fits on a single GPU.

  • A developer agent that fixes tickets from the tracker
  • Extending internal systems from a description
  • Automating routine code edits
Sizes
24B – 123B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Music and sound 4

Music and soundGGUF2025–2026

YuE

M-A-P and HKUST · China

Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.

  • Songs and jingles from lyrics
  • Demo versions of tracks
  • Music for videos
Sizes
0.5B – 7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Music and sound2026

HeartMuLa

HeartMuLa Team · not disclosed

An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.

  • Songs and jingles from lyrics
  • Music for videos
  • Transcribing song lyrics
Sizes
3B
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and sound2025–2026

ThinkSound / PrismAudio

Alibaba Tongyi (FunAudioLLM) · China

Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.

  • Audio for video based on the scene
  • Editing individual sounds in a track
  • Sound effects from a description
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and sound2025

HunyuanVideo-Foley

Tencent Hunyuan · China

Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.

  • Foley and sound effects for video
  • Sound for AI-generated ads
  • Sound design for short videos
Sizes
not stated on the model card (weights about 10 GB; XL version with memory offloading)
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Photo editing 4

Photo editing2025

SeedVR / SeedVR2

ByteDance Seed · China

ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.

  • Upscaling photos and video to 2K–4K
  • Restoring old videos and photos
  • Enhancing user photos before publishing
Sizes
3B – 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Photo editing2024

Flux.1-dev ControlNet Upscaler

Jasper AI · USA

A FLUX add-on for upscaling small and blurry images with detail reconstruction. Popular, but under the non-commercial FLUX dev license.

  • Upscaling small images with detail reconstruction
  • Enhancing generated images
  • Upscaling pilots for a catalog
Sizes
add-on for FLUX.1 dev 12B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Photo editingNot maintained2024

SUPIR

XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shanghai AI Lab and others) · China

Powerful SDXL-based restoration of badly damaged photos: it recreates details rather than just upscaling. Hardware-hungry; non-commercial license.

  • Restoring old and blurry photos
  • Upscaling with detail reconstruction
  • Archive restoration pilots
Sizes
based on SDXL, plus LLaVA 13B for captions
Hardware
from: 1 GPU
Non-commercial onlyDetails
Photo editingNot maintained2023

StableSR

S-Lab, Nanyang Technological University · Singapore

One of the first Stable Diffusion-based photo upscalers: restores realistic details. Non-commercial license.

  • Upscaling photos with detail reconstruction
  • Research restoration pilots
  • Comparison with classic upscalers
Sizes
based on SD 2.1
Hardware
from: 1 GPU
Non-commercial onlyDetails

Computer-use agents 4

Computer-use agents2025–2026

OpenCUA / Qwen-CUA

XLANG Lab (University of Hong Kong) · China

Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.

  • Working in desktop software without an API
  • Moving data between systems
  • Running user scenarios for tests
Sizes
7B – about 400B (MoE)
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2026

EvoCUA

Meituan · China

Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.

  • Working in office and legacy software without an API
  • Moving data between systems
  • Running test scenarios
Sizes
8B – 32B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agents2023–2025

CogAgent / AutoGLM

Zhipu AI (Z.ai) and Tsinghua University · China

One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.

  • Automating actions in mobile apps
  • Working in web interfaces without an API
  • Testing apps against scenarios
Sizes
9B – 18B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer-use agentsGGUF2025

Magma

Microsoft Research · USA

An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.

  • Pilots in interface control
  • Research projects spanning screens and robotics
  • Analyzing screenshots with an action plan
Sizes
8B
Hardware
from: 1 GPU
Commercial use allowedDetails

Biology and chemistry 4

Biology and chemistry2022–2026

OpenFold / OpenFold3

AlQuraishi Lab (Columbia University) and the OpenFold consortium · USA

A fully open reproduction of AlphaFold 2 and then AlphaFold 3 under Apache 2.0, with training data. OpenFold3 predicts complexes of proteins, nucleic acids and ligands.

  • Predicting structures of proteins and ligand complexes
  • Fine-tuning on the company's own data (training code is open)
  • An in-house structural analysis service without sending data outside
Sizes
a single set of weights per version
Hardware
from: 1 GPU
Commercial use allowedDetails
Biology and chemistry2024–2026

Arc Institute Evo / Evo 2

Arc Institute (with Together AI, Stanford, NVIDIA) · USA

DNA language models with context up to a million nucleotides: they assess the impact of mutations, annotate genomes and generate sequences. Evo 2 is trained on genomes from all domains of life.

  • Assessing the likely harmfulness of genetic variants for research
  • Annotating genomes of microorganisms and plants
  • Finding promising sequences in breeding and synthetic biology
Sizes
1B – 40B
Hardware
from: 1 GPU
Commercial use allowedDetails
Biology and chemistry2024–2025

Boltz (Boltz-1, Boltz-2, BoltzGen)

MIT (Jameel Clinic) and Recursion · USA

An open MIT-licensed alternative to AlphaFold 3: predicts structures of protein, DNA and small-molecule complexes; Boltz-2 estimates binding strength, BoltzGen designs new binding proteins.

  • Predicting how a candidate molecule binds to a target protein
  • Ranking compounds by predicted binding strength before synthesis
  • Designing binder proteins for a given target
Sizes
checkpoints of about 2 GB
Hardware
from: 1 GPU
Commercial use allowedDetails
Biology and chemistry2024

AlphaFold 3

Google DeepMind and Isomorphic Labs · UK

The reference model for the structure of biomolecules and their complexes. Weights are provided for non-commercial research only; companies need commercial access via Google Cloud or open alternatives (Boltz, OpenFold3).

  • Academic research on protein and complex structures
  • Benchmarking open alternatives against the reference on your own targets
Sizes
a single set of weights
Hardware
from: 1 GPU
Non-commercial onlyDetails

Finance 4

Finance2025

Fino1 / Fin-o1

The Fin AI · international project

Models that spell out their reasoning before answering financial questions with numbers and tables. Not investment advice: decisions are made by a specialist.

  • Calculation questions on statements with the working shown
  • Working through tasks with tables and numbers from documents
  • Checking calculations made by hand
Sizes
8B и 14B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Finance2023–2024

XuanYuan

Du Xiaoman (Duxiaoman-DI) · China

A large Chinese model family for the financial industry: advice, document reading and long texts up to 8k-16k. Not investment advice: decisions are made by a specialist.

  • Answering customer questions about banking products
  • Working through long financial documents
  • An internal assistant for a finance company regulations
Sizes
6B – 176B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Finance2023–2024

FinGPT

AI4Finance Foundation · USA

An open set of lightweight add-ons for ordinary language models that work with financial texts and news. Not investment advice: decisions are made by a specialist.

  • Assessing the tone of financial news and reports
  • Tagging mentions of companies and instruments in text
  • Preparing digests from a stream of business news
Sizes
adapters for 6B - 20B base models
Hardware
from: 1 GPU
Commercial use with conditionsDetails
FinanceNot maintained2023

DISC-FinLLM

Fudan-DISC, Fudan University · China

A financial assistant made of several fine-tuned experts: advice, calculations, document reading and knowledge-base search. Not investment advice: decisions are made by a specialist.

  • In-house advice on financial questions
  • Reading financial documents and news
  • Prompts for front-office staff
Sizes
13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Medicine 3

MedicineGGUF2025–2026

AntAngelMed

Ant Healthcare (Ant Group) and Zhejiang Provincial Medical Information Center · China

A large medical MoE model based on Ling-flash-2.0: 100B parameters with 6B active, so it answers quickly. Does not replace a doctor; decisions are made by a specialist.

  • Reference answers to staff on clinical questions
  • Draft discharge summaries for a doctor to review
  • Searching medical literature
Sizes
100B-A6B
Hardware
from: 1 GPU
Commercial use allowedDetails
MedicineNot maintained2023

Clinical Camel

Bo Wang's lab (University of Toronto, Vector Institute) · Canada

An early medical model on Llama 2 70B, trained on dialogues based on medical texts. Now mainly of research interest. Does not replace a doctor; decisions are made by a specialist.

  • Research pilots on medical dialogue
  • Training materials for staff
  • Comparison with newer medical models
Sizes
70B
Hardware
from: 1 GPU
Non-commercial onlyDetails
MedicineNot maintained2023

RadFM

Shanghai Jiao Tong University and Shanghai AI Lab · China

An early general-purpose radiology model: understands 2D and 3D images (CT, MRI) together with text. More of a research base. Does not replace a doctor; decisions are made by a specialist.

  • Research pilots on CT and MRI analysis
  • Hints for doctors when reviewing images
  • A base for fine-tuning on the clinic's own images
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Documents and OCR 3

Documents and OCR2026

jina-ocr-v1

Jina AI · Germany

Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.

  • Converting scans and PDFs to Markdown
  • Recognizing tables and formulas
  • Parsing invoices, acts and reports
Sizes
3.4B-A0.6B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Documents and OCRGGUF2025

olmOCR

Ai2 (Allen Institute for AI) · USA

A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.

  • Bulk digitisation of a PDF archive
  • Converting contracts and reports into text
  • Preparing documents for search and RAG
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCRNot maintained2024

Kosmos-2.5

Microsoft · USA

Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Converting resume scans into text that keeps its structure
  • Preparing documents for field extraction
  • Digitising paper forms
Sizes
about 1.4B
Hardware
from: 1 GPU
Commercial use allowedDetails

Tabular data 3

Tabular data2026

TabFM

Google Research · USA

Google's large tabular model: classification and regression from examples without training, with numeric and categorical columns. Weights are for non-commercial use only.

  • Pilot comparison with current scoring models
  • Exploring customer data
  • Testing churn hypotheses
Sizes
about 1.6B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Tabular dataGGUF2024–2025

TableGPT2 / TableGPT-R1

Zhejiang University · China

A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.

  • Answering questions about tables and data exports
  • Automated data analysis with generated code
  • A helper for BI and internal reporting
Sizes
7B – 72B
Hardware
from: 1 GPU
Commercial use allowedDetails
Tabular dataNot maintained2024

TabuLa-8B

ML Foundations · USA

A foundation model for predictions on tables: it classifies and forecasts from a handful of examples, with no separate task-specific training.

  • Classifying table rows from a few examples
  • Predicting a value from a data row
  • Quickly testing hypotheses on new datasets
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Faces 3

Faces2025

InfiniteYou

ByteDance · China

FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.

  • Portraits from one photo with a precise scene description
  • Testing characters for advertising
  • Comparing face-preservation methods
Sizes
adapter for FLUX.1-dev
Hardware
from: 1 GPU
Non-commercial onlyDetails
Faces2024

PuLID

ByteDance · China

Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.

  • Portraits and avatars from one photo
  • Ad characters with a recognizable face
  • Photoshoot prototypes
Sizes
adapters for SDXL and FLUX.1-dev
Hardware
from: 1 GPU
Commercial use with conditionsDetails
FacesNot maintained2024

InstantID

InstantX (Xiaohongshu) · China

Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.

  • Portraits in different styles from one photo
  • Avatar and character sketches
  • Photoshoot prototypes
Sizes
adapter for SDXL
Hardware
from: 1 GPU
Non-commercial onlyDetails

Deepfake detection 3

Deepfake detection2024–2025

AIDE

Xiaohongshu, USTC and Shanghai Jiao Tong University · China

An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.

  • Checking realistic AI images without obvious artifacts
  • Comparing detectors on hard examples
  • Fine-tuning for your own type of content
Sizes
several experts based on ConvNeXt and CLIP
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Deepfake detection2022–2025

TruFor

GRIP, University Federico II of Naples · Italy

Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.

  • Checking scans of certificates and contracts for edits
  • Highlighting altered photo regions for an expert
  • Filtering out obviously redrawn documents before manual review
Sizes
a transformer model producing a map of suspicious regions
Hardware
from: 1 GPU
Non-commercial onlyDetails
Deepfake detectionNot maintained2023–2024

IML-ViT

Sichuan University and co-authors · China

An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.

  • Finding pasted and erased fragments in photos
  • Checking document scans for edits
  • A baseline when comparing manipulation-localization models
Sizes
a Vision Transformer based model
Hardware
from: 1 GPU
Commercial use allowedDetails

Fact-checking and judges 3

Fact-checking and judges2025

Atla Selene

Atla · UK

An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.

  • Scoring chatbot answers against your own criteria
  • Comparing two versions of a prompt or model
  • Filtering out weak answers before they reach a person
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Fact-checking and judgesNot maintained2024

Patronus Lynx

Patronus AI · USA

Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.

  • Finding invented facts in AI assistant answers
  • Checking that answers rest on the attached documents
  • Filtering out answers before they go to a customer
Sizes
8B and 70B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Fact-checking and judgesNot maintained2024

ArmoRM (RLHFlow)

RLHFlow · USA

An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.

  • Choosing the best of several candidate answers
  • Preparing data for model fine-tuning
  • Scoring assistant answers across several attributes
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Moderation and safety 2

Moderation and safetyOllama2025

gpt-oss-safeguard

OpenAI · USA

Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.

  • Moderation by internal company rules
  • Labeling disputed messages with an explanation
  • Checking reviews and listings before publishing
Sizes
20B – 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
Moderation and safetyGGUFNot maintained2024

WildGuard

Ai2 (Allen Institute for AI) · USA

An open Ai2 filter: in a single pass it determines whether a request is harmful, whether a reply is harmful, and whether the bot refused needlessly. Works in English.

  • Checking requests to the bot
  • Checking bot replies
  • Finding unnecessary bot refusals on harmless questions
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails

Voice assistants 2

Voice assistantsGGUF2026

PersonaPlex

NVIDIA · USA

A voice conversation partner based on Moshi that listens and speaks at the same time and can be interrupted. The role is set by text, the voice by a sample recording. English only.

  • A voice assistant with a set role
  • A conversation simulator for staff training
  • Voice interfaces without delay
Sizes
7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Voice assistants2025

Kimi-Audio

Moonshot AI · China

A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.

  • Speech recognition
  • Detecting emotions and sound events
  • Speech-to-speech voice dialogue
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails

Autonomous driving 2

Autonomous driving2025–2026

NVIDIA Alpamayo

NVIDIA · USA

Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.

  • Auto-labeling camera recordings to train your own driver assistance systems
  • Analyzing complex road scenes with text explanations
  • Testing autopilot systems in simulation on rare scenarios
Sizes
10B – 34B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Autonomous driving2026

Qwen-Drive

Alibaba (Qwen team) · China

An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.

  • A perception and planning prototype for autonomous vehicles on closed sites
  • Answering questions about camera recordings when reviewing incidents
  • Labeling road scenes to train your own models
Sizes
4B
Hardware
from: 1 GPU
Commercial use allowedDetails

Math and reasoning 1

Math and reasoningOllama2024–2025

QwQ

Qwen (Alibaba) · China

Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Checking the logic of contracts and internal policies
Sizes
32B
Hardware
from: 1 GPU
Commercial use allowedDetails

Text to speech 1

Text to speechRUGGUF2025–2026

Higgs Audio

Boson AI · USA

Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.

  • Expressive video voiceover
  • Voicing dialogues
  • Voice cloning
Sizes
about 3B – 8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Translation 1

TranslationRUNot maintained2023–2024

ALMA и X-ALMA

Johns Hopkins University and Microsoft · USA

Research translators based on Llama 2. The first ALMA covered 5 pairs with English, including Russian; X-ALMA expanded coverage to 50 languages.

  • Translation between English and Russian
  • Experiments with LLM-based translation
  • Base for fine-tuning a translator
Sizes
7B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Voice: speakers and sound 1

Voice: speakers and sound2025

SAM Audio

Meta · USA

A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.

  • Isolating one person's voice from a noisy recording
  • Removing unwanted sound from a video
  • Splitting a recording into separate sound sources
Sizes
small, base, large (5 to 15 GB of weights)
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Cybersecurity 1

CybersecurityGGUF2025

Trendyol Cybersecurity LLM

Trendyol · Turkey

Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.

  • Reviewing alerts and first-pass incident assessment
  • Explaining suspicious activity in reports
  • Helping the on-duty shift of a monitoring centre
Sizes
32B и 70B
Hardware
from: 1 GPU
Commercial use allowedDetails

Rerankers 1

RerankersGGUFNot maintained2023–2024

RankVicuna / RankLLaMA / RankZephyr (Castorini)

University of Waterloo, Castorini group · Canada

Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.

  • Reordering a long list of search results
  • Selecting sources for an AI assistant answer
  • Research comparisons of retrieval approaches
Sizes
7B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails

Where people usually hit the wall

Plan by video memory, not by parameter count: the weights, the context and the request queue all consume it. A model that fits the weights exactly will fail on a long document. Leave headroom and test on your own scenario rather than a demo prompt.

By hardware

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment