Image generationGGUF2025–2026
Alibaba · China
Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.
- Infographics for product cards
- Photo editing by text command
- Ad creatives
- Sizes
- 7B – 20B
- Hardware
- from: 1 GPU
TextRU2022–2026
Yandex · Russia
Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.
- Russian-language assistant and chatbot
- Answers based on the company knowledge base
- Base for industry-specific fine-tuning
- Sizes
- 8B – 100B
- Hardware
- from: Laptop
Tabular data2022–2026
Prior Labs (University of Freiburg) · Germany
A ready-made model for tables: it takes example rows and immediately predicts for new ones, without lengthy training or tuning. Only v2 is free for business; newer versions are non-commercial.
- Predicting customer churn from a CRM export
- Scoring applications and leads
- Classifying customers from 1C data
- Sizes
- from a few to hundreds of millions of parameters
- Hardware
- from: Laptop
Tabular data2025–2026
Amazon (AutoGluon team) · USA
Amazon's tabular model built into AutoGluon: classification and regression from examples with brief fine-tuning. Mitra-v2 handles more rows and columns.
- Predicting churn and repeat purchases
- Scoring applications
- Predicting deal or order value
- Sizes
- about 76M
- Hardware
- from: Laptop
Tabular data2025–2026
Stable AI (Beijing, with Tsinghua University) · China
A table model that alone can classify, predict numbers and fill in missing data. The lightweight LimiX-2M runs on an ordinary computer.
- Filling gaps in 1C and CRM exports
- Churn prediction and scoring
- Classifying customers and products
- Sizes
- 2M – 16M and LimiX-2
- Hardware
- from: Laptop
Faces2021–2026
InsightFace (deepinsight) · China
The most widely used open toolkit for face detection and recognition. Many identity-preserving image generators are built on it. The pretrained weights are non-commercial.
- Detecting and comparing faces in photos
- Face-based access in prototypes
- Face processing as part of other AI systems
- Sizes
- packages from 16 MB to 407 MB
- Hardware
- from: Laptop
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
Video2025–2026
NVIDIA · USA
NVIDIA's lightweight, fast video model. Produces 720p clips on a single GPU; a 4-step version enables quick generation.
- Quick clips for social media
- Bulk video generation
- Video from an image
- Sizes
- 2B – 5B
- Hardware
- from: 1 GPU
Forecasting2024–2026
Google · USA
A ready-made Google forecasting model: forecasts any time series without training on your data.
- Sales and demand forecasting
- Purchase and inventory planning
- Load and traffic forecasting
- Sizes
- 200M – 500M
- Hardware
- from: Laptop
TextRUOllama2025–2026
Liquid AI · USA
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
- Offline assistant on a laptop or phone
- Data extraction from documents
- Tool calling in apps
- Sizes
- 230M – 24B-A2B
- Hardware
- from: Laptop
Video2025–2026
Alibaba · China
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
- Short promo videos
- Animating product photos
- Videos for social media
- Sizes
- 1,3B – 14B
- Hardware
- from: 1 GPU
Video2024–2026
Lightricks · Israel
A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.
- Ad videos with sound
- Video from a product photo
- Voiced scenes for social media
- Sizes
- 2B – 22B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Meituan · China
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
- Agents for orders and service processes
- Corporate assistant
- Analysis of long documents
- Sizes
- 1B – 1.6T-A48B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
Microsoft · USA
Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.
- Filling in web forms and applications
- Collecting data from web portals without an API
- Checking websites against scenarios
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
RerankersGGUF2024–2026
Jina AI · Germany
Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.
- Refining search results before a chatbot answers
- Sorting retrieved PDF pages and slides
- Catalog and knowledge base search
- Sizes
- 33M – 2.4B
- Hardware
- from: Laptop
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
Image generationGGUF2026
Krea · USA
A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.
- Realistic photos for advertising
- High-resolution images
- Style fine-tuning
- Sizes
- 12B
- Hardware
- from: 1 GPU
Computer vision2025–2026
Roboflow · USA
Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.
- Object detection in video and photos
- Precise outlines of parts and defects
- Fine-tuning for your own object classes
- Sizes
- Nano – 2XL
- Hardware
- from: Laptop
Forecasting2025–2026
NXAI · Austria
A compact forecasting model on the xLSTM architecture, a leader in open benchmarks despite its small size. Runs fast on a regular CPU.
- Demand and sales forecasting
- Energy consumption forecasting
- Forecasts on modest hardware and on site
- Sizes
- about 35M to 82M
- Hardware
- from: Laptop
Image + textOllama2024–2026
Moondream (M87 Labs) · USA
A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.
- Finding and counting objects in photos
- Checking photos from field reports
- Captions and tags for a catalogue
- Sizes
- 2B – 9B-A2B
- Hardware
- from: Laptop
Documents and OCRRU2022–2026
Baidu (PaddlePaddle) · China
Classic lightweight PaddleOCR models: detecting and recognizing lines of text plus page layout. They run on CPUs and phones; there is a separate model for East Slavic languages, including Russian.
- Recognizing text on scans, photos and screens
- Reading labels, displays and markings in production and warehouses
- Page layout: tables, formulas, stamps, headings
- Sizes
- from 1.5M to tens of millions of parameters
- Hardware
- from: Laptop
Image generationGGUF2025–2026
HiDream.ai · China
Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.
- Image generation from descriptions
- Editing images with words
- Variations of product photos
- Sizes
- about 9B – 17B
- Hardware
- from: 1 GPU
3D2024–2026
VAST (TripoSR together with Stability AI) · China
VAST family: a 3D model from a single photo. TripoSR runs in under a second, TripoSG gives cleaner geometry, TripoSplat builds a scene from Gaussian points.
- 3D product model from a photo
- Object assets for games and AR
- Quick 3D prototype for printing
- Sizes
- up to 1.5B
- Hardware
- from: Laptop
Search and RAGRUGGUF2023–2026
Jina AI · Germany
Strong multilingual embeddings with long context; v5-omni understands text, images and audio. Recent versions are open for non-commercial use only.
- Search across documents in many languages
- Search across images and scans
- Classification and clustering
- Sizes
- 33M – 3.8B
- Hardware
- from: Laptop
TranslationRU2025–2026
Tencent · China
Tencent translators for 33 languages; the first version won the WMT25 competition. Russian is supported. The small 1.8B version runs on a laptop; the new Hy-MT2 is under Apache 2.0.
- Translating documents while keeping formatting
- Translation with a set glossary of terms
- Translating correspondence with Chinese partners
- Sizes
- 1.8B – 30B-A3B
- Hardware
- from: Laptop
Photo editing2023–2026
BRIA AI · Israel
BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.
- Cutting products out onto a white background
- Staff and expert photos without background
- Background removal in video
- Sizes
- 44M – 220M
- Hardware
- from: Laptop
Image + textOllama2023–2026
LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China
The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.
- Answering questions about photos and screenshots
- Describing products from a photo
- Frame-by-frame video analysis
- Sizes
- 0.5B – 72B
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
H Company · France
A French model family for controlling a browser and computer: precisely finds the right element on screen and handles multi-step tasks. The latest Holo3 and 3.1 are open under Apache 2.0.
- Working in web portals and legacy software without an API
- Filling in forms and applications
- Testing interfaces against scenarios
- Sizes
- 0.8B – 235B-A22B
- Hardware
- from: Laptop
Computer vision2024–2026
Meta · USA
Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.
- Pose and body keypoint detection
- Segmentation of body parts and clothing
- Separating a person from the background
- Sizes
- 0.1B – 5B
- Hardware
- from: Laptop
Forecasting2024–2026
THUML, Tsinghua University · China
A compact forecasting foundation model from the Tsinghua lab: trained on a large set of diverse series and fine-tunable on your own data.
- Forecasting demand and load
- Forecasting sensor readings on the shop floor
- Fine-tuning forecasts on your own history
- Sizes
- 84M (timer-base)
- Hardware
- from: Laptop
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
TranslationRU2020–2026
Helsinki-NLP, University of Helsinki · Finland
More than a thousand small translators, each for its own language pair. Russian-English and back are available. Fast even on a regular CPU.
- Bulk translation of short texts
- Translation right on the server without a GPU
- Translating reviews and requests before analysis
- Sizes
- 25M – 240M
- Hardware
- from: Laptop
Computer-use agents2024–2026
Show Lab (National University of Singapore) · Singapore
A lightweight model for working with interfaces: finds buttons and fields by description and performs actions on the web and on a phone. ShowUI-π can drag with the mouse.
- Clicking and filling in forms from a task description
- Web UI autotests
- An assistant on a low-end computer without the cloud
- Sizes
- 2B (ShowUI), about 500M (ShowUI-π)
- Hardware
- from: Laptop
Text analysisOllama2024–2026
NuMind · France
Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.
- Extracting company details, amounts and dates from invoices and contracts into JSON
- Parsing receipts, waybills and forms against a set template
- Converting scans to Markdown for search
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Computer vision2023–2026
Meta · USA
Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.
- Background removal from product photos
- Counting objects in photos
- Data labeling for training
- Sizes
- 91M – ~0,85B
- Hardware
- from: Laptop
Text2025–2026
Reka AI · USA
Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.
- Photo and video analysis (Edge)
- Object detection in images
- Reasoning tasks (Flash)
- Sizes
- 7B – 21B
- Hardware
- from: Laptop
Tabular data2025–2026
Lexsi Labs · India
Recent open models for tabular data: they predict from a few examples given in the prompt, with no task-specific training.
- Classification and forecasting on tables with no separate training
- Quickly testing models on new datasets
- Assessing features in large tables
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Image generationGGUF2025–2026
Meituan · China
Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.
- Image generation from descriptions
- Instruction-based photo editing
- Visuals for product cards
- Sizes
- 6B
- Hardware
- from: 1 GPU
Voice assistantsGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.
- Voice assistant on your own server
- Analyzing videos and documents
- Voice answers about a camera image
- Sizes
- 8B – 9B
- Hardware
- from: Laptop
Virtual try-on2026
FASHN AI · Israel
A rare open try-on model with a commercial license: mask-free, accepts a photo of the item on a model or a flat lay. Weights are about 2 GB.
- Product cards on a model without a photo shoot
- Fitting room on a store website
- Catalog from flat-lay clothing photos
- Sizes
- 972M
- Hardware
- from: 1 GPU
Computer-use agents2025–2026
Alibaba (Tongyi Lab, X-PLUG) · China
Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.
- Automating actions in mobile apps
- Working in desktop software without an API
- Testing apps against scenarios
- Sizes
- 2B – 32B
- Hardware
- from: Laptop
Tabular data2025–2026
Inria (Soda team) · France
An open tabular model from the creators of scikit-learn: classifies and predicts from examples without training and handles tables of up to hundreds of thousands of rows. The license allows business use.
- Predicting customer churn
- Scoring applications and deals
- Classifying customers from 1C and CRM data
- Sizes
- about 25–30M
- Hardware
- from: Laptop
Image generationGGUF2024–2026
Black Forest Labs · Germany
Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.
- Images for product cards
- Banners and covers
- Photo editing by description (Kontext)
- Sizes
- 4B – 32B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
Alibaba (Tongyi-MAI) · China
A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.
- Photorealistic ad images
- Images with English and Chinese text
- Bulk visual generation
- Sizes
- 6B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Alibaba (Quark) · China
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
- Live avatar for customer dialogue
- Endless broadcasts with a presenter
- Interactive characters
- Sizes
- 14B
- Hardware
- from: 1 GPU
Computer vision2023–2026
Ultralytics · USA
The most widely used real-time object detector: finds and marks items in video even on modest hardware. YOLOv5 came out back in 2020; the catalog starts from YOLOv8.
- Counting people, cars and goods on video
- Checking hard hats and workwear
- Spotting defects on the production line
- Sizes
- 2.4M – 68M
- Hardware
- from: Laptop
TranslationRUOllama2026
Google · USA
Translators based on Gemma 3 for 55 languages that can also translate text in images. Russian is supported. The 4B version fits on a laptop.
- Translating documents and correspondence
- Translating text from screenshots and photos
- Localizing websites and apps
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
3DGGUF2024–2025
Microsoft · USA
One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.
- 3D models of products and interiors from photos
- Assets for games and AR/VR
- Prototypes for 3D printing
- Sizes
- up to 4B (TRELLIS.2)
- Hardware
- from: 1 GPU
Computer-use agents2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.
- Automating actions in mobile apps
- Working in web interfaces without an API
- Testing apps against scenarios
- Sizes
- 9B – 18B
- Hardware
- from: 1 GPU
Computer vision2025
Meta · USA
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
- Search photos and videos by description
- Catalog labeling and tagging
- Search across audio and video (PE-AV)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Tabular dataGGUF2024–2025
Zhejiang University · China
A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.
- Answering questions about tables and data exports
- Automated data analysis with generated code
- A helper for BI and internal reporting
- Sizes
- 7B – 72B
- Hardware
- from: 1 GPU
3DGGUF2025
Meta · USA
Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.
- 3D model of an item from a catalog photo
- Estimating body pose and shape from a photo
- Try-on and AR scenarios
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Computer vision2023–2025
Meta · USA
Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.
- Image search by text
- Image classification without training
- Search research and prototypes
- Sizes
- 0.15B – 3.6B
- Hardware
- from: Laptop
VideoGGUF2025
Meituan · China
A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.
- Long videos
- Video from a photo
- Continuing an existing video
- Sizes
- 13.6B
- Hardware
- from: 1 GPU
Computer vision2023–2025
IDEA Research · China
Finds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.
- Finding objects by description without labeling
- Automatic data labeling for training
- Checking photos against requirements
- Sizes
- 172M – 3B
- Hardware
- from: Laptop
ForecastingGGUF2024–2025
Amazon · USA
Amazon forecasting models, among the most downloaded. Chronos-2 takes external factors into account: prices, promotions, weather.
- Demand forecasting with promotions and prices
- Inventory planning
- Forecasting revenue and customer flow
- Sizes
- 8M – 710M
- Hardware
- from: Laptop
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Mixedbread · Germany
Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.
- Search across a knowledge base
- Reranking results before a bot answers
- Product catalog search
- Sizes
- 17M – 1.5B
- Hardware
- from: Laptop
TextRU2025
Avito Tech · Russia
Avito's model based on Qwen3-8B, retrained for Russian: its own tokenizer makes Russian text 15–25% faster. Supports function calling.
- Product and listing descriptions in Russian
- Chatbot that calls internal services
- Request analysis and classification
- Sizes
- 7.9B
- Hardware
- from: Laptop
Image + textRU2025
Avito Tech · Russia
Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.
- Product descriptions from photos in Russian
- Checking that a photo matches its description
- Reading brands and text in images
- Sizes
- 7.4B
- Hardware
- from: 1 GPU
3D2024–2025
Tencent · China
Tencent's open 3D line: shape and texture from an image, at the level of paid services. Omni adds control of pose and shape, Part splits a model into parts.
- Textured 3D product models
- Characters and objects for games
- Splitting a model into parts for printing
- Sizes
- set of models: shape and textures
- Hardware
- from: 1 GPU
Voice assistantsRUGGUF2025
Alibaba (Qwen) · China
Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.
- Voice assistant for customers
- Analyzing calls and videos
- Voice answers about documents and images
- Sizes
- 3B – 30B-A3B
- Hardware
- from: Laptop
Text to SQLGGUF2024–2025
Prem AI · UK
A text-to-SQL model of just 1B parameters, designed to run locally so the database never leaves for external services.
- Local translation of questions into SQL with no internet access
- Query hints on modest hardware
- Embedding into internal analytics tools
- Sizes
- 1B
- Hardware
- from: Laptop
ForecastingGGUF2024–2025
Salesforce · USA
Salesforce's universal forecasting model for series with different frequencies and many variables. Weights are open for research only.
- Research forecasting pilots
- Comparison with other forecasting models
- Forecasts across many related series
- Sizes
- 11M – 311M
- Hardware
- from: Laptop
Virtual try-on2025
Kunbyte AI · China
Try-on beyond clothing: glasses, earrings, bags, hats, watches and other accessories. Works without a mask. Needs a GPU with 28 GB or more.
- Trying accessories and jewelry on a photo
- Product cards with an accessory on a model
- Online fitting room for eyewear and jewelry
- Sizes
- add-on for FLUX.1 Fill dev 12B
- Hardware
- from: 1 GPU
Virtual try-on2025
LavieAI and Sun Yat-sen University · China
Tries on several items at once: top, bottom, shoes, bag. Faster than earlier models thanks to caching. Non-commercial license.
- Building an outfit from several products on one model
- "Build a look" pilot on a website
- Lookbook prototypes without a shoot
- Sizes
- based on SD 1.5 inpainting
- Hardware
- from: 1 GPU
Photo editingGGUF2024–2025
Nankai University · China
An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.
- Bulk background removal from product photos
- Precise masks for design and print
- Cutting out people with hair for advertising
- Sizes
- about 220M (lightweight lite versions available)
- Hardware
- from: Laptop
Image generationGGUF2024–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.
- Placing a product or person into a new scene
- Instruction-based photo editing
- Generation from multiple references
- Sizes
- about 4B
- Hardware
- from: 1 GPU
Photo editing2025
ByteDance Seed · China
ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.
- Upscaling photos and video to 2K–4K
- Restoring old videos and photos
- Enhancing user photos before publishing
- Sizes
- 3B – 7B
- Hardware
- from: 1 GPU
Text to SQLGGUF2025
IDEA Research · China
A text-to-SQL model trained with reinforcement learning: it works through the schema and the conditions step by step before producing a query.
- Database queries for questions with several conditions
- Reviewing and fixing other people SQL queries
- An analyst helper inside a BI system
- Sizes
- 3B – 14B
- Hardware
- from: Laptop
Image generationGGUF2025
ByteDance Seed · China
A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.
- Photo editing in a conversation
- Answering questions about an image
- Image generation with explanations
- Sizes
- 14B-A7B
- Hardware
- from: 1 GPU
Text to SQL2025
Snowflake · USA
A Snowflake model for turning questions into SQL, trained with reinforcement learning by checking query results. The open 7B version is based on Qwen2.5-Coder.
- Plain-language questions to a data warehouse
- Generating SQL for reports and dashboards
- Checking and fixing analysts' queries
- Sizes
- 7B
- Hardware
- from: Laptop
Forecasting2025
THUML, Tsinghua University · China
A forecasting model that returns a set of possible scenarios rather than a single line — useful when you need a range for demand or load, not one number.
- Forecasting demand with a range of values
- Planning stock while accounting for spread
- Forecasting load on services and staff
- Sizes
- 128M (sundial-base)
- Hardware
- from: Laptop
Virtual try-on2024–2025
Sun Yat-sen University and Pixocial · China
A lightweight try-on model that runs on a regular GPU. There is a mask-free version and CatV2TON, which also tries clothes on in video.
- Trying a garment on a customer's photo
- Draft product cards on a model
- Try-on in a short video
- Sizes
- 899M
- Hardware
- from: Laptop
Image + textGGUF2024–2025
Hugging Face · France / USA
The smallest vision models from Hugging Face, starting at 256M; they run in a browser and on a phone. SmolVLM2 also understands video.
- Describing photos and video on low-end hardware
- Reading simple documents
- Embedding in mobile and offline apps
- Sizes
- 256M – 2.2B
- Hardware
- from: Laptop
Text to SQL2025
Alibaba · China
Alibaba models for turning questions into SQL, based on Qwen2.5-Coder. They work with different SQL dialects; a small 3B version suits modest hardware.
- Plain-language database questions
- Queries for different databases (PostgreSQL, MySQL, SQLite)
- Automating routine reports
- Sizes
- 3B – 32B
- Hardware
- from: Laptop
Moderation and safety2023–2025
Falconsai and Freepik · USA and Spain
Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.
- Filtering user photos and avatars
- Checking generated images before publishing
- Labeling a media library
- Sizes
- 86M
- Hardware
- from: Laptop
Image generationGGUF2024–2025
NVIDIA · USA
NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.
- Bulk image generation
- High-resolution visuals
- Real-time generation inside apps
- Sizes
- 0.6B – 4.8B
- Hardware
- from: Laptop
Text to SQL2025
Renmin University of China (RUC) · China
Models for turning questions into SQL, trained on millions of synthetic query examples across different databases. Three sizes for different hardware.
- Database questions without knowing SQL
- Generating queries for reports
- A base for fine-tuning on your own database schema
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer vision2023–2025
Google · USA
Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.
- Image search by text query
- Automatic catalog labeling and tagging
- Filtering prohibited content
- Sizes
- about 0.2B to 2B
- Hardware
- from: Laptop
Virtual try-on2025
Beijing University of Posts and Telecommunications and others · China
An all-round FLUX-based apparel toolkit: try-on, generating a model wearing a given item, and "taking off" an item from a person into a separate product photo.
- Try-on from a product photo
- Photo of a model wearing an item from a text description
- Clean product photo extracted from a shot of a person
- Sizes
- add-ons (LoRA) for FLUX.1 dev 12B
- Hardware
- from: 1 GPU
3D2024–2025
Stability AI · UK
Stability AI models that turn a single photo into a textured 3D model in about a second. SPAR3D lets you adjust the shape through a point cloud.
- 3D product cards from photos
- Assets for games and AR
- Quick mock-ups for design
- Sizes
- 1B – 2B
- Hardware
- from: 1 GPU
Photo editing2024–2025
Prama LLC · USA
A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.
- Cutting out products and people from photos
- Background removal in video
- Preparing photos for a catalog
- Sizes
- about 95M
- Hardware
- from: Laptop
Search and RAGRUOllama2019–2025
UKP Lab (TU Darmstadt), later Hugging Face · Germany
The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.
- Search across a knowledge base and FAQ
- Finding similar tickets and duplicates
- Grouping reviews and requests by topic
- Sizes
- about 20M – 470M
- Hardware
- from: Laptop
Virtual try-on2024
Meta AI (with King's College London) · USA
Meta's model for virtual try-on and changing a person's pose in a photo. Carefully transfers fine fabric details and lettering. MIT license, but the training data is non-commercial.
- Trying clothes on a model photo
- Changing the model pose in an existing shot
- Adding extra angles for a product card
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
Virtual try-on2024
Tencent and Fudan University · China
Tencent's transformer-based try-on: more accurately reproduces fabric texture, fine prints and garment length. Non-commercial license.
- Try-on of items with complex prints and textures
- Test product cards on a model
- Checking length and fit on a photo
- Sizes
- based on SD3
- Hardware
- from: 1 GPU
Visual document search2024
Alibaba (Tongyi Lab) · China
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
- Finding a product by photo
- Search across a catalogue of images and cards
- Search across document pages as images
- Sizes
- 2B and 7B
- Hardware
- from: 1 GPU
Image generationGGUF2022–2024
Stability AI · UK
The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.
- Illustrations and banners for advertising
- Backgrounds and scenes for product cards
- Fine-tuning to a brand style
- Sizes
- 0.9B – 8B
- Hardware
- from: Laptop
Visual document search2024
TIGER-Lab · Canada
Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.
- Search across a mixed archive of texts and images
- Search across document pages as images
- Finding similar cards and illustrations
- Sizes
- about 4B (based on Phi-3.5-V)
- Hardware
- from: 1 GPU
Forecasting2024
Auton Lab, Carnegie Mellon University · USA
A foundation model for numeric series: one engine is used for forecasting, anomaly detection, filling gaps and classification.
- Forecasting demand and load
- Detecting anomalies in sensor readings and metrics
- Filling gaps in historical data
- Sizes
- about 40M – 385M
- Hardware
- from: Laptop
Forecasting2024
IBM Research · USA
Tiny forecasting models from IBM: they run on an ordinary CPU and sit next to the business system without a separate GPU server.
- Forecasting sales and warehouse stock
- Forecasting energy use and equipment load
- Fast forecasts right on the company server
- Sizes
- very small: TinyTimeMixers have about 1M parameters
- Hardware
- from: Laptop
Forecasting2024
The Time-MoE team · not disclosed
A forecasting model with a sparse architecture: only part of the network runs at each step, so it stays fast at a small size.
- Forecasting sales and stock levels
- Forecasting load on services and staff
- Planning purchases from history
- Sizes
- 50M and 200M
- Hardware
- from: Laptop
Tabular dataNot maintained2024
ML Foundations · USA
A foundation model for predictions on tables: it classifies and forecasts from a handful of examples, with no separate task-specific training.
- Classifying table rows from a few examples
- Predicting a value from a data row
- Quickly testing hypotheses on new datasets
- Sizes
- 8B
- Hardware
- from: 1 GPU
Image + textNot maintained2024
Microsoft · USA
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Sizes
- 0.23B – 0.77B
- Hardware
- from: Laptop
Photo editingNot maintained2023–2024
Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China
All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.
- Removing and replacing objects in photos
- Extending the frame to a required format
- Inserting a product or detail from a text description
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Photo editingNot maintained2024
Lvmin Zhang (author of ControlNet) · USA
Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.
- Matching product lighting to a new background
- Studio lighting for portraits without a reshoot
- Consistent lighting style across a catalog
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2023–2024
Defog · USA
One of the first open models that turn a plain-language question into an SQL query against a database. Available in Ollama, but newer competitors are already stronger.
- Answering managers' questions from the sales database without an analyst
- Drafting SQL queries for reports
- An assistant inside a BI system
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
3DNot maintained2024
Tencent ARC · China
Builds a 3D mesh from a single image in about 10 seconds: first it draws the object from several angles, then assembles the model from them.
- 3D model of an object from a photo
- Assets for games and visualizations
- Prototypes for 3D printing
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Virtual try-onNot maintained2024
KAIST and OMNIOUS.AI · South Korea
One of the best-known open virtual try-on models: moves a garment from a product photo onto a photo of a person, keeping prints and logos well. Non-commercial license.
- Pilot of a fitting room on a store website
- Prototype product cards on a model without a photo shoot
- Comparison with commercial try-on services
- Sizes
- based on SDXL
- Hardware
- from: 1 GPU
Text to SQLGGUFNot maintained2024
Chat2DB · China
A text-to-SQL model from the open Chat2DB database client: it supports different SQL dialects, with an English and Chinese model card.
- Turning a question into SQL inside a database client
- Drafting queries for different database engines
- Hints for developers working with a schema
- Sizes
- 7B
- Hardware
- from: Laptop
Virtual try-onNot maintained2024
Xiao-i Research · China
An early popular open try-on model: one version for upper-body garments, another for full-length outfits. Non-commercial license.
- Trying tops on a model photo
- Full-length try-on: tops, bottoms, dresses
- Fitting room prototype for testing
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
Image generationNot maintained2023–2024
Playground AI · USA
An SDXL-based model focused on aesthetics: vivid colors, contrast, portraits. Compatible with SDXL ecosystem tools.
- Aesthetic ad visuals
- Portraits and lifestyle images
- Post covers
- Sizes
- about 2.6B
- Hardware
- from: 1 GPU
Search and RAGRUNot maintained2022–2024
Microsoft · USA
Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.
- Search across a knowledge base and documents
- Finding answers for a chatbot (RAG)
- Finding similar requests and duplicates
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Virtual try-onNot maintained2024
KAIST · South Korea
A research try-on model from CVPR 2024, one of the first built on Stable Diffusion. Now mostly used as a comparison baseline.
- Pilot of upper-body garment try-on
- Comparing quality of different try-on models
- Training your own try-on on the open code
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
Text to SQLGGUFNot maintained2024
ChatDB · USA
A text-to-SQL model built on DeepSeek-Coder, aimed at complex questions spanning several tables and conditions.
- Complex queries joining several tables
- Answering database questions without an analyst
- Drafting SQL for reports and exports
- Sizes
- 7B
- Hardware
- from: Laptop
VideoNot maintained2023
Stability AI · UK
Stability AI's first open video model: turns a photo into a 2–4 second clip. Now behind newer models in quality.
- Animating product photos
- Short video intros
- Animating illustrations
- Sizes
- about 1.5B
- Hardware
- from: 1 GPU
Text to SQLNot maintained2023
RUCKBReasoning, Renmin University of China · China
An early line of open text-to-SQL models starting at 1B, including variants fine-tuned for specific database schemas.
- Turning an employee question into an SQL query
- Drafting warehouse queries for a report
- Embedding into a BI dashboard as a helper
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
Taehoon Kim (POSTECH) · South Korea
A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.
- Batch background removal from photos
- Replacing the background with a color or blur
- Background removal in video
- Sizes
- small (based on Swin-B)
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and others) · China
Transformer-based photo upscaling, more accurate than SwinIR on fine details. Versions for real noisy photos and a lightweight HAT-S.
- Upscaling product and interior photos
- Preparing images for print
- Sharpening archival photos
- Sizes
- 9M – 40M
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2023
Numbers Station · USA
One of the first open text-to-SQL lines, including very small versions from 350M that run on an ordinary PC.
- Turning a question into SQL from a table description
- Hints while writing queries
- A local analyst helper with no data leaving the company
- Sizes
- 350M – 7B
- Hardware
- from: Laptop
Text analysisRUNot maintained2022–2023
David Dale (cointegrated) and the community · Russia
Ready-made tiny rubert-tiny models for Russian text: detect rudeness and insults, sentiment and emotions. They run on a CPU in milliseconds.
- Filtering insults in Russian chats and comments
- Labeling reviews as positive, neutral or negative
- Spotting irritated customers in requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
RerankersRUNot maintained2022
UKP Lab and the Sentence Transformers community · Germany
The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.
- Reordering knowledge base search results
- Selecting passages before a chatbot answers
- Finding duplicates among tickets and product cards
- Sizes
- about 4M – 120M
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021–2022
Tencent ARC Lab · China
The classic for upscaling photos 2–4x while cleaning noise and compression artifacts. Lightweight, runs even on a CPU. Versions for drawings and anime.
- Upscaling old and small product photos
- Cleaning images of compression artifacts
- Preparing images for print
- Sizes
- about 17M
- Hardware
- from: Laptop
Computer visionNot maintained2021–2022
OpenAI · USA
The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.
- Image search by text query
- Automatic tags for a catalog
- Finding similar images
- Sizes
- about 0.15B – 0.6B
- Hardware
- from: Laptop
Computer visionRUNot maintained2022
Sber AI and SberDevices (ai-forever) · Russia
A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.
- Product search by photo and by Russian description
- Sorting images into categories without labeling
- Checking that a photo matches its caption
- Sizes
- 150M – 430M
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021
Samsung AI Center Moscow (with Skoltech) · Russia
Removes unwanted objects, text and watermarks from photos with clean background fill. Lightweight and fast; still the standard for this task.
- Removing price tags, people and clutter from photos
- Cleaning interior and real estate photos
- Removing text and dates from archival photos
- Sizes
- about 51M
- Hardware
- from: Laptop
Photo editingNot maintained2021
ETH Zurich · Switzerland
A transformer model for upscaling, denoising and removing JPEG artifacts from photos. Lightweight and proven; often embedded in other systems.
- Photo upscaling
- Image denoising
- Removing compression artifacts
- Sizes
- about 12M
- Hardware
- from: Laptop