MedicineGGUF2023–2026
FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China
A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Hints for doctors when reviewing images (Vision)
- Sizes
- 7B – 72B
- Hardware
- from: Laptop
Speech to textRUGGUF2025–2026
Microsoft · USA
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
- Transcribing long meetings with speaker labels
- Voicing podcasts and dialogues
- Real-time voice for assistants
- Sizes
- 0.5B – 9B
- Hardware
- from: Laptop
Music and sound2026
HeartMuLa Team · not disclosed
An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.
- Songs and jingles from lyrics
- Music for videos
- Transcribing song lyrics
- Sizes
- 3B
- Hardware
- from: 1 GPU
TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextGGUF2025–2026
Xiaomi · China
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
- Logic and calculation tasks
- Agents with tools
- Help for developers
- Sizes
- 7B – 1,02T-A42B
- Hardware
- from: Laptop
TextRU2024–2026
Sber · Russia
Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.
- Russian-language employee assistant on your own server
- Customer replies and request handling in Russian
- Working with contracts and internal policies
- Sizes
- 10B-A1.8B – 702B-A36B
- Hardware
- from: Laptop
Tabular data2025–2026
Amazon (AutoGluon team) · USA
Amazon's tabular model built into AutoGluon: classification and regression from examples with brief fine-tuning. Mitra-v2 handles more rows and columns.
- Predicting churn and repeat purchases
- Scoring applications
- Predicting deal or order value
- Sizes
- about 76M
- Hardware
- from: Laptop
Tabular data2025–2026
Layer 6 AI (TD Bank) · Canada
A tabular model from a Canadian bank's AI lab, trained on real tables rather than only synthetic ones. Version 1.2 Turbo made computation orders of magnitude faster.
- Scoring applications and customers
- Predicting churn
- Classifying transactions and customers
- Sizes
- about 60–80M
- Hardware
- from: Laptop
Text2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact text models that run directly on a device: laptop, phone or mini PC. The 1B and 2B MiniCPM5 models focus on tool calling and long context.
- A local chat assistant without the cloud
- Data extraction and text classification
- Tool calling and simple agents on low-end hardware
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Text2024–2026
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
- Reasoning, maths and technical questions
- Analysing long documents
- Agents and writing code
- Sizes
- 0.9B – 375B-A23B
- Hardware
- from: Laptop
TextGGUF2025–2026
Renmin University of China (GSAI) and Ant Group (inclusionAI) · China
Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.
- Fast generation of code and text
- Agent scenarios with long context
- Research into alternatives to standard LLMs
- Sizes
- 8B – 100B (MoE)
- Hardware
- from: 1 GPU
Autonomous driving2020–2026
comma.ai · USA
An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.
- A research testbed for driver assistance systems
- Studying driver attention monitoring with an in-cabin camera
- Comparison with your own lane-keeping algorithms
- Sizes
- compact, designed for an in-vehicle device
- Hardware
- from: Laptop
Visual document searchGGUF2026
Tencent · China
Tencent models based on Qwen3.5 for searching scans and PDFs as images. According to the model card, among the top of the ViDoRe leaderboard at release.
- Search across scans and PDFs without OCR
- RAG over reports with tables and charts
- Search across document archives
- Sizes
- 4.5B – 8B
- Hardware
- from: 1 GPU
TextRU2026
SberDevices (ai-forever) · Russia
A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.
- Experiments with faster generation
- Fine-tuning small models for your own tasks
- Research
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
Deepfake detection2023–2026
Adobe Research and University of Surrey · USA
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
- Marking images on the way out of your own pipeline
- Checking the provenance of a submitted image
- Linking with content provenance metadata
- Sizes
- model types Q and P with different mark capacity
- Hardware
- from: Laptop
VideoGGUF2025–2026
Sand AI · China
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
- Long videos with continuation
- Video with sound
- Animating images
- Sizes
- 4.5B – 114B-A6B
- Hardware
- from: 1 GPU
Video2025–2026
NVIDIA · USA
NVIDIA's lightweight, fast video model. Produces 720p clips on a single GPU; a 4-step version enables quick generation.
- Quick clips for social media
- Bulk video generation
- Video from an image
- Sizes
- 2B – 5B
- Hardware
- from: 1 GPU
Computer vision2024–2026
Microsoft Research · USA
Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.
- Measuring rooms and objects from photos
- 3D point cloud from a single shot
- Preparing data for robots and AR
- Sizes
- ViT-S – ViT-G
- Hardware
- from: Laptop
Search and RAGRU2024–2026
Sber (SberDevices) · Russia
Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.
- Search across Russian-language documents
- RAG for chatbots in Russian
- Classifying requests and reviews
- Sizes
- 480M – 10B-A1.8B
- Hardware
- from: Laptop
Robotics2025–2026
GigaAI · China
A robot control model trained mostly on synthetic data from a world model. It reduces spending on collecting data from real robots.
- Controlling a robot arm
- Fine-tuning on a small amount of your own data
- Sorting and assembly pilots
- Sizes
- 3.5B
- Hardware
- from: 1 GPU
Speech to textGGUF2024–2026
Moonshine AI (Useful Sensors) · USA
Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.
- Voice control of devices
- Offline recognition on a phone
- Live subtitles
- Sizes
- 27M – 245M
- Hardware
- from: Laptop
Speech to textGGUF2025–2026
IBM · USA
IBM speech models for recognizing and translating speech in English, several European languages and Japanese. Designed for enterprise use.
- Transcribing business meetings
- Translating speech into text in another language
- Voice assistants
- Sizes
- 470M – 8B
- Hardware
- from: Laptop
Text2025–2026
Ant Group (inclusionAI) · China
An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.
- Corporate assistant
- Agents for office processes
- Financial analytics (Fin version available)
- Sizes
- 7.9B-A1.3B – 1T
- Hardware
- from: Laptop
TextOllama2024–2026
IBM · USA
IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.
- Answers based on internal documents (RAG)
- Tool calling and agent work
- Data extraction and classification
- Sizes
- 350M – 34B
- Hardware
- from: Laptop
TextOllama2026
Meta Superintelligence Labs · USA
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
- Agents with tool calling
- Analysis of screenshots, charts and documents
- Multilingual assistant
- Sizes
- 30B
- Hardware
- from: 1 GPU
Text analysis2024–2026
Urchade Zaratiana and Fastino AI · France / USA
Finds the entities you need in text without training: just list what to look for (name, amount, date). GLiNER2 also classifies text. Multilingual versions understand Russian.
- Extracting names, amounts and dates from emails and contracts
- Parsing requests into CRM fields
- Classifying requests by topic
- Sizes
- about 50M to 500M
- Hardware
- from: Laptop
Documents and OCR2026
TeleAI (China Telecom) · China
A new lightweight document parsing model that led the OmniDocBench v1.6 benchmark at release. Handles pages photographed on a phone and crumpled pages well. Languages on the card: Chinese, English, Japanese.
- Recognising invoices and delivery notes photographed on a phone
- Recognising tables and formulas
- Converting documents to Markdown for RAG
- Sizes
- about 1.2B
- Hardware
- from: Laptop
Computer-use agents2025–2026
XLANG Lab (University of Hong Kong) · China
Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.
- Working in desktop software without an API
- Moving data between systems
- Running user scenarios for tests
- Sizes
- 7B – about 400B (MoE)
- Hardware
- from: 1 GPU
CodeOllama2026
DeepReinforce · not disclosed
Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.
- A developer agent in the terminal
- Fixing bugs from a task description
- Understanding and extending a large repository
- Sizes
- 9B – 397B
- Hardware
- from: 1 GPU
Weather and climate2024–2026
Allen Institute for AI (Ai2) · USA
A fast climate model emulator: simulates the atmosphere years and decades ahead on a single GPU. Coupled with an ocean model (SamudrACE) for long-term scenarios.
- Decades-long climate scenarios to assess long-term asset risks
- Large-scale what-if runs on temperature and precipitation
- Preparing data for crop yield and energy demand models
- Sizes
- checkpoint of about 1.8 GB
- Hardware
- from: 1 GPU
Weather and climate2023–2026
Google DeepMind · UK
Google DeepMind's family of global weather models: GraphCast (10-day forecast), GenCast (probabilistic ensemble) and WeatherNext 2 with cyclone forecasting. Since August 2026 the weights are cleared for commercial use.
- Medium-range weather forecasts for planning shifts, voyages and deliveries
- Probabilistic assessment of extreme weather for insurance portfolios
- Tropical cyclone track forecasts for marine and port operations
- Sizes
- from lightweight 1° versions to full 0.25°
- Hardware
- from: 1 GPU
Biology and chemistry2022–2026
AlQuraishi Lab (Columbia University) and the OpenFold consortium · USA
A fully open reproduction of AlphaFold 2 and then AlphaFold 3 under Apache 2.0, with training data. OpenFold3 predicts complexes of proteins, nucleic acids and ligands.
- Predicting structures of proteins and ligand complexes
- Fine-tuning on the company's own data (training code is open)
- An in-house structural analysis service without sending data outside
- Sizes
- a single set of weights per version
- Hardware
- from: 1 GPU
Autonomous driving2026
Alibaba (Qwen team) · China
An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.
- A perception and planning prototype for autonomous vehicles on closed sites
- Answering questions about camera recordings when reviewing incidents
- Labeling road scenes to train your own models
- Sizes
- 4B
- Hardware
- from: 1 GPU
Deepfake detection2025–2026
University of Michigan · USA
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
- Checking submitted photos and illustrations
- Filtering AI images in a content flow
- Flagging suspicious images for manual review
- Sizes
- 22M
- Hardware
- from: Laptop
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
Video2025–2026
Alibaba · China
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
- Short promo videos
- Animating product photos
- Videos for social media
- Sizes
- 1,3B – 14B
- Hardware
- from: 1 GPU
Code2025–2026
Kwaipilot (Kuaishou) · China
Kuaishou models for agentic development, trained to solve real tasks in repositories. KAT-Coder-V2.5-Dev (35B, 3B active) is the open version of their closed flagship.
- An agent that fixes tasks in the repository
- Code generation and refactoring
- Automating routine development tasks
- Sizes
- 32B – 72B, 35B-A3B
- Hardware
- from: 1 GPU
Image generationRU2022–2026
Sber (Kandinsky Lab) · Russia
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
- Images from Russian-language descriptions
- Short promo videos from text or a photo
- Instruction-based image editing
- Sizes
- 2B – 19B
- Hardware
- from: 1 GPU
Robotics2025–2026
NVIDIA · USA
"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.
- Synthetic video for training robots and self-driving vehicles
- Testing scenarios in simulation
- Robot control (Policy versions)
- Sizes
- 2B – 65B
- Hardware
- from: 1 GPU
Robotics2026
Ant Group (Robbyant) · China
A robot control model from Ant Group trained on a large volume of data from real robots. Version 2.0 works with different types of robot arms.
- Controlling a two-armed robot
- Fine-tuning for your own operation
- Assembly and sorting pilots
- Sizes
- 4B – 6B
- Hardware
- from: 1 GPU
Robotics2026
Xiaomi · China
Open robot control models from Xiaomi. Robotics-1 is designed for household and kitchen tasks, U0 combines scene understanding and action.
- Controlling a robot arm by command
- Household and service scenarios
- Base for fine-tuning
- Sizes
- 4B – 5B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2024–2026
Sber · Russia
Sber's models for Russian speech recognition, among the most accurate for Russian. Includes emotion recognition, v3 with punctuation, and a multilingual version (Russian, Kazakh, Kyrgyz, Uzbek).
- Transcribing calls in Russian
- Meeting minutes
- Voice control of services
- Sizes
- 220M – 600M
- Hardware
- from: Laptop
TextGGUF2025–2026
Meituan · China
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
- Agents for orders and service processes
- Corporate assistant
- Analysis of long documents
- Sizes
- 1B – 1.6T-A48B
- Hardware
- from: Laptop
TextRU2024–2026
T-Bank · Russia
T-Bank models fine-tuned from Qwen for Russian: they write and reason in Russian noticeably better than the original. T-Lite is 8B, T-Pro 32B on one GPU; T-Search is a multi-step search agent in Russian and English.
- Russian-language support chatbot
- Analysis of requests and documents in Russian
- Answers based on the company knowledge base
- Sizes
- 7B – 36B-A3B
- Hardware
- from: Laptop
TextGGUF2025–2026
Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland
Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.
- Multilingual assistant
- Answers based on documents
- Analysis of images and scans (v1.5)
- Sizes
- 0.5B – 70B
- Hardware
- from: Laptop
TextGGUF2026
Thinking Machines Lab · USA
Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.
- Flagship-level corporate assistant
- Analysis of documents, images and audio
- Programming help
- Sizes
- 276B-A12B, 975B-A41B
- Hardware
- from: Cluster
CodeGGUF2026
Poolside · USA
Models for agentic programming: they edit code in a repository on their own. The small XS runs on a Mac with 36 GB of memory; S 2.1 has a 1M-token context.
- Coding agent for in-house development
- Bug fixing and code improvements
- Working with large codebases
- Sizes
- 33B-A3B – 225B-A23B
- Hardware
- from: 1 GPU
Image + textGGUF2024–2026
Alibaba (AIDC-AI) · China
Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.
- Extracting data from invoices, contracts and delivery notes
- Table recognition
- Answering questions about photos and charts
- Sizes
- 0.9B – 80B-A3B
- Hardware
- from: Laptop
Voice: speakers and soundGGUF2022–2026
WeNet community · China
A set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.
- Voice verification of a customer during a call
- Finding repeat calls from the same person
- Splitting a recording by speaker
- Sizes
- from a few to tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2022–2026
Meta AI, then Alexandre Défossez · France
A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.
- Separating vocals from music in a recording
- Backing tracks and karaoke stems
- Cleaning speech in videos with background music
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2023–2026
RVC-Project community · China
The most widely used open voice conversion tool: a model for a specific voice trains on 10–30 minutes of recording and works in real time. Use only with the voice owner's consent.
- Voicing content with one brand voice
- Covers and vocal work
- Real-time voice changing
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
Microsoft · USA
Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.
- Filling in web forms and applications
- Collecting data from web portals without an API
- Checking websites against scenarios
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Image + text2026
OpenMOSS (Fudan University) · China
An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.
- Analyzing long videos and finding events by time
- Real-time streaming video analysis
- Understanding photos and documents
- Sizes
- about 11B
- Hardware
- from: 1 GPU
Weather and climate2024–2026
Microsoft Research · USA
A foundation model of Earth's atmosphere: global weather forecasts, plus separate versions for air quality and ocean waves. Computes a forecast in seconds instead of hours on a supercomputer.
- Your own forecast of temperature, wind and precipitation for company locations
- Sea state estimates for planning voyages and port operations
- Air pollution forecasts for industrial sites
- Sizes
- about 1.3B (a small test version is available)
- Hardware
- from: 1 GPU
Rerankers2026
Tencent · China
A pair of small Tencent models based on Qwen3 that pick the right skill for an AI agent for a given request: the embedding model finds candidates, the reranker chooses the best one.
- Choosing a tool or skill for an AI agent
- Routing requests between bot scenarios
- Search across a catalog of internal tools
- Sizes
- 0.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
CodeGGUF2025–2026
JetBrains · Czech Republic
JetBrains models for fast code autocompletion. Mellum2 (12B, 2.5B active) is already a full assistant: it writes and edits code, calls tools and reasons.
- Fast code autocompletion on your own server
- A developer assistant that does not send code to the cloud
- Fine-tuning on the company's code
- Sizes
- 4B – 12B-A2.5B
- Hardware
- from: Laptop
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Computer vision2025–2026
Roboflow · USA
Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.
- Object detection in video and photos
- Precise outlines of parts and defects
- Fine-tuning for your own object classes
- Sizes
- Nano – 2XL
- Hardware
- from: Laptop
Forecasting2025–2026
NXAI · Austria
A compact forecasting model on the xLSTM architecture, a leader in open benchmarks despite its small size. Runs fast on a regular CPU.
- Demand and sales forecasting
- Energy consumption forecasting
- Forecasts on modest hardware and on site
- Sizes
- about 35M to 82M
- Hardware
- from: Laptop
Speech to textRUGGUF2026
Alibaba (Qwen) · China
Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.
- Transcribing calls and meetings
- Video subtitles
- Multilingual recognition
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
Speech to textGGUF2026
Cohere · Canada
Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.
- Transcribing meetings and interviews
- Subtitles
- Searching an audio archive
- Sizes
- 2B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Zyphra · USA
Speech synthesis with voice cloning and fine control over emotion, speed and pitch.
- Voice cloning
- Emotional voiceover
- Voicing videos
- Sizes
- about 1.6B
- Hardware
- from: Laptop
Music and soundRUGGUF2025–2026
ACE Studio and StepFun · China
Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.
- Songs and jingles for ads
- Background music for videos
- Demo versions of tracks
- Sizes
- about 2B – 4B
- Hardware
- from: Laptop
Moderation and safety2024–2026
GLiNER community (Fastino, Knowledgator, NVIDIA and others) · USA
Small GLiNER-based models for finding personal data: passports, phone numbers, accounts, addresses. Data types are set in words. Russian is not officially supported.
- Masking personal data before cloud AI
- Finding passport data and bank details in documents
- Checking data exports for leaks
- Sizes
- about 200M to 500M
- Hardware
- from: Laptop
Documents and OCR2026
Baidu · China
Baidu's OCR model building on DeepSeek-OCR ideas: processes multi-page documents and PDFs in a single pass and outputs structured text. Claimed to be multilingual, but the language list is not published.
- Converting multi-page PDFs and scans to text and Markdown
- Recognizing contracts, invoices and reports
- Preparing document archives for search and RAG
- Sizes
- 3.3B
- Hardware
- from: Laptop
Documents and OCRRU2022–2026
Baidu (PaddlePaddle) · China
Classic lightweight PaddleOCR models: detecting and recognizing lines of text plus page layout. They run on CPUs and phones; there is a separate model for East Slavic languages, including Russian.
- Recognizing text on scans, photos and screens
- Reading labels, displays and markings in production and warehouses
- Page layout: tables, formulas, stamps, headings
- Sizes
- from 1.5M to tens of millions of parameters
- Hardware
- from: Laptop
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
VideoGGUF2025–2026
Zhipu AI (Z.ai) and Tsinghua University · China
Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.
- Transferring an actor's motion to a character
- Replacing a character in a finished video
- Animating mascots and illustrations
- Sizes
- 14B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
HiDream.ai · China
Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.
- Image generation from descriptions
- Editing images with words
- Variations of product photos
- Sizes
- about 9B – 17B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Meituan · China
Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.
- News or course presenter videos
- Promo videos with a talking character
- Singing and voice-over
- Sizes
- based on LongCat-Video 13.6B
- Hardware
- from: 1 GPU
3D2024–2026
VAST (TripoSR together with Stability AI) · China
VAST family: a 3D model from a single photo. TripoSR runs in under a second, TripoSG gives cleaner geometry, TripoSplat builds a scene from Gaussian points.
- 3D product model from a photo
- Object assets for games and AR
- Quick 3D prototype for printing
- Sizes
- up to 1.5B
- Hardware
- from: Laptop
Robotics2025–2026
Ai2 (Allen Institute for AI) · USA
A fully open robot control model that first "reasons" about space and trajectory, then acts. Its reasoning can be checked.
- Controlling a robot arm with explainable steps
- Fine-tuning for your own robot
- Research pilots
- Sizes
- 5B – 8B
- Hardware
- from: 1 GPU
Text to speechRUGGUF2025–2026
Resemble AI · USA
Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- about 350M – 500M
- Hardware
- from: Laptop
Text to speechRU2025–2026
OpenMOSS (Fudan University) · China
A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.
- Voicing podcasts and dialogues
- Voice for an assistant
- Voice cloning
- Sizes
- 100M – 8.5B
- Hardware
- from: Laptop
TextGGUF2025–2026
StepFun · China
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
- High-load agents
- Analysis of documents with diagrams and screenshots
- Help for developers
- Sizes
- 8B – 321B
- Hardware
- from: 1 GPU
Image + textOllama2023–2026
LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China
The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.
- Answering questions about photos and screenshots
- Describing products from a photo
- Frame-by-frame video analysis
- Sizes
- 0.5B – 72B
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Image + textGGUF2025–2026
Kuaishou · China
Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.
- Analysing and describing short videos
- Reviewing clips and content
- Finding the right moment in a video
- Sizes
- 8B – 671B-A37B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Baidu (PaddlePaddle) · China
A compact document parsing model from the popular PaddleOCR toolkit. Per the model card it supports 109 languages, including Russian; version 1.6 leads the OmniDocBench benchmark.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables, formulas and stamps
- Converting scans to Markdown and JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
Shanghai AI Laboratory (OpenDataLab) · China
A popular open tool for converting PDFs to Markdown with its own small model. MinerU2.5-Pro was improved through data alone, without growing in size. Languages on the card: Chinese and English.
- Converting PDF reports and contracts to Markdown
- Recognising tables and formulas
- Preparing documents for RAG and search
- Sizes
- 0.9B – 1.2B
- Hardware
- from: Laptop
Biology and chemistry2022–2026
EvolutionaryScale / Chan Zuckerberg Biohub (ESM-2 — Meta AI) · USA
Protein language models: they understand amino acid sequences, predict structure (ESMFold2) and help with protein design. Since 2026 all open versions are under MIT.
- Protein embeddings for predicting properties (stability, solubility)
- Predicting 3D structures of proteins and complexes
- Screening enzyme and antibody design candidates before lab work
- Sizes
- 8M – 15B (ESM-2), 300M – 6B (ESM C), 1.4B (open ESM3)
- Hardware
- from: Laptop
Forecasting2024–2026
THUML, Tsinghua University · China
A compact forecasting foundation model from the Tsinghua lab: trained on a large set of diverse series and fine-tunable on your own data.
- Forecasting demand and load
- Forecasting sensor readings on the shop floor
- Fine-tuning forecasts on your own history
- Sizes
- 84M (timer-base)
- Hardware
- from: Laptop
Avatars2024–2026
Fudan University · China
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
- Presenter video from a photo and audio
- Long training videos
- Live avatar
- Sizes
- about 1B – 5B
- Hardware
- from: 1 GPU
Search and RAGRUOllama2024–2026
IBM · USA
Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.
- Search across corporate documents
- RAG on a regular server without a GPU
- Reranking results
- Sizes
- 30M – 311M
- Hardware
- from: Laptop
ForecastingGGUF2025–2026
Datadog · USA
A Datadog forecasting model trained on server and application metrics. Especially strong for IT monitoring: load, latency, errors.
- Server load forecasting
- Anomaly detection in metrics
- Capacity planning
- Sizes
- 4M – 2.5B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.
- Voice cloning
- Voicing videos and audiobooks
- Voice for an assistant
- Sizes
- 0.5B – 2.3B
- Hardware
- from: Laptop
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
Arcee AI · USA
An American family of MoE models trained from scratch: Nano, Mini and Large. Trinity-Large-Thinking (398B) reasons before answering.
- Agents with tool calling
- Reasoning tasks
- Corporate assistant on your own servers
- Sizes
- 6B – 398B-A13B
- Hardware
- from: Laptop
TranslationRU2020–2026
Helsinki-NLP, University of Helsinki · Finland
More than a thousand small translators, each for its own language pair. Russian-English and back are available. Fast even on a regular CPU.
- Bulk translation of short texts
- Translation right on the server without a GPU
- Translating reviews and requests before analysis
- Sizes
- 25M – 240M
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2026
IBM · USA
IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.
- Checking bot requests and replies
- Finding made-up facts in knowledge-base answers
- Checking your own rules written as text
- Sizes
- 38M – 8B
- Hardware
- from: Laptop
Moderation and safety2026
OpenAI · USA
Finds and hides personal data: names, addresses, phone numbers, emails, account numbers, passwords. Runs even in the browser. Trained mostly on English.
- Removing personal data from text before sending it to cloud AI
- Finding passwords and keys in texts
- Anonymizing correspondence for analytics
- Sizes
- 1.5B (50M active)
- Hardware
- from: Laptop
Image + textOllama2025–2026
IBM · USA
Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.
- Extracting fields from forms and invoices
- Turning charts and tables into data
- Answering questions about documents
- Sizes
- 2B – 4B
- Hardware
- from: Laptop
Text analysisOllama2024–2026
NuMind · France
Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.
- Extracting company details, amounts and dates from invoices and contracts into JSON
- Parsing receipts, waybills and forms against a set template
- Converting scans to Markdown for search
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Text to speech2025–2026
Kyutai · France
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
- Streaming speech transcription for voice bots
- Voicing replies with minimal delay
- Speech synthesis on a server without a GPU (Pocket TTS)
- Sizes
- 100M (Pocket TTS) – 2.6B
- Hardware
- from: Laptop
TextGGUF2025–2026
ServiceNow · USA
ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.
- A reasoning assistant for internal services
- Tool calling and enterprise agents
- Analysing screenshots and documents with images
- Sizes
- 5B – 15B
- Hardware
- from: Laptop
CodeOllama2025–2026
Essential AI · USA
An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.
- Writing and fixing code
- A developer agent on a single GPU
- Solving technical and scientific problems
- Sizes
- 8B
- Hardware
- from: Laptop
Search and RAGGGUF2025–2026
Octen · USA / Singapore
Qwen3-Embedding models fine-tuned by the startup Octen for search in legal, financial and medical texts. As of January 2026 the 8B version topped the RTEB leaderboard.
- Search across contracts and case law
- Search across financial reports
- Search across long documents up to 32K tokens
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
TextRU2024–2026
Ivan Bondarenko (bond005), Novosibirsk State University · Russia
Russian-language models for working with documents rather than chatting: knowledge-base answers, extraction of entities and facts from Russian text, long context.
- Answers to questions based on internal documents
- Extracting names, dates and amounts from contracts
- Short summaries of long Russian texts
- Sizes
- 1.5B – 7.6B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Zhejiang University · China
A medical model for text, images, 3D scans and video: from a light 4B to a large MoE. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images and CT scans
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 4B – 235B-A22B
- Hardware
- from: Laptop
Math and reasoningGGUF2025–2026
Princeton University · USA
Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.
- Formal verification of mathematical workings
- Verifying code correctness
- Training
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Search and RAGGGUF2026
Perplexity · USA
Embeddings from the Perplexity search service. Some versions take into account the context of the whole document, not just a single fragment.
- Search across large document collections
- RAG that accounts for document context
- Website and catalog search
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
TextGGUF2024–2026
Sarvam AI · India
Indian models focused on 22 languages of India. Sarvam 30B and 105B (2026) are MoE models with strong reasoning and agent skills.
- Multilingual customer support
- Reasoning and calculation tasks
- Agents with tool calling
- Sizes
- 2B – 105B-A10B
- Hardware
- from: Laptop
Image + textGGUF2023–2026
Shanghai AI Laboratory (OpenGVLab) · China
A large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.
- Understanding documents, diagrams and charts
- Answering questions about photos
- Video analysis
- Sizes
- 1B – 241B-A28B
- Hardware
- from: Laptop
Image + text2024–2026
Ai2 (Allen Institute for AI) · USA
Fully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.
- Counting products and objects in photos
- Pointing to where an item is in an image
- Video analysis
- Sizes
- 1B-A7B – 72B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
rednote hilab (Xiaohongshu) · China
A multilingual document parsing model: text, tables, formulas and reading order in one pass. dots.mocr also turns charts and diagrams into vector SVG.
- Recognising invoices, contracts and delivery notes
- Converting tables into an editable format
- Converting charts and diagrams into vector format
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2026
Baidu (Qianfan) · China
A Baidu model that not only recognises a document but also answers questions about it. Per the model card it supports 192 languages, including Cyrillic.
- Recognising invoices, contracts and delivery notes, including in Russian
- Page layout analysis and table recognition
- Answering questions about a document
- Sizes
- 4B
- Hardware
- from: Laptop
Voice: speakers and soundRU2026
FireRedTeam (Xiaohongshu) · China
A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.
- Cutting recordings before speech recognition
- Separating speech from music and singing in broadcasts and videos
- Speech detection in voice bots
- Sizes
- compact, exact size not stated
- Hardware
- from: Laptop
Video2025–2026
Skywork AI (Kunlun) · China
An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.
- Game world prototypes without an engine
- Interactive demos and simulations
- Generating data to train agents
- Sizes
- 1.8B – 17B
- Hardware
- from: 1 GPU
Music and sound2025–2026
Alibaba Tongyi (FunAudioLLM) · China
Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
- Audio for video based on the scene
- Editing individual sounds in a track
- Sound effects from a description
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Avatars2026
SII-GAIR and Sand.ai · China
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
- Presenter video from a script
- Ad videos with a talking character
- Training videos with a narrator
- Sizes
- 15B
- Hardware
- from: 1 GPU
Search and RAGRU2026
Microsoft · USA
Microsoft's 2026 multilingual embeddings with context up to 32K tokens; Russian is on the language list. The 270M and 0.6B versions run on a regular server, 27B is the most accurate.
- Multilingual knowledge base search
- Picking passages for RAG
- Search across long documents
- Sizes
- 270M – 27B
- Hardware
- from: Laptop
Tabular data2025–2026
Lexsi Labs · India
Recent open models for tabular data: they predict from a few examples given in the prompt, with no task-specific training.
- Classification and forecasting on tables with no separate training
- Quickly testing models on new datasets
- Assessing features in large tables
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
CodeOllama2024–2026
Alibaba (Qwen team) · China
The broadest open coding family: from 0.5B for autocompletion to 480B for agents. Qwen3-Coder-Next (80B, 3B active) works as a developer agent on a single GPU.
- Code autocompletion in the editor
- An agent that edits code in the repository on its own
- Writing and refining scripts, SQL and integrations
- Sizes
- 0.5B – 480B-A35B
- Hardware
- from: Laptop
Code2026
Ai2 (Allen Institute for AI) · USA
Fully open developer agents from Ai2: weights, data and training recipe are all public. Designed so a company can cheaply fine-tune the agent on its own repository.
- An agent for fixing issues in code
- Fine-tuning the agent on an internal repository
- Automating small edits and tests
- Sizes
- 8B – 32B
- Hardware
- from: Laptop
Image generationGGUF2025–2026
Meituan · China
Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.
- Image generation from descriptions
- Instruction-based photo editing
- Visuals for product cards
- Sizes
- 6B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2026
LM Provers (CMU, Hugging Face, ETH Zurich, Project Numina) · USA, Switzerland, France
A small 4B model on Qwen3 that writes mathematical proofs in plain language almost at the level of large models. Runs on a laptop.
- Checking the logic of reasoning and workings
- Step-by-step explanations of solutions
- Training and olympiad preparation
- Sizes
- 4B
- Hardware
- from: Laptop
Voice assistantsGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.
- Voice assistant on your own server
- Analyzing videos and documents
- Voice answers about a camera image
- Sizes
- 8B – 9B
- Hardware
- from: Laptop
Virtual try-on2026
FASHN AI · Israel
A rare open try-on model with a commercial license: mask-free, accepts a photo of the item on a model or a flat lay. Weights are about 2 GB.
- Product cards on a model without a photo shoot
- Fitting room on a store website
- Catalog from flat-lay clothing photos
- Sizes
- 972M
- Hardware
- from: 1 GPU
Computer-use agents2025–2026
Alibaba (Tongyi Lab, X-PLUG) · China
Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.
- Automating actions in mobile apps
- Working in desktop software without an API
- Testing apps against scenarios
- Sizes
- 2B – 32B
- Hardware
- from: Laptop
Tabular data2025–2026
Inria (Soda team) · France
An open tabular model from the creators of scikit-learn: classifies and predicts from examples without training and handles tables of up to hundreds of thousands of rows. The license allows business use.
- Predicting customer churn
- Scoring applications and deals
- Classifying customers from 1C and CRM data
- Sizes
- about 25–30M
- Hardware
- from: Laptop
Biology and chemistry2024–2026
Arc Institute (with Together AI, Stanford, NVIDIA) · USA
DNA language models with context up to a million nucleotides: they assess the impact of mutations, annotate genomes and generate sequences. Evo 2 is trained on genomes from all domains of life.
- Assessing the likely harmfulness of genetic variants for research
- Annotating genomes of microorganisms and plants
- Finding promising sequences in breeding and synthetic biology
- Sizes
- 1B – 40B
- Hardware
- from: 1 GPU
Image + text2026
Sukhrob Nurali · not disclosed
A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.
- Moving a resume from PDF into a candidate record
- Filling a candidate database without manual typing
- Parsing resumes with different layouts and styling
- Sizes
- 8B, a fine-tune of Qwen3-VL-8B-Instruct
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
Alibaba (Tongyi-MAI) · China
A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.
- Photorealistic ad images
- Images with English and Chinese text
- Bulk visual generation
- Sizes
- 6B
- Hardware
- from: 1 GPU
Image generation2026
Zhipu AI (Z.ai) · China
A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.
- Posters and banners with text
- Infographics
- Illustrations for presentations
- Sizes
- 9B + 7B
- Hardware
- from: 1 GPU
Avatars2024–2026
Ant Group · China
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
- Presenter video from a photo and voice
- Avatar with gestures for presentations
- Voiced characters
- Sizes
- up to 1.3B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Alibaba (Quark) · China
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
- Live avatar for customer dialogue
- Endless broadcasts with a presenter
- Interactive characters
- Sizes
- 14B
- Hardware
- from: 1 GPU
MedicineGGUF2025–2026
Alibaba DAMO Academy · China
Alibaba's medical model based on Qwen2.5-VL: understands many types of medical images and medical text, and can reason step by step. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Ant Healthcare (Ant Group) and Zhejiang Provincial Medical Information Center · China
A large medical MoE model based on Ling-flash-2.0: 100B parameters with 6B active, so it answers quickly. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff on clinical questions
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 100B-A6B
- Hardware
- from: 1 GPU
Search and RAGRUOllama2025–2026
Alibaba (Qwen) · China
Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.
- Knowledge base search for RAG
- Reranking results before answering
- Search across scans, slides and screenshots
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Search and RAG2026
Voyage AI (MongoDB) · USA
The only open model in the Voyage 4 line: its vectors are compatible with the paid larger versions, so you can start locally and move to the API later.
- Document search on your own server
- RAG for small knowledge bases
- Finding similar texts
- Sizes
- about 340M
- Hardware
- from: Laptop
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Allen Institute for AI (Ai2) · USA
Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.
- Assistant and answers based on documents
- Reasoning tasks (Think versions)
- Fine-tuning on your data with a clear model history
- Sizes
- 1B – 32B
- Hardware
- from: Laptop
TextGGUF2024–2026
Prime Intellect · USA
Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.
- Reasoning and math tasks
- Programming help
- Agents with tool calling
- Sizes
- 10B – 106B-A12B
- Hardware
- from: 1 GPU
TextRUGGUF2024–2026
UTTER consortium (Unbabel, universities of Lisbon, Edinburgh, Amsterdam and others) · European Union
European language models trained on all EU languages and several others, with a focus on translation. Russian is supported. Permissive license.
- Translation and localization of texts
- Answering questions in different languages
- Draft emails for foreign partners
- Sizes
- 1.7B – 22B
- Hardware
- from: Laptop
Documents and OCROllama2025–2026
DeepSeek · China
An OCR model that compresses a page into a small number of visual tokens, so it processes large volumes quickly. Version 2 better understands reading order.
- Bulk recognition of scanned invoices and contracts
- Table recognition
- Converting PDFs to Markdown for search and RAG
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUOllama2026
Zhipu AI (Z.ai) · China
A lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables and formulas
- Extracting fields to JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
LightOn · France
A French 1B OCR model that converts a page into text in one pass and is fast on high volumes. Languages on the card: European languages, Chinese and Japanese; no Russian.
- Recognising invoices and contracts in European languages
- Table recognition
- Converting PDFs to text for search and RAG
- Sizes
- 0.9B – 1B
- Hardware
- from: Laptop
Computer-use agentsGGUF2026
Meituan · China
Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.
- Working in office and legacy software without an API
- Moving data between systems
- Running test scenarios
- Sizes
- 8B – 32B
- Hardware
- from: 1 GPU
Voice: speakers and soundRU2025–2026
Daily (Pipecat) · USA
Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.
- Voice bot does not interrupt the customer during pauses
- Fast reply when the customer has really finished
- An add-on to a standard speech detector in voice assistants
- Sizes
- 8M (v3) – 580M (v1)
- Hardware
- from: Laptop
VideoGGUF2026
OpenMOSS / MOSI · China
Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.
- Short clips with speech and sound from a description
- Ad scenes with dialogue
- Video prototypes for storyboards
- Sizes
- 32B-A18B
- Hardware
- from: 1 GPU
Weather and climate2024–2026
European Centre for Medium-Range Weather Forecasts (ECMWF) · Europe (intergovernmental organization)
ECMWF's weather neural network running operationally: a 15-day forecast four times a day, an ensemble version with 51 scenarios, and since version 2, ocean waves.
- Running your own forecast from open initial data
- Ensemble forecasts to estimate the probability of frost, downpours and storms
- Wave forecasts for marine operations
- Sizes
- checkpoint of about 1 GB
- Hardware
- from: 1 GPU
3DGGUF2024–2025
Microsoft · USA
One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.
- 3D models of products and interiors from photos
- Assets for games and AR/VR
- Prototypes for 3D printing
- Sizes
- up to 4B (TRELLIS.2)
- Hardware
- from: 1 GPU
Speech to textRU2025
Meta · USA
Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.
- Transcription in rare and local languages
- Digitizing oral archives
- Subtitles in many languages
- Sizes
- 300M – 7B
- Hardware
- from: Laptop
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
TextRUGGUF2024–2025
Vikhr Models · Russia
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
- Russian-language assistant on your own PC or server
- Knowledge-base answers (RAG)
- Texts and emails in Russian
- Sizes
- 0.5B – 24B
- Hardware
- from: Laptop
Image + text2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
Vision models from Zhipu: first CogVLM, then the GLM-V line. GLM-4.6V can call tools based on images and act as an agent operating an interface.
- Answering questions about photos and documents
- An agent that operates an interface from screenshots
- Analysing charts and reports
- Sizes
- 9B – 106B-A12B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
Alibaba (Tongyi-MAI) · China
Compact Alibaba models for working in smartphone and computer interfaces: they find elements and complete multi-step tasks. The small size allows running on an ordinary GPU.
- Automating actions in mobile apps
- Working in software without an API
- UI autotests
- Sizes
- 2B – 8B
- Hardware
- from: Laptop
Moderation and safety2025
ServiceNow · USA
A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.
- Screening chatbot requests for attacks and jailbreaks
- Filtering harmful model answers
- Monitoring the actions of AI agents that use tools
- Sizes
- 8B
- Hardware
- from: Laptop
TextGGUF2023–2025
Inception (G42), MBZUAI and Cerebras · UAE
A model family for Arabic and English, including Gulf dialects. Suits companies working with Arabic-speaking customers and government bodies in the region.
- A chatbot in Arabic and English
- Translating and summarising documents in Arabic
- Classifying customer requests
- Sizes
- 256M – 70B
- Hardware
- from: Laptop
3D2025
Meta and Carnegie Mellon University · USA
A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.
- 3D reconstruction of an object or room from photos
- Exporting the scene to COLMAP format for further processing
- Depth and camera pose estimation
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
Computer vision2025
Meta · USA
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
- Search photos and videos by description
- Catalog labeling and tagging
- Search across audio and video (PE-AV)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Tabular dataGGUF2024–2025
Zhejiang University · China
A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.
- Answering questions about tables and data exports
- Automated data analysis with generated code
- A helper for BI and internal reporting
- Sizes
- 7B – 72B
- Hardware
- from: 1 GPU
Deepfake detection2024–2025
Meta · USA
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
- Marking video created or processed by AI
- Finding your own mark in re-uploaded clips
- Protecting ad materials from being reused as someone else's
- Sizes
- a mark of 96 to 1024 bits
- Hardware
- from: Laptop
Text to speech2025
Nari Labs · South Korea
A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.
- Voicing dialogues and podcasts
- Ads with natural speech
- Training role-plays
- Sizes
- 1B – 2B
- Hardware
- from: Laptop
Voice: speakers and soundRU2020–2025
Silero · Russia
The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.
- Cutting calls and recordings before speech recognition
- Detecting when the customer is speaking in a voice bot
- Filtering out silence and noise to save on transcription
- Sizes
- about 2 MB
- Hardware
- from: Laptop
Video2025
Character.AI · USA
Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
- Short clips with talking characters
- Animating an image with voice-over
- Ad scene prototypes
- Sizes
- 11B
- Hardware
- from: 1 GPU
Satellite and geo2025
IBM and the European Space Agency (ESA) · USA / Europe
A multimodal Earth model: understands optical and radar imagery, terrain, vegetation index and land use maps, and can generate a missing data type (for example, a "see-through-clouds" image from radar).
- Analyzing fields and forests even in cloudy weather using radar imagery
- Land use maps for assessing plots
- Flood and wildfire assessment (ready-made fine-tunes available)
- Sizes
- tiny – large (checkpoints from ~200 MB to ~3.8 GB)
- Hardware
- from: Laptop
Rerankers2025
ZeroEntropy · USA
Rerankers built on Qwen3. The model card lists the target domains — finance, law, code, medicine, science; the stated language is English.
- Refining results before an AI assistant answers
- Sorting search results across contracts and reports
- Search across technical and scientific documentation
- Sizes
- zerank-2 — 4B (based on Qwen3-4B), plus a smaller "small" version
- Hardware
- from: Laptop
VideoGGUF2025
Meituan · China
A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.
- Long videos
- Video from a photo
- Continuing an existing video
- Sizes
- 13.6B
- Hardware
- from: 1 GPU
ForecastingGGUF2024–2025
Amazon · USA
Amazon forecasting models, among the most downloaded. Chronos-2 takes external factors into account: prices, promotions, weather.
- Demand forecasting with promotions and prices
- Inventory planning
- Forecasting revenue and customer flow
- Sizes
- 8M – 710M
- Hardware
- from: Laptop
Music and sound2025
ASLP-lab (Northwestern Polytechnical University) · China
Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.
- Songs and jingles from lyrics
- Music for videos
- Demo versions of tracks
- Sizes
- about 1.1B
- Hardware
- from: Laptop
Moderation and safetyOllama2025
OpenAI · USA
Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.
- Moderation by internal company rules
- Labeling disputed messages with an explanation
- Checking reviews and listings before publishing
- Sizes
- 20B – 120B
- Hardware
- from: 1 GPU
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Documents and OCRGGUF2025
Ai2 (Allen Institute for AI) · USA
A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.
- Bulk digitisation of a PDF archive
- Converting contracts and reports into text
- Preparing documents for search and RAG
- Sizes
- 7B
- Hardware
- from: 1 GPU
Biology and chemistry2024–2025
MIT (Jameel Clinic) and Recursion · USA
An open MIT-licensed alternative to AlphaFold 3: predicts structures of protein, DNA and small-molecule complexes; Boltz-2 estimates binding strength, BoltzGen designs new binding proteins.
- Predicting how a candidate molecule binds to a target protein
- Ranking compounds by predicted binding strength before synthesis
- Designing binder proteins for a given target
- Sizes
- checkpoints of about 2 GB
- Hardware
- from: 1 GPU
Search and RAGRUOllama2024–2025
Mixedbread · Germany
Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.
- Search across a knowledge base
- Reranking results before a bot answers
- Product catalog search
- Sizes
- 17M – 1.5B
- Hardware
- from: Laptop
Visual document search2025
Illuin Technology, EPFL, CentraleSupélec · France
A compact (250M) model for searching document pages as images. According to the authors, it matches models 10 times larger and runs without a GPU.
- Search across scans and PDFs on a modest server
- Indexing document archives
- Search across slides and manuals
- Sizes
- 250M
- Hardware
- from: Laptop
TextRU2025
Avito Tech · Russia
Avito's model based on Qwen3-8B, retrained for Russian: its own tokenizer makes Russian text 15–25% faster. Supports function calling.
- Product and listing descriptions in Russian
- Chatbot that calls internal services
- Request analysis and classification
- Sizes
- 7.9B
- Hardware
- from: Laptop
Image + textRU2025
Avito Tech · Russia
Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.
- Product descriptions from photos in Russian
- Checking that a photo matches its description
- Reading brands and text in images
- Sizes
- 7.4B
- Hardware
- from: 1 GPU
Speech to textRU2023–2025
Alpha Cephei · Russia
Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.
- Transcribing Russian calls and recordings without the cloud
- Voice control in apps and kiosks
- Low-latency streaming speech recognition
- Sizes
- about 45 MB – 1.8 GB
- Hardware
- from: Laptop
Voice assistantsRUGGUF2025
Alibaba (Qwen) · China
Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.
- Voice assistant for customers
- Analyzing calls and videos
- Voice answers about documents and images
- Sizes
- 3B – 30B-A3B
- Hardware
- from: Laptop
Moderation and safetyRUGGUF2025
Alibaba (Qwen) · China
Safety filters for 119 languages, Russian among them. The Stream version checks a bot's reply while it is being generated and can cut it off on the fly.
- Filtering bot requests in Russian
- Stopping a dangerous reply during generation
- Labeling messages by risk category
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Documents and OCR2025
IBM and Hugging Face · USA
Tiny models for the open Docling document converter: they turn a page into markup with tables, formulas and code. Run on an ordinary laptop.
- Converting PDFs and scans to Markdown for search and RAG
- Recognising tables in reports
- Processing invoices and contracts on an ordinary PC
- Sizes
- 256M – 258M
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
pyannoteAI (Hervé Bredin) · France
The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.
- Tagging calls: which part is the agent, which is the customer
- Meeting minutes with speaker labels
- Preparing recordings for transcription and analysis
- Sizes
- a few million parameters
- Hardware
- from: Laptop
Satellite and geo2023–2025
IBM and NASA · USA
Foundation models for Landsat and Sentinel-2 satellite imagery that account for image time series. Ready-made fine-tunes for floods, burn scars and crop types, plus a separate weather model, WxC.
- Mapping crops and field condition over the season
- Assessing flood zones and burn scars after natural disasters
- Monitoring changes in buildings and land use
- Sizes
- tiny – 600M (imagery), 2.3B (Prithvi WxC weather)
- Hardware
- from: Laptop
Search and RAGOllama2023–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.
- Search across English-language documents
- Picking passages for chatbot answers (RAG)
- Code search (bge-code)
- Sizes
- 24M – 9B
- Hardware
- from: Laptop
Text to SQLGGUF2024–2025
Prem AI · UK
A text-to-SQL model of just 1B parameters, designed to run locally so the database never leaves for external services.
- Local translation of questions into SQL with no internet access
- Query hints on modest hardware
- Embedding into internal analytics tools
- Sizes
- 1B
- Hardware
- from: Laptop
TextOllama2025
OpenAI · USA
OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.
- AI agent that calls internal systems
- Answers based on internal policies
- Drafts of emails and reports
- Sizes
- 20B, 120B
- Hardware
- from: 1 GPU
Avatars2025
Meituan · China
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
- Video dubbing with matched facial expressions
- Dialogue between two characters from audio
- Long videos with a presenter
- Sizes
- 14B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Moonshot AI and Project Numina · China, France
Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.
- Formal verification of mathematical workings
- Translating a problem from plain language into Lean
- Training and olympiad preparation
- Sizes
- 0.6B – 72B
- Hardware
- from: Laptop
TextRU2024–2025
Lomonosov Moscow State University Research Computing Center, LAIR lab (RefalMachine) · Russia
Qwen models adapted for Russian: a new tokenizer plus further training on Russian texts. As a result, Russian text is generated up to twice as fast as with the original model of the same size.
- Russian-language assistant on your own server
- Answers based on company documents (RAG) in Russian
- Analysis and summaries of long Russian texts
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
TextGGUF2025
ByteDance · China
An open ByteDance 36B model with up to 512K tokens of context and an adjustable thinking budget. Fits on a single powerful GPU.
- Analysis of long documents
- Agents with tools
- Corporate assistant
- Sizes
- 36B
- Hardware
- from: 1 GPU
Voice: speakers and sound2024–2025
Alibaba (Tongyi Lab) · China
Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.
- Noise suppression in conversation recordings
- Separating two voices speaking at once
- Improving old and phone recordings
- Sizes
- under 1B
- Hardware
- from: Laptop
MedicineGGUF2025
Intelligent Internet · UK
Reasoning medical models on Qwen3, designed to run on an ordinary computer. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff with the reasoning shown
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Speech to textRU2025
T-Bank · Russia
A compact T-Bank streaming model for recognizing Russian speech in phone calls. Works in real time even without a GPU.
- Transcribing phone calls
- Voice robots on the line
- Call quality control
- Sizes
- 72M
- Hardware
- from: Laptop
TextRUOllama2024–2025
Hugging Face · USA
Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.
- Simple on-device assistant
- Classification and routing of requests
- Base for fine-tuning on a narrow task
- Sizes
- 135M – 3B
- Hardware
- from: Laptop
TranslationRUGGUF2025
ByteDance Seed · China
A compact ByteDance translator for 28 languages, close in quality to large closed systems. Russian is supported. Ready-made compressed versions are available.
- Translating business correspondence and documents
- Translating product cards
- Translating technical and legal texts
- Sizes
- 7B
- Hardware
- from: Laptop
Photo editingGGUF2024–2025
Nankai University · China
An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.
- Bulk background removal from product photos
- Precise masks for design and print
- Cutting out people with hair for advertising
- Sizes
- about 220M (lightweight lite versions available)
- Hardware
- from: Laptop
Finance2025
Tsinghua University (NeoQuasar) · China
A foundation model for market candlestick data: trained on data from more than 45 exchanges, it forecasts prices and volumes. The largest version, large, is not open.
- Forecasting candlesticks and trading volumes
- Volatility estimation
- A base for fine-tuning on your own series
- Sizes
- 4.1M – 102M
- Hardware
- from: Laptop
Math and reasoningOllama2025
Agentica (Berkeley, Sky Computing Lab) and Together AI · USA
Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.
- Solving maths problems with step-by-step working
- Generating and checking code
- An agent for fixing bugs in a repository
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
- Marking generated and edited images
- Finding a marked fragment inside a collage
- Tracking which parts of a picture were made by AI
- Sizes
- a mark encoder and decoder for images
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
OpenCompass (Shanghai AI Laboratory) · China
A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring model answers against set criteria
- Checking an answer against a reference solution
- Comparing several models on your own data
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Image generationGGUF2024–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.
- Placing a product or person into a new scene
- Instruction-based photo editing
- Generation from multiple references
- Sizes
- about 4B
- Hardware
- from: 1 GPU
Robotics2025
Hugging Face · USA
A small robot control model that runs on a regular laptop. Trained on open data from the LeRobot community, suited to low-cost robot arms.
- Controlling a low-cost robot arm
- Quick robotization pilots and demos
- Training staff and students
- Sizes
- 450M
- Hardware
- from: Laptop
Photo editing2025
ByteDance Seed · China
ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.
- Upscaling photos and video to 2K–4K
- Restoring old videos and photos
- Enhancing user photos before publishing
- Sizes
- 3B – 7B
- Hardware
- from: 1 GPU
Image + textGGUF2025
Moonshot AI · China
An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.
- Analysing long PDFs and presentations
- Answering questions about video
- Operating interfaces from screenshots
- Sizes
- 16B-A3B
- Hardware
- from: 1 GPU
Voice: speakers and sound2024–2025
RVC-Boss and community · China
Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.
- Voicing texts with a specific narrator's voice
- Voice for a bot or assistant
- Dubbing training videos
- Sizes
- under 1B
- Hardware
- from: Laptop
Satellite and geo2025
MBZUAI · UAE
A compact research model for satellite imagery, trained on both optical (Sentinel-2) and radar (Sentinel-1) data. Narrower in scope and community than Prithvi and TerraMind.
- Classification and segmentation of satellite imagery after fine-tuning
- A base for a land monitoring prototype
- Sizes
- TerraFM-B (ViT-Base)
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Clouditera · China
A Chinese open family for cybersecurity: reviewing vulnerabilities, analysing logs and traffic, explaining commands and scripts.
- Reviewing vulnerabilities and drafting fix recommendations
- Analysing logs and reconstructing an attack chain
- Explaining suspicious commands and scripts
- Sizes
- 1.5B – 14B
- Hardware
- from: Laptop
CybersecurityGGUF2025
Trendyol · Turkey
Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.
- Reviewing alerts and first-pass incident assessment
- Explaining suspicious activity in reports
- Helping the on-duty shift of a monitoring centre
- Sizes
- 32B и 70B
- Hardware
- from: 1 GPU
Text to SQLGGUF2025
IDEA Research · China
A text-to-SQL model trained with reinforcement learning: it works through the schema and the conditions step by step before producing a query.
- Database queries for questions with several conditions
- Reviewing and fixing other people SQL queries
- An analyst helper inside a BI system
- Sizes
- 3B – 14B
- Hardware
- from: Laptop
Search and RAGGGUF2024–2025
TechWolf · Belgium
A model from a Belgian HR company: it turns job titles into vectors so you can find similar vacancies and resumes. A human makes the decision about a candidate; automatic screening without review must not be used.
- Matching job titles coming from different sources
- Finding similar vacancies and resumes by meaning
- Cleaning up the company job title reference list
- Sizes
- 109M – 278M
- Hardware
- from: Laptop
TextOllama2025
DeepSeek · China
A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.
- Complex calculations and logic checks
- Analysis of contracts and internal policies
- Help for developers
- Sizes
- 1,5B – 671B
- Hardware
- from: Laptop
CodeGGUF2025
ByteDance Seed · China
A compact 8B coding model from ByteDance in base, instruct and reasoning versions. Its training data was selected by the model itself, with almost no hand-written rules.
- Code autocompletion and generation
- Solving algorithmic problems
- A base for fine-tuning on your own stack
- Sizes
- 8B
- Hardware
- from: Laptop
Image generationGGUF2025
ByteDance Seed · China
A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.
- Photo editing in a conversation
- Answering questions about an image
- Image generation with explanations
- Sizes
- 14B-A7B
- Hardware
- from: 1 GPU
Text analysisRU2023–2025
deepvk (VK) · Russia
Russian encoders from the VK team: RuModernBERT reads long texts, USER produces vectors for search, GeRaCl classifies texts by topic without training.
- Classifying requests without labeled data
- Knowledge base search in Russian
- Analyzing long contracts
- Sizes
- 35M – 360M
- Hardware
- from: Laptop
Text to SQL2025
Snowflake · USA
A Snowflake model for turning questions into SQL, trained with reinforcement learning by checking query results. The open 7B version is based on Qwen2.5-Coder.
- Plain-language questions to a data warehouse
- Generating SQL for reports and dashboards
- Checking and fixing analysts' queries
- Sizes
- 7B
- Hardware
- from: Laptop
CodeRU2025
MTS AI (MWS AI) · Russia
A small coding assistant from MTS AI that understands requests in Russian. Runs locally, with plugins for VS Code and JetBrains.
- Code suggestions and completion in the editor
- Code explanations in Russian
- Drafts of tests and documentation
- Sizes
- 1.5B
- Hardware
- from: Laptop
Forecasting2025
THUML, Tsinghua University · China
A forecasting model that returns a set of possible scenarios rather than a single line — useful when you need a range for demand or load, not one number.
- Forecasting demand with a range of values
- Planning stock while accounting for spread
- Forecasting load on services and staff
- Sizes
- 128M (sundial-base)
- Hardware
- from: Laptop
Image + textGGUF2024–2025
Hugging Face · France / USA
The smallest vision models from Hugging Face, starting at 256M; they run in a browser and on a phone. SmolVLM2 also understands video.
- Describing photos and video on low-end hardware
- Reading simple documents
- Embedding in mobile and offline apps
- Sizes
- 256M – 2.2B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
ByteDance Seed · China
A model that looks at a screenshot and controls the mouse and keyboard itself: clicks, fills in fields, navigates menus. The first generation and 1.5-7B are open; UI-TARS-2 weights were not released.
- Working in legacy software without an API
- Filling in forms and moving data between systems
- UI autotests from plain-language scenarios
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Text to SQL2025
Alibaba · China
Alibaba models for turning questions into SQL, based on Qwen2.5-Coder. They work with different SQL dialects; a small 3B version suits modest hardware.
- Plain-language database questions
- Queries for different databases (PostgreSQL, MySQL, SQLite)
- Automating routine reports
- Sizes
- 3B – 32B
- Hardware
- from: Laptop
Moderation and safety2023–2025
Falconsai and Freepik · USA and Spain
Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.
- Filtering user photos and avatars
- Checking generated images before publishing
- Labeling a media library
- Sizes
- 86M
- Hardware
- from: Laptop
Voice assistants2025
Moonshot AI · China
A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.
- Speech recognition
- Detecting emotions and sound events
- Speech-to-speech voice dialogue
- Sizes
- 7B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2022–2025
OpenAI · USA
Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.
- Transcription of calls and video meetings
- Video subtitles
- Voice messages to text
- Sizes
- 39M – 1,5B
- Hardware
- from: Laptop
Image generationGGUF2024–2025
NVIDIA · USA
NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.
- Bulk image generation
- High-resolution visuals
- Real-time generation inside apps
- Sizes
- 0.6B – 4.8B
- Hardware
- from: Laptop
Video2024–2025
HPC-AI Tech · Singapore
A fully open video generation project: weights, code and training recipe. Version 2.0 at 11B makes video from text and from an image.
- Video from a text description
- Animating images
- Training your own video model
- Sizes
- up to 11B
- Hardware
- from: 1 GPU
Video2025
StepFun · China
A large 30B video model producing clips of up to 204 frames. Needs server hardware, but is open under MIT.
- Video from a description
- Animating images
- Sizes
- 30B
- Hardware
- from: Cluster
Avatars2024–2025
Tencent Music (Lyra Lab) · China
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
- Dubbing video into another language
- Live avatar in a video chat
- Editing speech in a finished video
- Sizes
- under 1B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of contracts and internal policies
- Sizes
- 32B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Stanford University · USA
A reasoning model trained on just a thousand problems. It can be told to think longer to answer a hard question more accurately.
- Calculations and formula checks
- Working through complex problems step by step
- Training
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Qihoo 360 · China
Reasoning models from Qihoo 360: a standard Qwen2.5 was fine-tuned for long reasoning using an open recipe; data and code are published.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Nomic AI · USA
Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.
- Search across documents and a knowledge base
- Code search
- Search across scans and PDFs without text recognition
- Sizes
- 137M – 7B
- Hardware
- from: Laptop
Text to speechGGUF2025
Canopy Labs · USA
Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.
- Real-time voice for an assistant
- Emotional voiceover
- Voice cloning
- Sizes
- 3B
- Hardware
- from: Laptop
Text to speechGGUF2025
Sesame · USA
A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.
- Voice for a conversational assistant
- Voicing dialogues
- Voice product prototypes
- Sizes
- 1B
- Hardware
- from: Laptop
Text to SQL2025
Renmin University of China (RUC) · China
Models for turning questions into SQL, trained on millions of synthetic query examples across different databases. Three sizes for different hardware.
- Database questions without knowing SQL
- Generating queries for reports
- A base for fine-tuning on your own database schema
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Finance2025
Shanghai University of Finance and Economics (SUFE) · China
A reasoning model for financial tasks based on Qwen2.5-7B: calculations, report analysis, regulatory questions. Trained on Chinese and English data.
- Financial calculations with step-by-step explanations
- Answering questions about financial statements
- Analyzing tables of financial data
- Sizes
- 7B
- Hardware
- from: Laptop
Deepfake detection2025
Desklib · India
A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.
- Checking submitted articles and reports
- Filtering templated reviews and applications
- First-pass check of student work
- Sizes
- 0.4B (DeBERTa-v3-large)
- Hardware
- from: Laptop
Math and reasoningGGUF2025
NovaSky (Sky Computing Lab, Berkeley) · USA
A Berkeley reasoning model trained for under 450 dollars. It showed that o1-preview-level reasoning can be reproduced with modest resources.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer vision2023–2025
Google · USA
Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.
- Image search by text query
- Automatic catalog labeling and tagging
- Filtering prohibited content
- Sizes
- about 0.2B to 2B
- Hardware
- from: Laptop
Robotics2024–2025
Stanford, Berkeley and partners · USA
The first large open vision-language-action model: a robot arm carries out commands like "put the apple in the bowl". OFT makes it several times faster.
- Controlling a robot arm by text command
- Pilots for robotizing simple operations
- Base for fine-tuning to your own robot
- Sizes
- 7B
- Hardware
- from: 1 GPU
Text to speechGGUF2024–2025
hexgrad (independent developer) · not disclosed
A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.
- Voicing articles and notifications
- Voice for apps without a GPU
- Bulk text voiceover
- Sizes
- 82M
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
Microsoft Research · USA
An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.
- Pilots in interface control
- Research projects spanning screens and robotics
- Analyzing screenshots with an action plan
- Sizes
- 8B
- Hardware
- from: 1 GPU
TextGGUF2025
HUMAIN (formerly SDAIA) · Saudi Arabia
A Saudi model for Arabic and English, trained from scratch. One 7B version is openly available.
- An Arabic-language assistant
- Answering questions about documents
- Writing and editing texts in Arabic
- Sizes
- 7B
- Hardware
- from: Laptop
Search and RAGRU2023–2025
Alibaba · China
Alibaba embeddings and rerankers for search: from tiny to 7B based on Qwen2. There is a multilingual mGTE version with long context.
- Semantic search across documents
- Reranking search results
- Clustering and classifying texts
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Text analysisGGUF2024–2025
Answer.AI and LightOn · USA / France
A modern replacement for classic BERT: faster, reads up to 8 thousand tokens at once. A base for your own classifiers. Trained on English and code; for Russian there is RuModernBERT.
- Classifying requests and documents
- Finding relevant passages in long texts
- Base for your own classifier after fine-tuning
- Sizes
- 150M – 395M
- Hardware
- from: Laptop
Photo editing2024–2025
Prama LLC · USA
A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.
- Cutting out products and people from photos
- Background removal in video
- Preparing photos for a catalog
- Sizes
- about 95M
- Hardware
- from: Laptop
Computer vision2022–2025
University of Sydney and JD Explore Academy · Australia / China
A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.
- Body keypoints in photos and video
- Motion analysis in sports and rehabilitation
- Monitoring work postures and safety practices
- Sizes
- 33M – about 1B
- Hardware
- from: Laptop
Image + text2023–2025
Alibaba DAMO Academy · China
Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
- Video description and short summary
- Finding a moment in a recording by question
- Tagging a video archive
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Search and RAGRUOllama2019–2025
UKP Lab (TU Darmstadt), later Hugging Face · Germany
The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.
- Search across a knowledge base and FAQ
- Finding similar tickets and duplicates
- Grouping reviews and requests by topic
- Sizes
- about 20M – 470M
- Hardware
- from: Laptop
Visual document search2025
LlamaIndex · USA
A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.
- Search across scans and PDFs without OCR
- Search across invoices, acts and contracts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Search and RAGRUOllama2024
Snowflake · USA
Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.
- Search across documents and knowledge bases
- Picking passages for RAG
- Search across reports and internal data
- Sizes
- 22M – 568M
- Hardware
- from: Laptop
Deepfake detection2024
Meta · USA
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
- Marking speech synthesized by your service
- Finding your own mark in third-party publications
- Checking whether synthesis was mixed into a call recording
- Sizes
- a watermark generator and detector, 16-bit message
- Hardware
- from: Laptop
Visual document search2024
Alibaba (Tongyi Lab) · China
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
- Finding a product by photo
- Search across a catalogue of images and cards
- Search across document pages as images
- Sizes
- 2B and 7B
- Hardware
- from: 1 GPU
Text to speechGGUF2024
Hugging Face · USA
Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.
- Choosing a voice by description
- Voicing videos
- Voice service prototypes
- Sizes
- 880M – 2.2B
- Hardware
- from: Laptop
Documents and OCR2024
StepFun · China
One of the first general-purpose new-generation OCR models: text, formulas, tables, sheet music and diagrams. Small and runs on low-end hardware, but already behind newer models.
- Recognising scanned invoices and contracts
- Converting tables into an editable format
- Recognising formulas and diagrams
- Sizes
- 580M
- Hardware
- from: Laptop
TranslationRU2024
deepvk (VK) · Russia
Compact Kazakh-Russian translators from VK. At 197M they translate as well as the 600M NLLB and run on a regular CPU.
- Translating requests from Kazakh to Russian
- Translating documents and instructions into Kazakh
- Bilingual customer support
- Sizes
- 197M
- Hardware
- from: Laptop
VideoGGUF2024
Genmo · USA
An open 10B video model with realistic motion. At release it was among the strongest open models; no updates now.
- Video from a description
- Short ad scenes
- Sizes
- 10B
- Hardware
- from: 1 GPU
Video2022–2024
hzwer (Zhewei Huang) and co-authors · China
Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.
- Increasing video frame rate
- Smooth slow-motion video
- Smoothing clips from AI generators
- Sizes
- lightweight model (size not stated on the model card)
- Hardware
- from: Laptop
Cybersecurity2024
cybersectony · not disclosed
A very light classifier for emails and links showing signs of phishing. It errs in both directions, so borderline emails are still reviewed by a person.
- Flagging suspicious incoming emails
- Checking links from correspondence before opening them
- A first-level filter in a mail gateway
- Sizes
- about 66M
- Hardware
- from: Laptop
Visual document search2024
TIGER-Lab · Canada
Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.
- Search across a mixed archive of texts and images
- Search across document pages as images
- Finding similar cards and illustrations
- Sizes
- about 4B (based on Phi-3.5-V)
- Hardware
- from: 1 GPU
Visual document search2024
LightOn · France
A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.
- Refining search results over scans and PDFs
- Selecting pages before an AI assistant answers
- Sorting retrieved slides and reports
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Forecasting2024
Auton Lab, Carnegie Mellon University · USA
A foundation model for numeric series: one engine is used for forecasting, anomaly detection, filling gaps and classification.
- Forecasting demand and load
- Detecting anomalies in sensor readings and metrics
- Filling gaps in historical data
- Sizes
- about 40M – 385M
- Hardware
- from: Laptop
Forecasting2024
IBM Research · USA
Tiny forecasting models from IBM: they run on an ordinary CPU and sit next to the business system without a separate GPU server.
- Forecasting sales and warehouse stock
- Forecasting energy use and equipment load
- Fast forecasts right on the company server
- Sizes
- very small: TinyTimeMixers have about 1M parameters
- Hardware
- from: Laptop
CodeOllama2024
01.AI · China
Coding models from 01.AI at 1.5B and 9B with a 128K-token context and support for 52 programming languages. A separate line next to the text Yi models.
- Code autocompletion and generation
- Explaining and refactoring code
- A programming assistant without the cloud
- Sizes
- 1.5B – 9B
- Hardware
- from: Laptop
Visual document search2024
University of Waterloo, Tevatron project · Canada
Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.
- Search across scans and PDFs without OCR
- Search across presentations and reports with complex layouts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Fact-checking and judges2024
Flow AI · not disclosed
A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.
- Checking AI assistant answers against the instruction
- Bulk scoring of exported conversations
- Quality control before rolling out changes
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Forecasting2024
The Time-MoE team · not disclosed
A forecasting model with a sparse architecture: only part of the network runs at each step, so it stays fast at a small size.
- Forecasting sales and stock levels
- Forecasting load on services and staff
- Planning purchases from history
- Sizes
- 50M and 200M
- Hardware
- from: Laptop
Image + textNot maintained2023–2024
Hugging Face · France / USA
Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.
- Answering questions about images
- Analysing documents and screenshots
- A base for fine-tuning
- Sizes
- 8B – 80B
- Hardware
- from: 1 GPU
Documents and OCRNot maintained2024
Microsoft · USA
Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.
- Converting resume scans into text that keeps its structure
- Preparing documents for field extraction
- Digitising paper forms
- Sizes
- about 1.4B
- Hardware
- from: 1 GPU
Image generationNot maintained2023–2024
Lvmin Zhang (Stanford) and the community · USA
An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.
- Image from a sketch or outline
- Keeping pose and composition
- Interior visualization from a floor plan
- Sizes
- 0.4B – 1.3B
- Hardware
- from: Laptop
VideoNot maintained2023–2024
Shanghai AI Lab and CUHK · China
A module that brings Stable Diffusion image models to life, turning them into short animations. One of the first open video technologies.
- Short animations in brand style
- Animated covers and banners
- Animated stickers
- Sizes
- motion module on top of SD 1.5 / SDXL
- Hardware
- from: Laptop
AvatarsNot maintained2024
Kuaishou (Kling) · China
Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.
- Animating portraits
- Transferring an actor's expressions to a character
- Mascot animation
- Sizes
- under 1B
- Hardware
- from: Laptop
Text analysisRUNot maintained2020–2024
SberDevices (ai-forever) · Russia
Sber's Russian-language encoders trained on large Russian corpora. A base for classifiers, NER and semantic search in Russian.
- Classifying requests in Russian
- Extracting names, amounts and dates after fine-tuning
- Detecting review sentiment
- Sizes
- about 30M to 430M
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
Vectara · USA
A small model that checks whether an AI answer is grounded in the source text or made up. Runs on a CPU and works well as a filter in RAG systems.
- Checking knowledge base chatbot answers for fabrications
- Quality control of document summaries
- Comparing language models by their tendency to make errors
- Sizes
- 110M
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
BAAI (Beijing Academy of Artificial Intelligence) · China
Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.
- Refining search results before a chatbot answers
- Sorting knowledge base search results
- Selecting the most relevant clauses of contracts and policies
- Sizes
- 278M – 9B
- Hardware
- from: Laptop
Moderation and safetyGGUFNot maintained2024
Ai2 (Allen Institute for AI) · USA
An open Ai2 filter: in a single pass it determines whether a request is harmful, whether a reply is harmful, and whether the bot refused needlessly. Works in English.
- Checking requests to the bot
- Checking bot replies
- Finding unnecessary bot refusals on harmless questions
- Sizes
- 7B
- Hardware
- from: 1 GPU
Image + textNot maintained2024
Microsoft · USA
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Sizes
- 0.23B – 0.77B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
01.AI · China
Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.
- Chat assistant on a single GPU
- Analysis of long documents
- Classification and data extraction from text
- Sizes
- 6B – 34B
- Hardware
- from: Laptop
Photo editingNot maintained2023–2024
Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China
All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.
- Removing and replacing objects in photos
- Extending the frame to a required format
- Inserting a product or detail from a text description
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Photo editingNot maintained2024
Lvmin Zhang (author of ControlNet) · USA
Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.
- Matching product lighting to a new background
- Studio lighting for portraits without a reshoot
- Consistent lighting style across a catalog
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
TextNot maintained2024
Equall · France
Language models for legal texts, fine-tuned on US and European legal corpora (based on Mistral and Mixtral). English only.
- Reviewing English-language contracts
- Spotting risks and non-standard terms
- Drafting legal memos
- Sizes
- 7B – 141B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2024
UC Santa Barbara and co-authors · USA
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
- Checking long articles and reports as a whole
- Filtering machine text in a publication flow
- Comparing detectors on your own data
- Sizes
- about 150M (Longformer-base)
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Huawei Noah's Ark Lab and partners · China
A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.
- Illustrations for articles and social media
- Backgrounds for product cards
- Quick visual drafts
- Sizes
- 0.6B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2024
Avignon University and Nantes University · France
Mistral 7B fine-tuned on PubMed Central papers, plus several merges with the general model. Compact and easy to run. Does not replace a doctor; decisions are made by a specialist.
- Searching and summarising medical papers
- Draft reference materials for staff
- Explaining medical terminology
- Sizes
- 7B
- Hardware
- from: Laptop
3DNot maintained2024
Tencent ARC · China
Builds a 3D mesh from a single image in about 10 seconds: first it draws the object from several angles, then assembles the model from them.
- 3D model of an object from a photo
- Assets for games and visualizations
- Prototypes for 3D printing
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
Hugging Face (H4) · USA
Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.
- Lightweight chat assistant
- Reference and starting point for your own fine-tuning
- Drafts of texts and replies
- Sizes
- 7B – 141B-A35B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
KAIST and LG AI Research (prometheus-eval) · South Korea
An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.
- Scoring chatbot answers on your own scale
- Comparing two answer options
- Quality checks before launching an AI service
- Sizes
- 7B – 8x7B
- Hardware
- from: Laptop
Text to speechNot maintained2024
MyShell and MIT · USA
Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.
- Voicing bot replies in foreign languages
- Voicing training materials
- Reading texts aloud on a server without a GPU
- Sizes
- small, runs in real time on a CPU
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2024
Chat2DB · China
A text-to-SQL model from the open Chat2DB database client: it supports different SQL dialects, with an English and Chinese model card.
- Turning a question into SQL inside a database client
- Drafting queries for different database engines
- Hints for developers working with a schema
- Sizes
- 7B
- Hardware
- from: Laptop
Text analysisNot maintained2024
University of Southern Denmark · Denmark
Turns a free-form occupation description into a standard HISCO code in 13 languages. Built for historical archives, but also useful for cleaning up job title reference lists. A human makes the decision about a candidate; automatic screening without review must not be used.
- Mapping mixed occupation names onto a single code
- Processing archives of HR and statistical data
- Preparing data for reporting
- Sizes
- based on CANINE-s, size not stated on the model card
- Hardware
- from: Laptop
TextRUNot maintained2023–2024
Sber (ai-forever) · Russia
Sber's Russian text-to-text model, successor to ruT5 (2021). Small and fast: fine-tuned for summarizing, paraphrasing and fixing errors in Russian text; ready-made SAGE spell-checking versions exist.
- Fixing spelling mistakes and typos in Russian text
- Short summaries and paraphrasing
- Normalizing requests and inquiries before processing
- Sizes
- 95M – 1.7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
CybersecurityGGUFNot maintained2023–2024
ZySec AI · India
A small open assistant for security professionals: questions about standards, reviewing threats and vulnerabilities, drafting internal documents.
- Answering questions about security policies and standards
- First-pass review of threat reports
- Drafting internal protection guidelines
- Sizes
- 2.8B и 7B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023–2024
Sichuan University and co-authors · China
An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.
- Finding pasted and erased fragments in photos
- Checking document scans for edits
- A baseline when comparing manipulation-localization models
- Sizes
- a Vision Transformer based model
- Hardware
- from: 1 GPU
Search and RAGRUNot maintained2022–2024
Microsoft · USA
Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.
- Search across a knowledge base and documents
- Finding answers for a chatbot (RAG)
- Finding similar requests and duplicates
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
ForecastingNot maintained2024
ServiceNow, Mila and partners · Canada
One of the first open out-of-the-box forecasting models. Tiny, gives a probabilistic forecast, now behind Chronos and TimesFM.
- Probabilistic sales forecast
- Quick forecasting pilots
- Baseline model for comparison
- Sizes
- 2.4M
- Hardware
- from: Laptop
Documents and OCRNot maintained2024
Microsoft · USA
One model for every document task: reading, answering questions about a page, extracting fields, classification. In HR it is used to parse resumes and attached scans. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from a resume and its attachments
- Answering questions about document content
- Classifying incoming documents
- Sizes
- 742M
- Hardware
- from: Laptop
Search and RAGRUOllamaNot maintained2024
BAAI · China
A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.
- Search across a document base
- RAG for a chatbot
- Finding similar requests and duplicates
- Sizes
- 568M
- Hardware
- from: Laptop
Satellite and geoNot maintained2023–2024
Allen Institute for AI (Ai2) · USA
Pretrained models from the Satlas project for Sentinel-2, Landsat and high-resolution aerial imagery. The predecessor of OlmoEarth, still used in TorchGeo.
- Detecting objects in imagery: solar farms, wind turbines, ships
- Mapping roads and buildings from aerial photos
- A starting point for fine-tuning your own geo model
- Sizes
- Swin-v2 and ResNet backbones (Base)
- Hardware
- from: Laptop
Photo editingNot maintained2023
Alibaba DAMO Academy · China
Colorizes black-and-white photos in natural colors. A lightweight model with a commercial-friendly license; a compact tiny version is available.
- Colorizing archival photos
- Color versions of historical photos for publications
- Family photo restoration service
- Sizes
- DDColor-T (tiny) and DDColor-L
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2023
Resemble AI · USA
A speech enhancement model: removes noise and restores lost frequencies so a muffled recording sounds studio-quality. Good for preparing a voice for voiceover.
- Restoring old and phone recordings
- Cleaning a voice before voiceover and cloning
- Improving audio in videos and podcasts
- Sizes
- under 1B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Intel · USA
A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.
- A simple chat assistant
- Experiments with running on Intel hardware
- A base for fine-tuning
- Sizes
- 7B
- Hardware
- from: Laptop
RerankersNot maintained2023
NetEase Youdao · China
An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.
- Search across a knowledge base and reference materials
- Reordering retrieved passages
- Picking answers for a support chatbot
- Sizes
- about 280M
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2023
Google · USA
Google's translator for more than 400 languages under a permissive license. Russian is supported. A good substitute for NLLB when commercial use is needed.
- Translating documents and emails
- Translating catalogs and product descriptions
- Translating into CIS and Asian languages
- Sizes
- 3B – 10B
- Hardware
- from: Laptop
Documents and OCRNot maintained2022–2023
Microsoft · USA
Small models that find tables on PDF and scanned pages and restore their structure: rows, columns, headers. The text inside is read by a separate OCR.
- Finding tables in reports, statements and invoices
- Restoring rows and columns for export to Excel
- Preparing tabular data for analysis and RAG
- Sizes
- 29M
- Hardware
- from: Laptop
Music and soundNot maintained2023
LAION · Germany
CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.
- Search sounds and music by description
- Automatic tags for an audio library
- Recognizing sound types (siren, breaking glass, voice)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Text to SQLNot maintained2023
RUCKBReasoning, Renmin University of China · China
An early line of open text-to-SQL models starting at 1B, including variants fine-tuned for specific database schemas.
- Turning an employee question into an SQL query
- Drafting warehouse queries for a report
- Embedding into a BI dashboard as a helper
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
Taehoon Kim (POSTECH) · South Korea
A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.
- Batch background removal from photos
- Replacing the background with a color or blur
- Background removal in video
- Sizes
- small (based on Swin-B)
- Hardware
- from: Laptop
Deepfake detectionNot maintained2022–2023
EURECOM · France
A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.
- Spotting synthetic speech in calls
- Checking voice messages and recordings
- Fine-tuning for your own data and codecs
- Sizes
- about 0.3B (wav2vec 2.0 XLS-R encoder)
- Hardware
- from: Laptop
TextRUGGUFNot maintained2022–2023
Sber (ai-forever) · Russia
Sber's multilingual model covering 61 languages, including languages of the peoples of Russia and the CIS. Separate fine-tunes exist for Buryat, Yakut, Tatar, Bashkir, Kazakh and others, rare for open models.
- Texts in languages of the peoples of Russia and the CIS
- Base for fine-tuning on a less common language
- Drafts and templates in several languages
- Sizes
- 1.3B – 13B
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and others) · China
Transformer-based photo upscaling, more accurate than SwinIR on fine details. Versions for real noisy photos and a lightweight HAT-S.
- Upscaling product and interior photos
- Preparing images for print
- Sharpening archival photos
- Sizes
- 9M – 40M
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2022–2023
Hendrik Schröter (University of Erlangen) · Germany
Lightweight real-time speech noise suppression that works even on a regular CPU and low-power devices. Removes hum, street and office noise while keeping the voice.
- Cleaning calls and voice messages of noise
- Preparing recordings before speech recognition
- Noise suppression for video calls
- Sizes
- about 2M
- Hardware
- from: Laptop
AvatarsNot maintained2023
Xi'an Jiaotong University and Tencent AI Lab · China
An older lightweight talking-head model: one photo plus audio becomes a video. Runs on weak hardware, but quality is noticeably below newer models.
- Talking photo for greetings
- Simple voiced avatars
- Sizes
- under 1B
- Hardware
- from: Laptop
TextRUGGUFNot maintained2023
Sber (ai-forever) · Russia
Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.
- Base for fine-tuning on a narrow Russian-language task
- Generating template Russian texts
- Experiments with Russian-language models without license restrictions
- Sizes
- 13B
- Hardware
- from: 1 GPU
Text to speechRUGGUFNot maintained2023
Suno · USA
One of the first open models to voice text with intonation, laughter and pauses. Supports about ten languages, including Russian. Now outdated.
- Draft voiceovers for videos
- Voice service prototypes
- Sound effects in speech
- Sizes
- about 300M – 1B
- Hardware
- from: Laptop
3DNot maintained2022–2023
OpenAI · USA
Early open OpenAI models that create a 3D object from text or an image in seconds. Quality is basic, but they are fast and easy to run.
- Rough 3D mock-ups from a description
- Quick object prototypes for games and AR
- Training and research pilots in 3D
- Sizes
- 40M – 1B
- Hardware
- from: Laptop
TextNot maintained2022–2023
Google · USA
Compact input-output models trained to follow instructions. Still used as a cheap base for classification, extraction and short answers.
- Classification of requests and documents
- Extracting fields from text
- Short answers and summaries
- Sizes
- 80M – 20B
- Hardware
- from: Laptop
Image + textNot maintained2023
Google · USA
Reads a document or a screenshot as an image and answers with structure: text, fields, answers to questions. In HR it is fine-tuned for resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting data from resumes and forms supplied as images
- Questions about the content of a scan
- Parsing tables and diagrams in documents
- Sizes
- 282M – 1.3B
- Hardware
- from: Laptop
TextNot maintained2022–2023
EleutherAI · USA
Fully open models from the non-profit lab EleutherAI: GPT-NeoX-20B and the Pythia series with published intermediate training checkpoints.
- Base model for fine-tuning
- Research into model behavior
- Simple text generation and completion
- Sizes
- 70M – 20B
- Hardware
- from: Laptop
Music and soundNot maintained2022
MIT · USA
A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.
- Sound event recognition
- Tagging an audio archive
- Detecting alarm sounds
- Sizes
- about 87M
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
SCUT DLVC Lab, South China University of Technology · China
A light model that takes both the text and the position of blocks on the page into account: trained in one language and transferable to others. Good for tagging fields in resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Tagging fields in resumes and forms
- Extracting data from forms and templates
- Parsing documents in several languages
- Sizes
- about 130M for the English version and about 280M for the multilingual one
- Hardware
- from: Laptop
Documents and OCRNot maintained2021–2022
Microsoft · USA
Recognizes a single line of text, including handwriting. The official weights are English only, but the model is often fine-tuned for other languages; there are community Russian versions.
- Recognizing handwritten lines in questionnaires and forms
- Recognizing printed lines after text detection on the page
- A base for fine-tuning to your own handwriting or font
- Sizes
- 62M – 608M
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
NAVER CLOVA · South Korea
Reads a scanned document and returns a filled-in field structure straight away, with no separate OCR step. In HR it is fine-tuned for parsing resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from forms and resumes
- Parsing scans of certificates and diplomas
- Detecting the type of an incoming document
- Sizes
- about 200M
- Hardware
- from: Laptop
RerankersRUNot maintained2022
UKP Lab and the Sentence Transformers community · Germany
The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.
- Reordering knowledge base search results
- Selecting passages before a chatbot answers
- Finding duplicates among tickets and product cards
- Sizes
- about 4M – 120M
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021–2022
Tencent ARC Lab · China
The classic for upscaling photos 2–4x while cleaning noise and compression artifacts. Lightweight, runs even on a CPU. Versions for drawings and anime.
- Upscaling old and small product photos
- Cleaning images of compression artifacts
- Preparing images for print
- Sizes
- about 17M
- Hardware
- from: Laptop
Computer visionNot maintained2021–2022
OpenAI · USA
The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.
- Image search by text query
- Automatic tags for a catalog
- Finding similar images
- Sizes
- about 0.15B – 0.6B
- Hardware
- from: Laptop
Photo editingNot maintained2021–2022
Tencent ARC Lab · China
Proven face restoration for old and compressed photos, with a commercial-friendly license. Often paired with Real-ESRGAN; the most used versions are 1.3 and 1.4.
- Restoring faces in old photos
- Enhancing avatars and profile photos
- Restoration in a photo shop or online service
- Sizes
- small, up to 0.1B
- Hardware
- from: Laptop
Text analysisRUNot maintained2019–2022
Meta · USA
A classic multilingual encoder for 100 languages, including Russian. The base of many sentiment, NER and embedding models, including BGE-M3.
- Detecting review sentiment in different languages
- Extracting names and organizations after fine-tuning
- Classifying requests
- Sizes
- 270M – 10.7B
- Hardware
- from: Laptop
Computer visionRUNot maintained2022
Sber AI and SberDevices (ai-forever) · Russia
A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.
- Product search by photo and by Russian description
- Sorting images into categories without labeling
- Checking that a photo matches its caption
- Sizes
- 150M – 430M
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
Microsoft · USA
A time-tested encoder behind many classifiers and NER models (including GLiNER). The multilingual mDeBERTa-v3 understands Russian.
- Classifying review sentiment
- Entity extraction after fine-tuning
- Checking whether a conclusion follows from a text
- Sizes
- 70M – 435M
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
David Dale (cointegrated) · Russia
A very small Russian-English BERT that runs fast on a regular CPU. Ready-made fine-tuned versions exist for sentiment, toxicity and emotions.
- Detecting review sentiment
- Filtering rude chat messages
- Fast classification of requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
Deepfake detectionNot maintained2021
NAVER Clova AI Research and EURECOM · South Korea
The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.
- Voice check during phone authentication
- Filtering replays and synthesis in a voice menu
- A baseline when comparing voice detectors
- Sizes
- weight files of 0.4 and 1.3 MB
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021
Samsung AI Center Moscow (with Skoltech) · Russia
Removes unwanted objects, text and watermarks from photos with clean background fill. Lightweight and fast; still the standard for this task.
- Removing price tags, people and clutter from photos
- Cleaning interior and real estate photos
- Removing text and dates from archival photos
- Sizes
- about 51M
- Hardware
- from: Laptop
Photo editingNot maintained2021
ETH Zurich · Switzerland
A transformer model for upscaling, denoising and removing JPEG artifacts from photos. Lightweight and proven; often embedded in other systems.
- Photo upscaling
- Image denoising
- Removing compression artifacts
- Sizes
- about 12M
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2020–2021
Meta · USA
An early Meta translator that translates directly between 100 languages, without English in the middle. Russian is supported. Old, but light and permissively licensed.
- Translation between any pair of 100 languages
- Quick draft translation on modest hardware
- Base for fine-tuning to your subject area
- Sizes
- 418M – 12B
- Hardware
- from: Laptop
FinanceNot maintained2020
Prosus · Netherlands
A classic model that determines the tone of financial news: positive, negative or neutral. English only, runs fast on a CPU.
- Scoring the tone of company news
- Labeling reports and press releases
- Signals for analytics dashboards
- Sizes
- 110M
- Hardware
- from: Laptop
Deepfake detectionNot maintained2020
MiniVision Technology · China
Practically the only fully open weight set for single-frame face liveness: it tells a live person from a photo, a screen or a mask. It errs in both directions - a person must be able to appeal a rejection.
- Liveness check when signing in by selfie
- Protecting an access system from a photo on a phone
- A check during remote customer identification
- Sizes
- two models of about 1.8 MB each
- Hardware
- from: Laptop