MedicineGGUF2023–2026
FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China
A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Hints for doctors when reviewing images (Vision)
- Sizes
- 7B – 72B
- Hardware
- from: Laptop
Speech to textRUGGUF2025–2026
Microsoft · USA
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
- Transcribing long meetings with speaker labels
- Voicing podcasts and dialogues
- Real-time voice for assistants
- Sizes
- 0.5B – 9B
- Hardware
- from: Laptop
TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextGGUF2025–2026
Xiaomi · China
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
- Logic and calculation tasks
- Agents with tools
- Help for developers
- Sizes
- 7B – 1,02T-A42B
- Hardware
- from: Laptop
TextRU2024–2026
Sber · Russia
Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.
- Russian-language employee assistant on your own server
- Customer replies and request handling in Russian
- Working with contracts and internal policies
- Sizes
- 10B-A1.8B – 702B-A36B
- Hardware
- from: Laptop
TextRU2022–2026
Yandex · Russia
Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.
- Russian-language assistant and chatbot
- Answers based on the company knowledge base
- Base for industry-specific fine-tuning
- Sizes
- 8B – 100B
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
Tabular data2022–2026
Prior Labs (University of Freiburg) · Germany
A ready-made model for tables: it takes example rows and immediately predicts for new ones, without lengthy training or tuning. Only v2 is free for business; newer versions are non-commercial.
- Predicting customer churn from a CRM export
- Scoring applications and leads
- Classifying customers from 1C data
- Sizes
- from a few to hundreds of millions of parameters
- Hardware
- from: Laptop
Tabular data2025–2026
Amazon (AutoGluon team) · USA
Amazon's tabular model built into AutoGluon: classification and regression from examples with brief fine-tuning. Mitra-v2 handles more rows and columns.
- Predicting churn and repeat purchases
- Scoring applications
- Predicting deal or order value
- Sizes
- about 76M
- Hardware
- from: Laptop
Tabular data2025–2026
Layer 6 AI (TD Bank) · Canada
A tabular model from a Canadian bank's AI lab, trained on real tables rather than only synthetic ones. Version 1.2 Turbo made computation orders of magnitude faster.
- Scoring applications and customers
- Predicting churn
- Classifying transactions and customers
- Sizes
- about 60–80M
- Hardware
- from: Laptop
Tabular data2025–2026
Stable AI (Beijing, with Tsinghua University) · China
A table model that alone can classify, predict numbers and fill in missing data. The lightweight LimiX-2M runs on an ordinary computer.
- Filling gaps in 1C and CRM exports
- Churn prediction and scoring
- Classifying customers and products
- Sizes
- 2M – 16M and LimiX-2
- Hardware
- from: Laptop
Faces2021–2026
InsightFace (deepinsight) · China
The most widely used open toolkit for face detection and recognition. Many identity-preserving image generators are built on it. The pretrained weights are non-commercial.
- Detecting and comparing faces in photos
- Face-based access in prototypes
- Face processing as part of other AI systems
- Sizes
- packages from 16 MB to 407 MB
- Hardware
- from: Laptop
Text2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact text models that run directly on a device: laptop, phone or mini PC. The 1B and 2B MiniCPM5 models focus on tool calling and long context.
- A local chat assistant without the cloud
- Data extraction and text classification
- Tool calling and simple agents on low-end hardware
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Text2024–2026
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
- Reasoning, maths and technical questions
- Analysing long documents
- Agents and writing code
- Sizes
- 0.9B – 375B-A23B
- Hardware
- from: Laptop
3D2024–2026
NAVER LABS Europe · France (NAVER, South Korea)
The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.
- 3D scene from several photos without calibration
- Point matching between images
- Mapping from video (SLAM)
- Sizes
- 0.57B – 0.69B
- Hardware
- from: Laptop
Music and soundGGUF2022–2026
m-a-p (Multimodal Art Projection) · UK / China
A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.
- Automatic tagging of a music catalog
- Finding similar tracks
- Detecting genre, mood and tempo
- Sizes
- 95M – 632M
- Hardware
- from: Laptop
Autonomous driving2020–2026
comma.ai · USA
An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.
- A research testbed for driver assistance systems
- Studying driver attention monitoring with an in-cabin camera
- Comparison with your own lane-keeping algorithms
- Sizes
- compact, designed for an in-vehicle device
- Hardware
- from: Laptop
TextRU2026
SberDevices (ai-forever) · Russia
A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.
- Experiments with faster generation
- Fine-tuning small models for your own tasks
- Research
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
Voice assistants2026
Samsung · South Korea
Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.
- Describing an audio recording in words
- Answering questions about a sound
- Identifying the type of sound and the setting
- Sizes
- 99M – 356M
- Hardware
- from: Laptop
Deepfake detection2023–2026
Adobe Research and University of Surrey · USA
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
- Marking images on the way out of your own pipeline
- Checking the provenance of a submitted image
- Linking with content provenance metadata
- Sizes
- model types Q and P with different mark capacity
- Hardware
- from: Laptop
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
Computer vision2024–2026
Microsoft Research · USA
Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.
- Measuring rooms and objects from photos
- 3D point cloud from a single shot
- Preparing data for robots and AR
- Sizes
- ViT-S – ViT-G
- Hardware
- from: Laptop
Search and RAGRU2024–2026
Sber (SberDevices) · Russia
Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.
- Search across Russian-language documents
- RAG for chatbots in Russian
- Classifying requests and reviews
- Sizes
- 480M – 10B-A1.8B
- Hardware
- from: Laptop
Forecasting2024–2026
Google · USA
A ready-made Google forecasting model: forecasts any time series without training on your data.
- Sales and demand forecasting
- Purchase and inventory planning
- Load and traffic forecasting
- Sizes
- 200M – 500M
- Hardware
- from: Laptop
Speech to textGGUF2024–2026
Moonshine AI (Useful Sensors) · USA
Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.
- Voice control of devices
- Offline recognition on a phone
- Live subtitles
- Sizes
- 27M – 245M
- Hardware
- from: Laptop
Speech to textGGUF2025–2026
IBM · USA
IBM speech models for recognizing and translating speech in English, several European languages and Japanese. Designed for enterprise use.
- Transcribing business meetings
- Translating speech into text in another language
- Voice assistants
- Sizes
- 470M – 8B
- Hardware
- from: Laptop
Text to speechGGUF2025–2026
bilibili · China
Speech synthesis with voice cloning and precise duration control, handy for video dubbing. Controls emotion separately from timbre.
- Video dubbing matched to timing
- Voice cloning
- Emotional voiceover
- Sizes
- about 1B – 2B
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
Text2024–2026
Tencent · China
Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.
- Corporate assistant
- Translation and multilingual texts
- Agents with tools
- Sizes
- 0.5B – 770B-A49B
- Hardware
- from: Laptop
Text2025–2026
Ant Group (inclusionAI) · China
An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.
- Corporate assistant
- Agents for office processes
- Financial analytics (Fin version available)
- Sizes
- 7.9B-A1.3B – 1T
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere · Canada
Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.
- Knowledge-base answers with source citations
- Agents that work with internal systems
- Translation and correspondence in different languages
- Sizes
- 2.5B – 218B-A25B
- Hardware
- from: Laptop
TextOllama2024–2026
NVIDIA · USA
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
- Agents with tool calling
- Reasoning and calculation tasks
- Answers based on long documents
- Sizes
- 4B – 550B-A55B
- Hardware
- from: Laptop
TextOllama2024–2026
IBM · USA
IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.
- Answers based on internal documents (RAG)
- Tool calling and agent work
- Data extraction and classification
- Sizes
- 350M – 34B
- Hardware
- from: Laptop
TextRUOllama2025–2026
Liquid AI · USA
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
- Offline assistant on a laptop or phone
- Data extraction from documents
- Tool calling in apps
- Sizes
- 230M – 24B-A2B
- Hardware
- from: Laptop
Text analysis2024–2026
Urchade Zaratiana and Fastino AI · France / USA
Finds the entities you need in text without training: just list what to look for (name, amount, date). GLiNER2 also classifies text. Multilingual versions understand Russian.
- Extracting names, amounts and dates from emails and contracts
- Parsing requests into CRM fields
- Classifying requests by topic
- Sizes
- about 50M to 500M
- Hardware
- from: Laptop
Documents and OCR2026
TeleAI (China Telecom) · China
A new lightweight document parsing model that led the OmniDocBench v1.6 benchmark at release. Handles pages photographed on a phone and crumpled pages well. Languages on the card: Chinese, English, Japanese.
- Recognising invoices and delivery notes photographed on a phone
- Recognising tables and formulas
- Converting documents to Markdown for RAG
- Sizes
- about 1.2B
- Hardware
- from: Laptop
Voice: speakers and sound2022–2026
Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community
A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.
- Clean vocals from a recording with music
- Backing tracks and stems for karaoke
- Removing background music and noise from videos
- Sizes
- from tens to hundreds of millions of parameters
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
Ant Group (inclusionAI) · China
An Ant Group family for finding elements on screen and completing tasks in phone and computer interfaces. UI-Venus-2 was specifically trained to refuse dangerous actions.
- Automating actions in mobile apps
- Filling in forms in web interfaces
- UI autotests
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Deepfake detection2025–2026
University of Michigan · USA
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
- Checking submitted photos and illustrations
- Filtering AI images in a content flow
- Flagging suspicious images for manual review
- Sizes
- 22M
- Hardware
- from: Laptop
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
TextRUOllama2023–2026
Mistral AI · France
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
- Fast chat responses
- Data extraction from text
- Translation and multilingual work
- Sizes
- 3B – 675B
- Hardware
- from: Laptop
Search and RAGRU2025–2026
NVIDIA · USA
NVIDIA embeddings for search and RAG. Nemotron-3-Embed, released in 2026, is under the permissive OpenMDW license and works in many languages.
- Search across corporate documents
- RAG for chatbots and assistants
- Search across images and pages (VL versions)
- Sizes
- 1B – 8B
- Hardware
- from: Laptop
Speech to textRUGGUF2024–2026
Sber · Russia
Sber's models for Russian speech recognition, among the most accurate for Russian. Includes emotion recognition, v3 with punctuation, and a multilingual version (Russian, Kazakh, Kyrgyz, Uzbek).
- Transcribing calls in Russian
- Meeting minutes
- Voice control of services
- Sizes
- 220M – 600M
- Hardware
- from: Laptop
TextGGUF2025–2026
Meituan · China
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
- Agents for orders and service processes
- Corporate assistant
- Analysis of long documents
- Sizes
- 1B – 1.6T-A48B
- Hardware
- from: Laptop
TextOllama2024–2026
LG AI Research · South Korea
Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.
- Corporate assistant
- Working with Korean and English texts
- Analysis of documents and images (4.5)
- Sizes
- 1.2B – 750B-A37B
- Hardware
- from: Laptop
TextGGUF2025–2026
Kakao · South Korea
Compact Korean-English models from Kakao. Kanana 2 30B-A3B is fast thanks to MoE; small 1–3B versions suit a regular PC.
- Support chatbot
- Customer request classification
- Lightweight assistant on your own PC
- Sizes
- 1.3B – 30B-A3B
- Hardware
- from: Laptop
TextRU2024–2026
T-Bank · Russia
T-Bank models fine-tuned from Qwen for Russian: they write and reason in Russian noticeably better than the original. T-Lite is 8B, T-Pro 32B on one GPU; T-Search is a multi-step search agent in Russian and English.
- Russian-language support chatbot
- Analysis of requests and documents in Russian
- Answers based on the company knowledge base
- Sizes
- 7B – 36B-A3B
- Hardware
- from: Laptop
TextGGUF2025–2026
Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland
Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.
- Multilingual assistant
- Answers based on documents
- Analysis of images and scans (v1.5)
- Sizes
- 0.5B – 70B
- Hardware
- from: Laptop
Image + textGGUF2024–2026
Alibaba (AIDC-AI) · China
Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.
- Extracting data from invoices, contracts and delivery notes
- Table recognition
- Answering questions about photos and charts
- Sizes
- 0.9B – 80B-A3B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Tencent · China
A lightweight OCR model from Tencent: document parsing, finding text in photos, field extraction and translating text from images. Version 1.5 is faster and runs on an ordinary PC.
- Extracting fields from invoices and delivery notes
- Recognising tables and formulas
- Translating text in photos and scans
- Sizes
- 1B
- Hardware
- from: Laptop
Voice: speakers and soundGGUF2022–2026
WeNet community · China
A set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.
- Voice verification of a customer during a call
- Finding repeat calls from the same person
- Splitting a recording by speaker
- Sizes
- from a few to tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2022–2026
Meta AI, then Alexandre Défossez · France
A classic model that splits a track into vocals, drums, bass and the rest. The v4 hybrid transformer version remains the benchmark; the project is now maintained by its author in his own repository.
- Separating vocals from music in a recording
- Backing tracks and karaoke stems
- Cleaning speech in videos with background music
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Voice: speakers and sound2023–2026
RVC-Project community · China
The most widely used open voice conversion tool: a model for a specific voice trains on 10–30 minutes of recording and works in real time. Use only with the voice owner's consent.
- Voicing content with one brand voice
- Covers and vocal work
- Real-time voice changing
- Sizes
- tens of millions of parameters
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
Microsoft · USA
Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.
- Filling in web forms and applications
- Collecting data from web portals without an API
- Checking websites against scenarios
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Tabular data2026
LG AI Research · South Korea
LG's small tabular model: with 21M parameters it nearly matches the leaders in classification and regression accuracy. Weights are for non-commercial use only.
- Pilot churn forecasts
- Testing scoring hypotheses
- Exploring customer data
- Sizes
- about 21M
- Hardware
- from: Laptop
Documents and OCRRU2024–2026
Datalab · USA
A compact OCR toolkit from the makers of Marker and Chandra: text recognition, page layout, reading order and tables. Surya OCR 2 (650M) also runs on a CPU; Russian scored 88.8% in benchmarks.
- Recognizing scans and PDFs, including in Russian
- Page layout: headings, tables, images, reading order
- Recognizing tables by rows and columns
- Sizes
- up to 650M
- Hardware
- from: Laptop
RerankersGGUF2024–2026
Jina AI · Germany
Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.
- Refining search results before a chatbot answers
- Sorting retrieved PDF pages and slides
- Catalog and knowledge base search
- Sizes
- 33M – 2.4B
- Hardware
- from: Laptop
Rerankers2026
Tencent · China
A pair of small Tencent models based on Qwen3 that pick the right skill for an AI agent for a given request: the embedding model finds candidates, the reranker chooses the best one.
- Choosing a tool or skill for an AI agent
- Routing requests between bot scenarios
- Search across a catalog of internal tools
- Sizes
- 0.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
CodeGGUF2025–2026
JetBrains · Czech Republic
JetBrains models for fast code autocompletion. Mellum2 (12B, 2.5B active) is already a full assistant: it writes and edits code, calls tools and reasons.
- Fast code autocompletion on your own server
- A developer assistant that does not send code to the cloud
- Fine-tuning on the company's code
- Sizes
- 4B – 12B-A2.5B
- Hardware
- from: Laptop
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Computer vision2025–2026
Roboflow · USA
Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.
- Object detection in video and photos
- Precise outlines of parts and defects
- Fine-tuning for your own object classes
- Sizes
- Nano – 2XL
- Hardware
- from: Laptop
Forecasting2025–2026
NXAI · Austria
A compact forecasting model on the xLSTM architecture, a leader in open benchmarks despite its small size. Runs fast on a regular CPU.
- Demand and sales forecasting
- Energy consumption forecasting
- Forecasts on modest hardware and on site
- Sizes
- about 35M to 82M
- Hardware
- from: Laptop
Speech to textRUGGUF2026
Alibaba (Qwen) · China
Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.
- Transcribing calls and meetings
- Video subtitles
- Multilingual recognition
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
Speech to textGGUF2026
Cohere · Canada
Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.
- Transcribing meetings and interviews
- Subtitles
- Searching an audio archive
- Sizes
- 2B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Zyphra · USA
Speech synthesis with voice cloning and fine control over emotion, speed and pitch.
- Voice cloning
- Emotional voiceover
- Voicing videos
- Sizes
- about 1.6B
- Hardware
- from: Laptop
Music and soundRUGGUF2025–2026
ACE Studio and StepFun · China
Fast generation of songs with vocals in 19 languages, including Russian: a full song in seconds, editing of individual parts and style changes.
- Songs and jingles for ads
- Background music for videos
- Demo versions of tracks
- Sizes
- about 2B – 4B
- Hardware
- from: Laptop
Moderation and safety2024–2026
GLiNER community (Fastino, Knowledgator, NVIDIA and others) · USA
Small GLiNER-based models for finding personal data: passports, phone numbers, accounts, addresses. Data types are set in words. Russian is not officially supported.
- Masking personal data before cloud AI
- Finding passport data and bank details in documents
- Checking data exports for leaks
- Sizes
- about 200M to 500M
- Hardware
- from: Laptop
Image + textOllama2024–2026
Moondream (M87 Labs) · USA
A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.
- Finding and counting objects in photos
- Checking photos from field reports
- Captions and tags for a catalogue
- Sizes
- 2B – 9B-A2B
- Hardware
- from: Laptop
Documents and OCR2026
Baidu · China
Baidu's OCR model building on DeepSeek-OCR ideas: processes multi-page documents and PDFs in a single pass and outputs structured text. Claimed to be multilingual, but the language list is not published.
- Converting multi-page PDFs and scans to text and Markdown
- Recognizing contracts, invoices and reports
- Preparing document archives for search and RAG
- Sizes
- 3.3B
- Hardware
- from: Laptop
Documents and OCRRU2022–2026
Baidu (PaddlePaddle) · China
Classic lightweight PaddleOCR models: detecting and recognizing lines of text plus page layout. They run on CPUs and phones; there is a separate model for East Slavic languages, including Russian.
- Recognizing text on scans, photos and screens
- Reading labels, displays and markings in production and warehouses
- Page layout: tables, formulas, stamps, headings
- Sizes
- from 1.5M to tens of millions of parameters
- Hardware
- from: Laptop
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
Satellite and geo2025–2026
Allen Institute for AI (Ai2) · USA
Ai2's family of models for Sentinel-1, Sentinel-2 and Landsat imagery, with ready-made fine-tunes for mangroves, deforestation and ecosystem types. The license excludes the extractive industries.
- Monitoring deforestation and forest condition across the supply chain
- Classifying land and crops from image series
- Image embeddings for finding similar plots
- Sizes
- Nano – Large (Base about 114M)
- Hardware
- from: Laptop
MedicineOllama2023–2026
EPFL · Switzerland
Open medical models from Swiss EPFL, fine-tuned on clinical guidelines on top of various base models. Does not replace a doctor; decisions are made by a specialist.
- Answering staff questions based on clinical guidelines
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 2B – 70B
- Hardware
- from: Laptop
3D2024–2026
VAST (TripoSR together with Stability AI) · China
VAST family: a 3D model from a single photo. TripoSR runs in under a second, TripoSG gives cleaner geometry, TripoSplat builds a scene from Gaussian points.
- 3D product model from a photo
- Object assets for games and AR
- Quick 3D prototype for printing
- Sizes
- up to 1.5B
- Hardware
- from: Laptop
Search and RAGRUGGUF2023–2026
Jina AI · Germany
Strong multilingual embeddings with long context; v5-omni understands text, images and audio. Recent versions are open for non-commercial use only.
- Search across documents in many languages
- Search across images and scans
- Classification and clustering
- Sizes
- 33M – 3.8B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Cognitive Computations (Eric Hartford) · USA
Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.
- Assistant that does not refuse legal but sensitive topics
- Internal tools under a strict system prompt
- Role-play and creative scenarios
- Sizes
- 0.5B – 405B
- Hardware
- from: Laptop
Speech to textRU2023–2026
NVIDIA · USA
Fast NVIDIA speech recognition models, including streaming ones for real-time use. Parakeet TDT v3 and Nemotron 3.5 ASR understand Russian.
- Transcribing calls and meetings
- Video subtitles
- Real-time voice input
- Sizes
- 110M – 2.5B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Resemble AI · USA
Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- about 350M – 500M
- Hardware
- from: Laptop
Text to speechRU2025–2026
OpenMOSS (Fudan University) · China
A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.
- Voicing podcasts and dialogues
- Voice for an assistant
- Voice cloning
- Sizes
- 100M – 8.5B
- Hardware
- from: Laptop
Music and soundGGUF2024–2026
Stability AI · UK
Generates short music clips and sound effects from a description. Version 3 is split into separate models for music and for sounds.
- Sound effects for videos and games
- Background music and jingles
- Interface sounds
- Sizes
- about 0.5B – 2.3B
- Hardware
- from: Laptop
TranslationRU2025–2026
Tencent · China
Tencent translators for 33 languages; the first version won the WMT25 competition. Russian is supported. The small 1.8B version runs on a laptop; the new Hy-MT2 is under Apache 2.0.
- Translating documents while keeping formatting
- Translation with a set glossary of terms
- Translating correspondence with Chinese partners
- Sizes
- 1.8B – 30B-A3B
- Hardware
- from: Laptop
Moderation and safety2024–2026
NVIDIA · USA
NVIDIA content filters for bots, with separate models for keeping the conversation on topic and detecting jailbreaks. Safety Guard v3 was trained on 9 languages; Russian was tested only without fine-tuning.
- Checking bot requests and replies
- Keeping the bot within its topic
- Detecting attempts to bypass rules
- Sizes
- 4B – 8B
- Hardware
- from: Laptop
Photo editing2023–2026
BRIA AI · Israel
BRIA's background removal, trained on licensed photos. Soft edges, hair, transparency. Video versions available. Business use requires a paid agreement.
- Cutting products out onto a white background
- Staff and expert photos without background
- Background removal in video
- Sizes
- 44M – 220M
- Hardware
- from: Laptop
Image + textOllama2023–2026
LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China
The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.
- Answering questions about photos and screenshots
- Describing products from a photo
- Frame-by-frame video analysis
- Sizes
- 0.5B – 72B
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Image + textGGUF2025–2026
Kuaishou · China
Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.
- Analysing and describing short videos
- Reviewing clips and content
- Finding the right moment in a video
- Sizes
- 8B – 671B-A37B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Baidu (PaddlePaddle) · China
A compact document parsing model from the popular PaddleOCR toolkit. Per the model card it supports 109 languages, including Russian; version 1.6 leads the OmniDocBench benchmark.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables, formulas and stamps
- Converting scans to Markdown and JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
Shanghai AI Laboratory (OpenDataLab) · China
A popular open tool for converting PDFs to Markdown with its own small model. MinerU2.5-Pro was improved through data alone, without growing in size. Languages on the card: Chinese and English.
- Converting PDF reports and contracts to Markdown
- Recognising tables and formulas
- Preparing documents for RAG and search
- Sizes
- 0.9B – 1.2B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
H Company · France
A French model family for controlling a browser and computer: precisely finds the right element on screen and handles multi-step tasks. The latest Holo3 and 3.1 are open under Apache 2.0.
- Working in web portals and legacy software without an API
- Filling in forms and applications
- Testing interfaces against scenarios
- Sizes
- 0.8B – 235B-A22B
- Hardware
- from: Laptop
Text to speechRU2025–2026
Supertone · South Korea
Very fast, lightweight speech synthesis that runs directly on the device, without a GPU or the cloud. Supertonic 3 speaks 31 languages, including Russian.
- Voicing voice bot replies on an ordinary server
- Voiceover in offline and mobile apps
- Reading texts and notifications aloud
- Sizes
- about 99M
- Hardware
- from: Laptop
Computer vision2024–2026
Meta · USA
Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.
- Pose and body keypoint detection
- Segmentation of body parts and clothing
- Separating a person from the background
- Sizes
- 0.1B – 5B
- Hardware
- from: Laptop
Biology and chemistry2022–2026
EvolutionaryScale / Chan Zuckerberg Biohub (ESM-2 — Meta AI) · USA
Protein language models: they understand amino acid sequences, predict structure (ESMFold2) and help with protein design. Since 2026 all open versions are under MIT.
- Protein embeddings for predicting properties (stability, solubility)
- Predicting 3D structures of proteins and complexes
- Screening enzyme and antibody design candidates before lab work
- Sizes
- 8M – 15B (ESM-2), 300M – 6B (ESM C), 1.4B (open ESM3)
- Hardware
- from: Laptop
Rerankers2025–2026
NVIDIA · USA
A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.
- Reordering passages before an AI assistant answers
- Sorting retrieved scan and PDF pages
- Search across internal policies and instructions
- Sizes
- 1B
- Hardware
- from: Laptop
Forecasting2024–2026
THUML, Tsinghua University · China
A compact forecasting foundation model from the Tsinghua lab: trained on a large set of diverse series and fine-tunable on your own data.
- Forecasting demand and load
- Forecasting sensor readings on the shop floor
- Fine-tuning forecasts on your own history
- Sizes
- 84M (timer-base)
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2026
IBM · USA
Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.
- Search across corporate documents
- RAG on a regular server without a GPU
- Reranking results
- Sizes
- 30M – 311M
- Hardware
- from: Laptop
ForecastingGGUF2025–2026
Datadog · USA
A Datadog forecasting model trained on server and application metrics. Especially strong for IT monitoring: load, latency, errors.
- Server load forecasting
- Anomaly detection in metrics
- Capacity planning
- Sizes
- 4M – 2.5B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.
- Voice cloning
- Voicing videos and audiobooks
- Voice for an assistant
- Sizes
- 0.5B – 2.3B
- Hardware
- from: Laptop
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
Arcee AI · USA
An American family of MoE models trained from scratch: Nano, Mini and Large. Trinity-Large-Thinking (398B) reasons before answering.
- Agents with tool calling
- Reasoning tasks
- Corporate assistant on your own servers
- Sizes
- 6B – 398B-A13B
- Hardware
- from: Laptop
TranslationRU2020–2026
Helsinki-NLP, University of Helsinki · Finland
More than a thousand small translators, each for its own language pair. Russian-English and back are available. Fast even on a regular CPU.
- Bulk translation of short texts
- Translation right on the server without a GPU
- Translating reviews and requests before analysis
- Sizes
- 25M – 240M
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2026
IBM · USA
IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.
- Checking bot requests and replies
- Finding made-up facts in knowledge-base answers
- Checking your own rules written as text
- Sizes
- 38M – 8B
- Hardware
- from: Laptop
Moderation and safety2026
OpenAI · USA
Finds and hides personal data: names, addresses, phone numbers, emails, account numbers, passwords. Runs even in the browser. Trained mostly on English.
- Removing personal data from text before sending it to cloud AI
- Finding passwords and keys in texts
- Anonymizing correspondence for analytics
- Sizes
- 1.5B (50M active)
- Hardware
- from: Laptop
Image + textOllama2025–2026
IBM · USA
Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.
- Extracting fields from forms and invoices
- Turning charts and tables into data
- Answering questions about documents
- Sizes
- 2B – 4B
- Hardware
- from: Laptop
Computer-use agents2024–2026
Show Lab (National University of Singapore) · Singapore
A lightweight model for working with interfaces: finds buttons and fields by description and performs actions on the web and on a phone. ShowUI-π can drag with the mouse.
- Clicking and filling in forms from a task description
- Web UI autotests
- An assistant on a low-end computer without the cloud
- Sizes
- 2B (ShowUI), about 500M (ShowUI-π)
- Hardware
- from: Laptop
Text analysisOllama2024–2026
NuMind · France
Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.
- Extracting company details, amounts and dates from invoices and contracts into JSON
- Parsing receipts, waybills and forms against a set template
- Converting scans to Markdown for search
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Text to speech2025–2026
Kyutai · France
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
- Streaming speech transcription for voice bots
- Voicing replies with minimal delay
- Speech synthesis on a server without a GPU (Pocket TTS)
- Sizes
- 100M (Pocket TTS) – 2.6B
- Hardware
- from: Laptop
TextGGUF2025–2026
ServiceNow · USA
ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.
- A reasoning assistant for internal services
- Tool calling and enterprise agents
- Analysing screenshots and documents with images
- Sizes
- 5B – 15B
- Hardware
- from: Laptop
CodeOllama2025–2026
Essential AI · USA
An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.
- Writing and fixing code
- A developer agent on a single GPU
- Solving technical and scientific problems
- Sizes
- 8B
- Hardware
- from: Laptop
Voice assistants2024–2026
NVIDIA · USA
Models that listen to speech, sounds and music and answer questions about them. Audio Flamingo Next handles recordings up to 30 minutes. Research use only.
- Detailed descriptions of audio recordings
- Questions and answers about a long recording
- Tagging music and sounds
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Search and RAGGGUF2025–2026
Octen · USA / Singapore
Qwen3-Embedding models fine-tuned by the startup Octen for search in legal, financial and medical texts. As of January 2026 the 8B version topped the RTEB leaderboard.
- Search across contracts and case law
- Search across financial reports
- Search across long documents up to 32K tokens
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
TextRU2024–2026
Ivan Bondarenko (bond005), Novosibirsk State University · Russia
Russian-language models for working with documents rather than chatting: knowledge-base answers, extraction of entities and facts from Russian text, long context.
- Answers to questions based on internal documents
- Extracting names, dates and amounts from contracts
- Short summaries of long Russian texts
- Sizes
- 1.5B – 7.6B
- Hardware
- from: Laptop
Computer vision2023–2026
Meta · USA
Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.
- Background removal from product photos
- Counting objects in photos
- Data labeling for training
- Sizes
- 91M – ~0,85B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Zhejiang University · China
A medical model for text, images, 3D scans and video: from a light 4B to a large MoE. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images and CT scans
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 4B – 235B-A22B
- Hardware
- from: Laptop
Math and reasoningGGUF2025–2026
Princeton University · USA
Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.
- Formal verification of mathematical workings
- Verifying code correctness
- Training
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Search and RAGGGUF2026
Perplexity · USA
Embeddings from the Perplexity search service. Some versions take into account the context of the whole document, not just a single fragment.
- Search across large document collections
- RAG that accounts for document context
- Website and catalog search
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
Speech to textRUGGUF2025–2026
Mistral AI · France
Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.
- Transcribing and summarizing recordings
- Asking questions about audio
- Real-time recognition
- Sizes
- 3B – 24B
- Hardware
- from: Laptop
Text to speechRU2024–2026
Fish Audio · USA / China
Speech synthesis with voice cloning and emotion control in 80+ languages, including Russian. Quality is close to paid services, but the weights are for research only.
- Voice cloning
- Emotional voiceover
- Multilingual voiceover
- Sizes
- 0.5B – about 4.5B
- Hardware
- from: Laptop
TextGGUF2024–2026
Sarvam AI · India
Indian models focused on 22 languages of India. Sarvam 30B and 105B (2026) are MoE models with strong reasoning and agent skills.
- Multilingual customer support
- Reasoning and calculation tasks
- Agents with tool calling
- Sizes
- 2B – 105B-A10B
- Hardware
- from: Laptop
Text2025–2026
Reka AI · USA
Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.
- Photo and video analysis (Edge)
- Object detection in images
- Reasoning tasks (Flash)
- Sizes
- 7B – 21B
- Hardware
- from: Laptop
Image + textGGUF2023–2026
Shanghai AI Laboratory (OpenGVLab) · China
A large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.
- Understanding documents, diagrams and charts
- Answering questions about photos
- Video analysis
- Sizes
- 1B – 241B-A28B
- Hardware
- from: Laptop
Image + text2024–2026
Ai2 (Allen Institute for AI) · USA
Fully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.
- Counting products and objects in photos
- Pointing to where an item is in an image
- Video analysis
- Sizes
- 1B-A7B – 72B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
rednote hilab (Xiaohongshu) · China
A multilingual document parsing model: text, tables, formulas and reading order in one pass. dots.mocr also turns charts and diagrams into vector SVG.
- Recognising invoices, contracts and delivery notes
- Converting tables into an editable format
- Converting charts and diagrams into vector format
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Datalab · USA
A strong OCR model from the authors of Marker and Surya: handwriting, forms, tables. Per the model card it supports 90+ languages, with Russian among the examples.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising handwritten forms and questionnaires
- Recognising complex tables
- Sizes
- 5B – 9B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2026
Baidu (Qianfan) · China
A Baidu model that not only recognises a document but also answers questions about it. Per the model card it supports 192 languages, including Cyrillic.
- Recognising invoices, contracts and delivery notes, including in Russian
- Page layout analysis and table recognition
- Answering questions about a document
- Sizes
- 4B
- Hardware
- from: Laptop
Voice: speakers and soundRU2026
FireRedTeam (Xiaohongshu) · China
A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.
- Cutting recordings before speech recognition
- Separating speech from music and singing in broadcasts and videos
- Speech detection in voice bots
- Sizes
- compact, exact size not stated
- Hardware
- from: Laptop
Text to speechRU2026
k2-fsa (Next-gen Kaldi) · China
Speech synthesis with voice cloning from a short sample in 646 languages, including Russian and languages of Russia's peoples. A voice can be described in words. Weights are for non-commercial use only.
- Voiceover in rare languages
- Voice cloning from a sample
- Research and prototypes of multilingual voiceover
- Sizes
- 0.6B
- Hardware
- from: Laptop
Photo editing2025–2026
S-Lab, Nanyang Technological University · Singapore
Cuts a person out of video with a precise alpha mask, including hair and edges, without a green screen. Needs a first-frame mask, for example from SAM.
- Background replacement in video without chroma key
- Cutting out a person for editing and effects
- Preparing videos for advertising and social media
- Sizes
- about 35M
- Hardware
- from: Laptop
Search and RAGRU2026
Microsoft · USA
Microsoft's 2026 multilingual embeddings with context up to 32K tokens; Russian is on the language list. The 270M and 0.6B versions run on a regular server, 27B is the most accurate.
- Multilingual knowledge base search
- Picking passages for RAG
- Search across long documents
- Sizes
- 270M – 27B
- Hardware
- from: Laptop
Tabular data2025–2026
Lexsi Labs · India
Recent open models for tabular data: they predict from a few examples given in the prompt, with no task-specific training.
- Classification and forecasting on tables with no separate training
- Quickly testing models on new datasets
- Assessing features in large tables
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
CodeOllama2024–2026
Alibaba (Qwen team) · China
The broadest open coding family: from 0.5B for autocompletion to 480B for agents. Qwen3-Coder-Next (80B, 3B active) works as a developer agent on a single GPU.
- Code autocompletion in the editor
- An agent that edits code in the repository on its own
- Writing and refining scripts, SQL and integrations
- Sizes
- 0.5B – 480B-A35B
- Hardware
- from: Laptop
CodeGGUF2026
IQuest Research · China
A family of coding models with standard and reasoning versions, including a Loop variant that runs through its layers a second time. Sizes from 7B to 40B.
- Writing and refining code
- Solving tasks with step-by-step reasoning
- Agentic work with a repository
- Sizes
- 7B – 40B
- Hardware
- from: Laptop
Code2026
Ai2 (Allen Institute for AI) · USA
Fully open developer agents from Ai2: weights, data and training recipe are all public. Designed so a company can cheaply fine-tune the agent on its own repository.
- An agent for fixing issues in code
- Fine-tuning the agent on an internal repository
- Automating small edits and tests
- Sizes
- 8B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2026
LM Provers (CMU, Hugging Face, ETH Zurich, Project Numina) · USA, Switzerland, France
A small 4B model on Qwen3 that writes mathematical proofs in plain language almost at the level of large models. Runs on a laptop.
- Checking the logic of reasoning and workings
- Step-by-step explanations of solutions
- Training and olympiad preparation
- Sizes
- 4B
- Hardware
- from: Laptop
Voice assistants2024–2026
Kyutai · France
A voice assistant that listens and speaks at the same time, with no delay for recognition and synthesis. Hibiki does simultaneous speech-to-speech translation between several European languages.
- Real-time voice conversation partner
- Simultaneous speech translation
- Zero-latency voice interfaces
- Sizes
- 2B – 7B
- Hardware
- from: Laptop
Voice assistantsGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.
- Voice assistant on your own server
- Analyzing videos and documents
- Voice answers about a camera image
- Sizes
- 8B – 9B
- Hardware
- from: Laptop
Computer-use agents2025–2026
Alibaba (Tongyi Lab, X-PLUG) · China
Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.
- Automating actions in mobile apps
- Working in desktop software without an API
- Testing apps against scenarios
- Sizes
- 2B – 32B
- Hardware
- from: Laptop
Tabular data2025–2026
Inria (Soda team) · France
An open tabular model from the creators of scikit-learn: classifies and predicts from examples without training and handles tables of up to hundreds of thousands of rows. The license allows business use.
- Predicting customer churn
- Scoring applications and deals
- Classifying customers from 1C and CRM data
- Sizes
- about 25–30M
- Hardware
- from: Laptop
MedicineOllama2025–2026
Google · USA
Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.
- Draft discharge summaries and reports for a doctor to review
- Hints for doctors when reviewing images
- Searching and summarising medical literature
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Medicine2023–2026
Stanford AIMI · USA
Stanford models for chest X-rays: they describe the image and prepare a draft report. Does not replace a doctor; decisions are made by a specialist.
- A draft X-ray description for the radiologist
- Hints for doctors when reviewing images
- Checking reports for completeness
- Sizes
- 3B – 8B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Alibaba DAMO Academy · China
Alibaba's medical model based on Qwen2.5-VL: understands many types of medical images and medical text, and can reason step by step. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer vision2023–2026
Ultralytics · USA
The most widely used real-time object detector: finds and marks items in video even on modest hardware. YOLOv5 came out back in 2020; the catalog starts from YOLOv8.
- Counting people, cars and goods on video
- Checking hard hats and workwear
- Spotting defects on the production line
- Sizes
- 2.4M – 68M
- Hardware
- from: Laptop
Search and RAGRUOllama2025–2026
Alibaba (Qwen) · China
Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.
- Knowledge base search for RAG
- Reranking results before answering
- Search across scans, slides and screenshots
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Search and RAG2026
Voyage AI (MongoDB) · USA
The only open model in the Voyage 4 line: its vectors are compatible with the paid larger versions, so you can start locally and move to the API later.
- Document search on your own server
- RAG for small knowledge bases
- Finding similar texts
- Sizes
- about 340M
- Hardware
- from: Laptop
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
TextGGUF2023–2026
Baichuan Intelligence · China
First general-purpose Chinese models, then a medical line from 2025. Baichuan-M3 (based on Qwen3-235B) is trained to model a doctor's clinical reasoning.
- Reference assistant for doctors
- Preliminary patient intake questions
- Analysis of medical documents
- Sizes
- 7B – 235B-A22B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Technology Innovation Institute (TII) · UAE
A family from Abu Dhabi: from the early Falcon 40B and 180B to hybrid Falcon-H1 and tiny Falcon-H1-Tiny models of 90–600M parameters for devices.
- Assistant and answers based on documents
- Running on low-end hardware and devices
- Tool calling in simple agents
- Sizes
- 90M – 180B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Allen Institute for AI (Ai2) · USA
Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.
- Assistant and answers based on documents
- Reasoning tasks (Think versions)
- Fine-tuning on your data with a clear model history
- Sizes
- 1B – 32B
- Hardware
- from: Laptop
TextGGUF2024–2026
AI21 Labs · Israel
A hybrid of Transformer and Mamba with a window of up to 256K tokens: handles long documents faster than conventional models. Jamba2 focuses on accurate, source-based answers.
- Answers based on long policies and contracts
- Knowledge-base search (RAG)
- Summaries of large documents
- Sizes
- 3B – 398B-A94B
- Hardware
- from: Laptop
TextRUGGUF2024–2026
UTTER consortium (Unbabel, universities of Lisbon, Edinburgh, Amsterdam and others) · European Union
European language models trained on all EU languages and several others, with a focus on translation. Russian is supported. Permissive license.
- Translation and localization of texts
- Answering questions in different languages
- Draft emails for foreign partners
- Sizes
- 1.7B – 22B
- Hardware
- from: Laptop
TranslationRUOllama2026
Google · USA
Translators based on Gemma 3 for 55 languages that can also translate text in images. Russian is supported. The 4B version fits on a laptop.
- Translating documents and correspondence
- Translating text from screenshots and photos
- Localizing websites and apps
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Documents and OCROllama2025–2026
DeepSeek · China
An OCR model that compresses a page into a small number of visual tokens, so it processes large volumes quickly. Version 2 better understands reading order.
- Bulk recognition of scanned invoices and contracts
- Table recognition
- Converting PDFs to Markdown for search and RAG
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUOllama2026
Zhipu AI (Z.ai) · China
A lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables and formulas
- Extracting fields to JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
LightOn · France
A French 1B OCR model that converts a page into text in one pass and is fast on high volumes. Languages on the card: European languages, Chinese and Japanese; no Russian.
- Recognising invoices and contracts in European languages
- Table recognition
- Converting PDFs to text for search and RAG
- Sizes
- 0.9B – 1B
- Hardware
- from: Laptop
Cybersecurity2025–2026
Cisco (Foundation AI) · USA
Cisco models for information security based on Llama 3.1 8B: analysis of vulnerabilities, threats and incidents. Can be deployed inside your own perimeter.
- Analyzing vulnerability and threat reports
- Helping SOC analysts during incidents
- Mapping threats to MITRE ATT&CK
- Sizes
- 8B
- Hardware
- from: Laptop
Voice: speakers and soundRU2025–2026
Daily (Pipecat) · USA
Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.
- Voice bot does not interrupt the customer during pauses
- Fast reply when the customer has really finished
- An add-on to a standard speech detector in voice assistants
- Sizes
- 8M (v3) – 580M (v1)
- Hardware
- from: Laptop
Computer visionGGUF2024–2025
ByteDance and the University of Hong Kong · China
Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.
- Estimating distances and volumes from a camera
- Depth effects for photo and video
- Navigation for robots and drones
- Sizes
- 25M – 1.4B
- Hardware
- from: Laptop
Speech to textRU2025
Meta · USA
Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.
- Transcription in rare and local languages
- Digitizing oral archives
- Subtitles in many languages
- Sizes
- 300M – 7B
- Hardware
- from: Laptop
Speech to text2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Alibaba's set of fast speech recognition models, primarily for Chinese and Asian languages. SenseVoice also detects emotions and sound events.
- Transcribing calls
- Detecting emotions in the voice
- Recognizing laughter, music and other sounds
- Sizes
- about 230M – 800M
- Hardware
- from: Laptop
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
Text2025
Naver · South Korea
Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.
- Lightweight Korean-English assistant
- Analysis of images and documents
- Text classification
- Sizes
- 0.5B – 32B
- Hardware
- from: Laptop
TextRUGGUF2024–2025
Vikhr Models · Russia
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
- Russian-language assistant on your own PC or server
- Knowledge-base answers (RAG)
- Texts and emails in Russian
- Sizes
- 0.5B – 24B
- Hardware
- from: Laptop
Image + text2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
Vision models from Zhipu: first CogVLM, then the GLM-V line. GLM-4.6V can call tools based on images and act as an agent operating an interface.
- Answering questions about photos and documents
- An agent that operates an interface from screenshots
- Analysing charts and reports
- Sizes
- 9B – 106B-A12B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
Alibaba (Tongyi-MAI) · China
Compact Alibaba models for working in smartphone and computer interfaces: they find elements and complete multi-step tasks. The small size allows running on an ordinary GPU.
- Automating actions in mobile apps
- Working in software without an API
- UI autotests
- Sizes
- 2B – 8B
- Hardware
- from: Laptop
Moderation and safety2025
ServiceNow · USA
A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.
- Screening chatbot requests for attacks and jailbreaks
- Filtering harmful model answers
- Monitoring the actions of AI agents that use tools
- Sizes
- 8B
- Hardware
- from: Laptop
TextGGUF2023–2025
Inception (G42), MBZUAI and Cerebras · UAE
A model family for Arabic and English, including Gulf dialects. Suits companies working with Arabic-speaking customers and government bodies in the region.
- A chatbot in Arabic and English
- Translating and summarising documents in Arabic
- Classifying customer requests
- Sizes
- 256M – 70B
- Hardware
- from: Laptop
Computer vision2025
Meta · USA
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
- Search photos and videos by description
- Catalog labeling and tagging
- Search across audio and video (PE-AV)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
- Marking video created or processed by AI
- Finding your own mark in re-uploaded clips
- Protecting ad materials from being reused as someone else's
- Sizes
- a mark of 96 to 1024 bits
- Hardware
- from: Laptop
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek's maths models. The first 7B version introduced the GRPO training method; the 685B V2 writes and checks its own olympiad-level proofs.
- Calculations and formula checks
- Checking mathematical workings in reports
- Working through problems step by step
- Sizes
- 7B – 685B
- Hardware
- from: Laptop
TextOllama2023–2025
Nous Research · USA
Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.
- Agents that call functions and APIs
- Data extraction in strict JSON format
- Assistant with flexible role and tone settings
- Sizes
- 3B – 405B
- Hardware
- from: Laptop
Text to speechRU2022–2025
Silero · Russia
Lightweight Russian speech synthesis that runs on a regular CPU. Version v5 added CIS languages and languages of Russia's peoples: Tatar, Bashkir, Yakut, Kazakh and others.
- Voicing voice bot replies
- Reading texts in Russian
- Voices in the languages of Russia's peoples
- Sizes
- tens of megabytes
- Hardware
- from: Laptop
Text to speech2025
Nari Labs · South Korea
A model that voices entire two-person dialogues with laughter, sighs and pauses. English only.
- Voicing dialogues and podcasts
- Ads with natural speech
- Training role-plays
- Sizes
- 1B – 2B
- Hardware
- from: Laptop
Voice: speakers and soundRU2020–2025
Silero · Russia
The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.
- Cutting calls and recordings before speech recognition
- Detecting when the customer is speaking in a voice bot
- Filtering out silence and noise to save on transcription
- Sizes
- about 2 MB
- Hardware
- from: Laptop
TextOllama2025
Deep Cogito · USA
Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.
- A chat assistant with a reasoning mode
- Writing code and calling tools
- Answering complex questions about documents
- Sizes
- 3B – 671B
- Hardware
- from: Laptop
Satellite and geo2025
IBM and the European Space Agency (ESA) · USA / Europe
A multimodal Earth model: understands optical and radar imagery, terrain, vegetation index and land use maps, and can generate a missing data type (for example, a "see-through-clouds" image from radar).
- Analyzing fields and forests even in cloudy weather using radar imagery
- Land use maps for assessing plots
- Flood and wildfire assessment (ready-made fine-tunes available)
- Sizes
- tiny – large (checkpoints from ~200 MB to ~3.8 GB)
- Hardware
- from: Laptop
Computer vision2023–2025
Meta · USA
Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.
- Image search by text
- Image classification without training
- Search research and prototypes
- Sizes
- 0.15B – 3.6B
- Hardware
- from: Laptop
Rerankers2025
ZeroEntropy · USA
Rerankers built on Qwen3. The model card lists the target domains — finance, law, code, medicine, science; the stated language is English.
- Refining results before an AI assistant answers
- Sorting search results across contracts and reports
- Search across technical and scientific documentation
- Sizes
- zerank-2 — 4B (based on Qwen3-4B), plus a smaller "small" version
- Hardware
- from: Laptop
Computer vision2023–2025
IDEA Research · China
Finds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.
- Finding objects by description without labeling
- Automatic data labeling for training
- Checking photos against requirements
- Sizes
- 172M – 3B
- Hardware
- from: Laptop
ForecastingGGUF2024–2025
Amazon · USA
Amazon forecasting models, among the most downloaded. Chronos-2 takes external factors into account: prices, promotions, weather.
- Demand forecasting with promotions and prices
- Inventory planning
- Forecasting revenue and customer flow
- Sizes
- 8M – 710M
- Hardware
- from: Laptop
Music and sound2025
ASLP-lab (Northwestern Polytechnical University) · China
Fast generation of a full song with vocals from lyrics and a style sample, up to several minutes long.
- Songs and jingles from lyrics
- Music for videos
- Demo versions of tracks
- Sizes
- about 1.1B
- Hardware
- from: Laptop
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025
Nanonets · USA / India
A model that converts documents to Markdown with tables, stamps, signatures, checkboxes and watermarks. The OCR2 model card lists Russian among its languages.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising stamps, signatures and marks
- Handwriting recognition
- Sizes
- 1.5B – 3B
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
NVIDIA · USA
NVIDIA models for "who is speaking": TitaNet recognizes a specific person's voice, Sortformer splits a recording into up to 4 speakers, including live during a call.
- Real-time speaker tagging in conversations
- Checking that the same person is calling (voiceprint)
- Preparing meeting transcripts
- Sizes
- 23M (TitaNet) – 117M (Sortformer)
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Mixedbread · Germany
Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.
- Search across a knowledge base
- Reranking results before a bot answers
- Product catalog search
- Sizes
- 17M – 1.5B
- Hardware
- from: Laptop
Visual document search2025
Illuin Technology, EPFL, CentraleSupélec · France
A compact (250M) model for searching document pages as images. According to the authors, it matches models 10 times larger and runs without a GPU.
- Search across scans and PDFs on a modest server
- Indexing document archives
- Search across slides and manuals
- Sizes
- 250M
- Hardware
- from: Laptop
TextRU2025
Avito Tech · Russia
Avito's model based on Qwen3-8B, retrained for Russian: its own tokenizer makes Russian text 15–25% faster. Supports function calling.
- Product and listing descriptions in Russian
- Chatbot that calls internal services
- Request analysis and classification
- Sizes
- 7.9B
- Hardware
- from: Laptop
Deepfake detection2025
National Institute of Informatics, Yamagishi Lab · Japan
Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.
- Checking audio recordings for synthesis
- Fine-tuning for your own language and recording channel
- Comparing several encoders on your own data
- Sizes
- 0,3B – 2B
- Hardware
- from: Laptop
Search and RAGOllama2025
Google · USA
A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.
- On-device document search
- RAG without sending data outside
- Text classification
- Sizes
- 300M
- Hardware
- from: Laptop
Speech to textRU2023–2025
Alpha Cephei · Russia
Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.
- Transcribing Russian calls and recordings without the cloud
- Voice control in apps and kiosks
- Low-latency streaming speech recognition
- Sizes
- about 45 MB – 1.8 GB
- Hardware
- from: Laptop
Voice assistantsRUGGUF2025
Alibaba (Qwen) · China
Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.
- Voice assistant for customers
- Analyzing calls and videos
- Voice answers about documents and images
- Sizes
- 3B – 30B-A3B
- Hardware
- from: Laptop
Moderation and safetyRUGGUF2025
Alibaba (Qwen) · China
Safety filters for 119 languages, Russian among them. The Stream version checks a bot's reply while it is being generated and can cut it off on the fly.
- Filtering bot requests in Russian
- Stopping a dangerous reply during generation
- Labeling messages by risk category
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Documents and OCR2025
IBM and Hugging Face · USA
Tiny models for the open Docling document converter: they turn a page into markup with tables, formulas and code. Run on an ordinary laptop.
- Converting PDFs and scans to Markdown for search and RAG
- Recognising tables in reports
- Processing invoices and contracts on an ordinary PC
- Sizes
- 256M – 258M
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
pyannoteAI (Hervé Bredin) · France
The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.
- Tagging calls: which part is the agent, which is the customer
- Meeting minutes with speaker labels
- Preparing recordings for transcription and analysis
- Sizes
- a few million parameters
- Hardware
- from: Laptop
Satellite and geo2023–2025
IBM and NASA · USA
Foundation models for Landsat and Sentinel-2 satellite imagery that account for image time series. Ready-made fine-tunes for floods, burn scars and crop types, plus a separate weather model, WxC.
- Mapping crops and field condition over the season
- Assessing flood zones and burn scars after natural disasters
- Monitoring changes in buildings and land use
- Sizes
- tiny – 600M (imagery), 2.3B (Prithvi WxC weather)
- Hardware
- from: Laptop
Search and RAGOllama2023–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.
- Search across English-language documents
- Picking passages for chatbot answers (RAG)
- Code search (bge-code)
- Sizes
- 24M – 9B
- Hardware
- from: Laptop
Text to SQLGGUF2024–2025
Prem AI · UK
A text-to-SQL model of just 1B parameters, designed to run locally so the database never leaves for external services.
- Local translation of questions into SQL with no internet access
- Query hints on modest hardware
- Embedding into internal analytics tools
- Sizes
- 1B
- Hardware
- from: Laptop
Deepfake detection2023–2025
IBM Research and The Chinese University of Hong Kong · USA
An AI-text detector trained together with a paraphraser: it was deliberately taught not to give up when the text has been rewritten. It errs in both directions; a human reviews the output.
- Checking texts that may have been rewritten after generation
- First-pass filtering in a newsroom or admissions office
- Comparison against simpler detectors
- Sizes
- about 355M (RoBERTa-large)
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Moonshot AI and Project Numina · China, France
Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.
- Formal verification of mathematical workings
- Translating a problem from plain language into Lean
- Training and olympiad preparation
- Sizes
- 0.6B – 72B
- Hardware
- from: Laptop
Computer vision2023–2025
Meta · USA
Foundation models that turn an image into a numeric "fingerprint". They are used to build similar-image search, classification and segmentation without large labeled datasets.
- Finding similar products and photos
- Image classification on small datasets
- Base for your own quality-control models
- Sizes
- 21M – 7B
- Hardware
- from: Laptop
ForecastingGGUF2024–2025
Salesforce · USA
Salesforce's universal forecasting model for series with different frequencies and many variables. Weights are open for research only.
- Research forecasting pilots
- Comparison with other forecasting models
- Forecasts across many related series
- Sizes
- 11M – 311M
- Hardware
- from: Laptop
TextRU2024–2025
Lomonosov Moscow State University Research Computing Center, LAIR lab (RefalMachine) · Russia
Qwen models adapted for Russian: a new tokenizer plus further training on Russian texts. As a result, Russian text is generated up to twice as fast as with the original model of the same size.
- Russian-language assistant on your own server
- Answers based on company documents (RAG) in Russian
- Analysis and summaries of long Russian texts
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
Alibaba (Tongyi Lab) · China
Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.
- Noise suppression in conversation recordings
- Separating two voices speaking at once
- Improving old and phone recordings
- Sizes
- under 1B
- Hardware
- from: Laptop
Faces2025
ByteDance · China
Transformer-based face recognition from ByteDance, one of the most accurate open models on benchmarks. Weights are published in ONNX format but are non-commercial.
- Comparing faces and searching a photo database
- Research on recognition accuracy
- Access control prototypes
- Sizes
- ViT-T – ViT-L
- Hardware
- from: Laptop
Text to speechRU2025
ESpeech (independent group of Russian-speaking developers) · Russia
Russian speech synthesis with voice cloning based on the F5-TTS architecture, trained on Russian speech datasets collected by the authors. Stress is placed automatically. Several variants, including a "podcaster" one.
- Voicing videos and audiobooks in Russian
- Cloning a narrator's voice from a sample
- Voice for a bot or assistant in Russian
- Sizes
- about 340M
- Hardware
- from: Laptop
Deepfake detection2023–2025
The Chinese University of Hong Kong, Shenzhen (SCLBD) · China
Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.
- First-pass check of a submitted video or selfie
- Comparing several detectors on your own data
- Fine-tuning a detector for your own flow of applications
- Sizes
- Xception- and EfficientNet-class detectors, tens of millions of parameters
- Hardware
- from: Laptop
Medicine2025
Google · USA
Google's lightweight encoder for medical images and text, the same one inside MedGemma. Sorts images and finds similar ones. Does not replace a doctor; decisions are made by a specialist.
- Finding similar images in a clinic's archive
- Pre-sorting images for a doctor
- A base for your own image classifiers
- Sizes
- about 0.9B
- Hardware
- from: Laptop
MedicineGGUF2025
Intelligent Internet · UK
Reasoning medical models on Qwen3, designed to run on an ordinary computer. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff with the reasoning shown
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Speech to textRU2025
T-Bank · Russia
A compact T-Bank streaming model for recognizing Russian speech in phone calls. Works in real time even without a GPU.
- Transcribing phone calls
- Voice robots on the line
- Call quality control
- Sizes
- 72M
- Hardware
- from: Laptop
TextRUOllama2024–2025
Hugging Face · USA
Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.
- Simple on-device assistant
- Classification and routing of requests
- Base for fine-tuning on a narrow task
- Sizes
- 135M – 3B
- Hardware
- from: Laptop
TranslationRUGGUF2025
ByteDance Seed · China
A compact ByteDance translator for 28 languages, close in quality to large closed systems. Russian is supported. Ready-made compressed versions are available.
- Translating business correspondence and documents
- Translating product cards
- Translating technical and legal texts
- Sizes
- 7B
- Hardware
- from: Laptop
Photo editingGGUF2024–2025
Nankai University · China
An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.
- Bulk background removal from product photos
- Precise masks for design and print
- Cutting out people with hair for advertising
- Sizes
- about 220M (lightweight lite versions available)
- Hardware
- from: Laptop
Finance2025
Tsinghua University (NeoQuasar) · China
A foundation model for market candlestick data: trained on data from more than 45 exchanges, it forecasts prices and volumes. The largest version, large, is not open.
- Forecasting candlesticks and trading volumes
- Volatility estimation
- A base for fine-tuning on your own series
- Sizes
- 4.1M – 102M
- Hardware
- from: Laptop
Voice: speakers and sound2025
Agora (TEN project) · USA / China
A lightweight speech detector for real-time voice assistants: it notices the start and end of a phrase faster than Silero VAD. Runs on servers, phones and in the browser.
- Zero-lag speech detection in a voice bot
- Fast assistant response at the end of a phrase
- Use in mobile apps and the browser
- Sizes
- very small, the library is smaller than Silero VAD
- Hardware
- from: Laptop
Math and reasoningOllama2025
Agentica (Berkeley, Sky Computing Lab) and Together AI · USA
Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.
- Solving maths problems with step-by-step working
- Generating and checking code
- An agent for fixing bugs in a repository
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Visual document search2024–2025
Illuin Technology (ViDoRe team) · France
Searches PDFs and scans as images: pages do not need to be OCR'd first, the model finds the right one for a question directly, including tables and charts. Trained on English.
- Search across scans, presentations and PDFs
- RAG over documents with tables and charts
- Search across technical documentation
- Sizes
- 256M – 3B
- Hardware
- from: Laptop
Voice: speakers and soundRU2023–2025
Community (xbgoose and others), Dusha dataset from SberDevices · Russia
Models that detect emotion from voice in Russian speech: neutral, anger, positive, sadness. Trained on the open Dusha dataset from SberDevices.
- Finding calls with irritated customers
- Assessing the tone of operator conversations
- Prioritizing complaints in a call center
- Sizes
- 21M – 316M
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
- Marking generated and edited images
- Finding a marked fragment inside a collage
- Tracking which parts of a picture were made by AI
- Sizes
- a mark encoder and decoder for images
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
OpenCompass (Shanghai AI Laboratory) · China
A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring model answers against set criteria
- Checking an answer against a reference solution
- Comparing several models on your own data
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoning2025
NVIDIA · USA
NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Working through programming problems
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Robotics2025
Hugging Face · USA
A small robot control model that runs on a regular laptop. Trained on open data from the LeRobot community, suited to low-cost robot arms.
- Controlling a low-cost robot arm
- Quick robotization pilots and demos
- Training staff and students
- Sizes
- 450M
- Hardware
- from: Laptop
TranslationRUGGUF2024–2025
Unbabel · Portugal
Language models tailored for translation and multilingual text work: translating, editing, and assessing translation quality. Russian is supported. Non-commercial license.
- Translation that respects context and terminology
- Post-editing machine translation
- Assessing the quality of a finished translation
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
RVC-Boss and community · China
Speech synthesis with voice cloning: a 5-second sample is enough, and after fine-tuning on a minute of recording the voice sounds noticeably more accurate. Use only with the voice owner's consent.
- Voicing texts with a specific narrator's voice
- Voice for a bot or assistant
- Dubbing training videos
- Sizes
- under 1B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
Skywork (Kunlun Tech) · China
Reward models: they score how good a language model's answer is for the user. Used for fine-tuning your own models and picking the best of several answers.
- Choosing the best of several bot answers
- Scoring answer quality during model fine-tuning
- Comparing models before rollout
- Sizes
- 0.6B – 27B
- Hardware
- from: Laptop
Avatars2025
ByteDance · China
Matches lip movements in an existing video to a new voice track. Version 1.6 works at 512 pixels and produces a sharper face.
- Dubbing videos into another language with lip sync
- Editing lines in finished video without reshooting
- Talking avatars for training courses
- Sizes
- requires 8–18 GB of VRAM
- Hardware
- from: Laptop
Satellite and geo2025
MBZUAI · UAE
A compact research model for satellite imagery, trained on both optical (Sentinel-2) and radar (Sentinel-1) data. Narrower in scope and community than Prithvi and TerraMind.
- Classification and segmentation of satellite imagery after fine-tuning
- A base for a land monitoring prototype
- Sizes
- TerraFM-B (ViT-Base)
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Clouditera · China
A Chinese open family for cybersecurity: reviewing vulnerabilities, analysing logs and traffic, explaining commands and scripts.
- Reviewing vulnerabilities and drafting fix recommendations
- Analysing logs and reconstructing an attack chain
- Explaining suspicious commands and scripts
- Sizes
- 1.5B – 14B
- Hardware
- from: Laptop
Text to SQLGGUF2025
IDEA Research · China
A text-to-SQL model trained with reinforcement learning: it works through the schema and the conditions step by step before producing a query.
- Database queries for questions with several conditions
- Reviewing and fixing other people SQL queries
- An analyst helper inside a BI system
- Sizes
- 3B – 14B
- Hardware
- from: Laptop
Search and RAGGGUF2024–2025
TechWolf · Belgium
A model from a Belgian HR company: it turns job titles into vectors so you can find similar vacancies and resumes. A human makes the decision about a candidate; automatic screening without review must not be used.
- Matching job titles coming from different sources
- Finding similar vacancies and resumes by meaning
- Cleaning up the company job title reference list
- Sizes
- 109M – 278M
- Hardware
- from: Laptop
TextOllama2025
DeepSeek · China
A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.
- Complex calculations and logic checks
- Analysis of contracts and internal policies
- Help for developers
- Sizes
- 1,5B – 671B
- Hardware
- from: Laptop
CodeGGUF2025
ByteDance Seed · China
A compact 8B coding model from ByteDance in base, instruct and reasoning versions. Its training data was selected by the model itself, with almost no hand-written rules.
- Code autocompletion and generation
- Solving algorithmic problems
- A base for fine-tuning on your own stack
- Sizes
- 8B
- Hardware
- from: Laptop
Text analysisRU2023–2025
deepvk (VK) · Russia
Russian encoders from the VK team: RuModernBERT reads long texts, USER produces vectors for search, GeRaCl classifies texts by topic without training.
- Classifying requests without labeled data
- Knowledge base search in Russian
- Analyzing long contracts
- Sizes
- 35M – 360M
- Hardware
- from: Laptop
Text to SQL2025
Snowflake · USA
A Snowflake model for turning questions into SQL, trained with reinforcement learning by checking query results. The open 7B version is based on Qwen2.5-Coder.
- Plain-language questions to a data warehouse
- Generating SQL for reports and dashboards
- Checking and fixing analysts' queries
- Sizes
- 7B
- Hardware
- from: Laptop
CodeRU2025
MTS AI (MWS AI) · Russia
A small coding assistant from MTS AI that understands requests in Russian. Runs locally, with plugins for VS Code and JetBrains.
- Code suggestions and completion in the editor
- Code explanations in Russian
- Drafts of tests and documentation
- Sizes
- 1.5B
- Hardware
- from: Laptop
Rerankers2023–2025
Stanford NLP, later Answer.AI and LightOn · USA and France
A different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.
- Search across a knowledge base of long documents
- Reordering retrieved passages
- Search across policies and technical documentation
- Sizes
- about 33M – 150M
- Hardware
- from: Laptop
Forecasting2025
THUML, Tsinghua University · China
A forecasting model that returns a set of possible scenarios rather than a single line — useful when you need a range for demand or load, not one number.
- Forecasting demand with a range of values
- Planning stock while accounting for spread
- Forecasting load on services and staff
- Sizes
- 128M (sundial-base)
- Hardware
- from: Laptop
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek models for formal proofs in Lean 4: the proof is checked by a program, not a person. A narrow tool for mathematicians and engineers.
- Formal verification of mathematical workings
- Verifying algorithm correctness
- Training and olympiad preparation
- Sizes
- 7B – 671B
- Hardware
- from: Laptop
TextRUGGUF2023–2025
Ilya Gusev (IlyaGusev) · Russia
The best-known Russian community fine-tune: open models (Llama, Mistral, Gemma, YandexGPT) trained to act as a Russian-speaking assistant. A convenient starting point for a Russian chatbot on your own server.
- Russian-language chat assistant
- Answers based on the company knowledge base
- Drafts of emails, descriptions and posts in Russian
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Moderation and safetyOllama2023–2025
Meta · USA
Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.
- Checking user questions to the bot
- Checking bot replies before sending
- Reporting which rule category was violated
- Sizes
- 1B – 12B
- Hardware
- from: Laptop
Moderation and safety2024–2025
Meta · USA
Tiny classifiers that catch attempts to hack a bot: prompt injections and rule bypassing. The 86M version is multilingual, 22M is English only.
- Protecting a bot from prompt injections
- Checking emails and documents that reach an AI agent
- Fast filter in front of a large model
- Sizes
- 22M – 86M
- Hardware
- from: Laptop
Virtual try-on2024–2025
Sun Yat-sen University and Pixocial · China
A lightweight try-on model that runs on a regular GPU. There is a mask-free version and CatV2TON, which also tries clothes on in video.
- Trying a garment on a customer's photo
- Draft product cards on a model
- Try-on in a short video
- Sizes
- 899M
- Hardware
- from: Laptop
Image + textGGUF2024–2025
Hugging Face · France / USA
The smallest vision models from Hugging Face, starting at 256M; they run in a browser and on a phone. SmolVLM2 also understands video.
- Describing photos and video on low-end hardware
- Reading simple documents
- Embedding in mobile and offline apps
- Sizes
- 256M – 2.2B
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
Songting Liu (Plachtaa) · China
Voice conversion without training: transfers timbre from a 1–30 second sample, can sing and work in real time; V2 also changes accent. Use only with the voice owner's consent.
- Re-voicing a video with a different voice
- Voice anonymization in recordings
- Real-time voice for streams
- Sizes
- about 70M – 200M
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
ByteDance Seed · China
A model that looks at a screenshot and controls the mouse and keyboard itself: clicks, fills in fields, navigates menus. The first generation and 1.5-7B are open; UI-TARS-2 weights were not released.
- Working in legacy software without an API
- Filling in forms and moving data between systems
- UI autotests from plain-language scenarios
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Text to SQL2025
Alibaba · China
Alibaba models for turning questions into SQL, based on Qwen2.5-Coder. They work with different SQL dialects; a small 3B version suits modest hardware.
- Plain-language database questions
- Queries for different databases (PostgreSQL, MySQL, SQLite)
- Automating routine reports
- Sizes
- 3B – 32B
- Hardware
- from: Laptop
Moderation and safety2023–2025
Falconsai and Freepik · USA and Spain
Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.
- Filtering user photos and avatars
- Checking generated images before publishing
- Labeling a media library
- Sizes
- 86M
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Kindo · USA
One of the best-known open families for security and DevSecOps work: reviewing code for weaknesses, test scenarios, explaining attacks.
- Finding weak spots in code and configurations
- Reviewing incidents and explaining attack techniques
- Drafting scripts and procedures for the security team
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Speech to textRUGGUF2022–2025
OpenAI · USA
Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.
- Transcription of calls and video meetings
- Video subtitles
- Voice messages to text
- Sizes
- 39M – 1,5B
- Hardware
- from: Laptop
Image generationGGUF2024–2025
NVIDIA · USA
NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.
- Bulk image generation
- High-resolution visuals
- Real-time generation inside apps
- Sizes
- 0.6B – 4.8B
- Hardware
- from: Laptop
Avatars2024–2025
Tencent Music (Lyra Lab) · China
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
- Dubbing video into another language
- Live avatar in a video chat
- Editing speech in a finished video
- Sizes
- under 1B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Stanford University · USA
A reasoning model trained on just a thousand problems. It can be told to think longer to answer a hard question more accurately.
- Calculations and formula checks
- Working through complex problems step by step
- Training
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Qihoo 360 · China
Reasoning models from Qihoo 360: a standard Qwen2.5 was fine-tuned for long reasoning using an open recipe; data and code are published.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Nomic AI · USA
Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.
- Search across documents and a knowledge base
- Code search
- Search across scans and PDFs without text recognition
- Sizes
- 137M – 7B
- Hardware
- from: Laptop
Text to speechRUGGUF2023–2025
Rhasspy / Open Home Foundation · USA
Very fast speech synthesis that runs even on a Raspberry Pi. Ready-made voices in 35+ languages, including several Russian ones.
- Voicing notifications and bot replies
- Voice for offline devices
- Voice menus
- Sizes
- about 5M – 30M
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
Shanghai Jiao Tong University and partners · China
A voice cloning model that needs only a few seconds of a sample, in English and Chinese. The community has released many fine-tuned versions for other languages, including Russian.
- Voice cloning
- Voicing audiobooks and videos
- Research and prototypes
- Sizes
- about 340M
- Hardware
- from: Laptop
Text to speechGGUF2025
Canopy Labs · USA
Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.
- Real-time voice for an assistant
- Emotional voiceover
- Voice cloning
- Sizes
- 3B
- Hardware
- from: Laptop
Text to speechGGUF2025
Sesame · USA
A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.
- Voice for a conversational assistant
- Voicing dialogues
- Voice product prototypes
- Sizes
- 1B
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2025
Google · USA
Gemma-based filters: they check text for dangerous and offensive content, and ShieldGemma 2 checks images. Focused on English.
- Moderating user messages
- Checking bot replies
- Checking generated images before publishing
- Sizes
- 2B – 27B
- Hardware
- from: Laptop
Text to SQL2025
Renmin University of China (RUC) · China
Models for turning questions into SQL, trained on millions of synthetic query examples across different databases. Three sizes for different hardware.
- Database questions without knowing SQL
- Generating queries for reports
- A base for fine-tuning on your own database schema
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer-use agents2024–2025
Salesforce · USA
Salesforce models for function calling and agents: they pick the right tool and fill in its parameters. Strong on benchmarks, but the license is non-commercial.
- Calling APIs and internal services on user request
- Multi-step agents with several tools
- Comparing approaches before choosing a commercial model
- Sizes
- 1B – 8x22B
- Hardware
- from: Laptop
Finance2025
Shanghai University of Finance and Economics (SUFE) · China
A reasoning model for financial tasks based on Qwen2.5-7B: calculations, report analysis, regulatory questions. Trained on Chinese and English data.
- Financial calculations with step-by-step explanations
- Answering questions about financial statements
- Analyzing tables of financial data
- Sizes
- 7B
- Hardware
- from: Laptop
Deepfake detection2025
Desklib · India
A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.
- Checking submitted articles and reports
- Filtering templated reviews and applications
- First-pass check of student work
- Sizes
- 0.4B (DeBERTa-v3-large)
- Hardware
- from: Laptop
Math and reasoningGGUF2025
NovaSky (Sky Computing Lab, Berkeley) · USA
A Berkeley reasoning model trained for under 450 dollars. It showed that o1-preview-level reasoning can be reproduced with modest resources.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer vision2023–2025
Google · USA
Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.
- Image search by text query
- Automatic catalog labeling and tagging
- Filtering prohibited content
- Sizes
- about 0.2B to 2B
- Hardware
- from: Laptop
TextOllama2023–2025
Ai2 · USA
Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.
- Employee assistant on your own server
- Math and precise instruction following
- Reference recipe for your own fine-tuning
- Sizes
- 7B – 405B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
hexgrad (independent developer) · not disclosed
A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.
- Voicing articles and notifications
- Voice for apps without a GPU
- Bulk text voiceover
- Sizes
- 82M
- Hardware
- from: Laptop
Computer-use agents2024–2025
Microsoft · USA
Breaks a screenshot down into buttons, fields and icons with labels so a regular language model can understand and control the screen. It does not click itself; it serves as the agent's eyes.
- Mapping legacy software screens for automation
- Preparing an agent to work in an interface
- Checking that the required elements are on screen
- Sizes
- under 1B (detector + captioning)
- Hardware
- from: Laptop
Cybersecurity2025
Trend Micro · Japan
Trend Micro cybersecurity models based on Llama 3.1 8B, fine-tuned on a corpus of security texts. A reasoning version is available.
- Answering questions about threats and vulnerabilities
- Analyzing cyberattack reports
- A base for fine-tuning for SOC tasks
- Sizes
- 8B
- Hardware
- from: Laptop
TextGGUF2025
HUMAIN (formerly SDAIA) · Saudi Arabia
A Saudi model for Arabic and English, trained from scratch. One 7B version is openly available.
- An Arabic-language assistant
- Answering questions about documents
- Writing and editing texts in Arabic
- Sizes
- 7B
- Hardware
- from: Laptop
Medicine2024–2025
Bioptimus · France
Pathology foundation models from France's Bioptimus with 1.1B parameters, plus the compact H0-mini. The first version is open under Apache 2.0. Does not replace a doctor; decisions are made by a specialist.
- Histology slide patch features for research models
- A prototype for sorting slides by tissue type
- Research projects linking morphology and molecular data
- Sizes
- 86M (H0-mini) – 1,1B
- Hardware
- from: Laptop
Search and RAG2025
The authors of the CareerBERT paper, German universities · Germany
A German-language model that matches a resume to occupations from the European ESCO reference list and suggests suitable directions. A human makes the decision about a candidate; automatic screening without review must not be used.
- Suggesting occupations that fit the candidate experience
- Matching resumes against job descriptions
- Hints on internal career moves
- Sizes
- 110M
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.
- Calculations and formula checks
- Checking calculations in estimates and reports
- Working through problems step by step
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Search and RAGRU2023–2025
Alibaba · China
Alibaba embeddings and rerankers for search: from tiny to 7B based on Qwen2. There is a multilingual mGTE version with long context.
- Semantic search across documents
- Reranking search results
- Clustering and classifying texts
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Text analysisGGUF2024–2025
Answer.AI and LightOn · USA / France
A modern replacement for classic BERT: faster, reads up to 8 thousand tokens at once. A base for your own classifiers. Trained on English and code; for Russian there is RuModernBERT.
- Classifying requests and documents
- Finding relevant passages in long texts
- Base for your own classifier after fine-tuning
- Sizes
- 150M – 395M
- Hardware
- from: Laptop
Photo editing2024–2025
Prama LLC · USA
A background removal model focused on difficult edges: hair, fur, fine details. The open version is MIT-licensed and can process video.
- Cutting out products and people from photos
- Background removal in video
- Preparing photos for a catalog
- Sizes
- about 95M
- Hardware
- from: Laptop
Image + textGGUF2024–2025
DeepSeek · China
A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.
- Answering questions about images
- Draft illustrations from a description
- Experiments with a unified vision and generation model
- Sizes
- 1B – 7B
- Hardware
- from: Laptop
Computer vision2022–2025
University of Sydney and JD Explore Academy · Australia / China
A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.
- Body keypoints in photos and video
- Motion analysis in sports and rehabilitation
- Monitoring work postures and safety practices
- Sizes
- 33M – about 1B
- Hardware
- from: Laptop
Image + text2023–2025
Alibaba DAMO Academy · China
Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
- Video description and short summary
- Finding a moment in a recording by question
- Tagging a video archive
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Medicine2024–2025
Mahmood Lab (Mass General Brigham, Harvard) · USA
A foundation model for histology slides: turns patches of digital slides into features for tissue classification. Does not replace a doctor; decisions are made by a specialist.
- Research classifiers of tissue types from digital slides
- Finding similar cases in a slide archive
- Preparing features for research prognosis models
- Sizes
- about 300M (UNI) – about 680M (UNI2-h)
- Hardware
- from: Laptop
Search and RAGRUOllama2019–2025
UKP Lab (TU Darmstadt), later Hugging Face · Germany
The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.
- Search across a knowledge base and FAQ
- Finding similar tickets and duplicates
- Grouping reviews and requests by topic
- Sizes
- about 20M – 470M
- Hardware
- from: Laptop
Text analysisRUOllama2024–2025
Jina AI · Germany
Small models that turn raw web page HTML into clean Markdown or JSON. Handy for preparing websites for a knowledge base. Non-commercial license only.
- Cleaning website pages for a knowledge base
- Extracting data from pages into JSON
- Preparing texts for RAG
- Sizes
- 0.5B – 1.5B
- Hardware
- from: Laptop
Visual document search2025
LlamaIndex · USA
A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.
- Search across scans and PDFs without OCR
- Search across invoices, acts and contracts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Fact-checking and judges2025
Atla · UK
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring chatbot answers against your own criteria
- Comparing two versions of a prompt or model
- Filtering out weak answers before they reach a person
- Sizes
- 8B
- Hardware
- from: 1 GPU
Music and sound2024
University of Illinois and Sony AI · USA / Japan
Adds sound to silent video: generates noises and sound effects in sync with the on-screen action, from the video and a text prompt. One of the first strong open Foley models.
- Sound effects for silent video
- Sound for clips from AI generators
- Draft sound design for editing
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Music and sound2024
Tencent AI Lab · China
The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.
- Searching music by text description
- Tagging tracks by genre and mood
- Finding similar music
- Sizes
- 300M – 700M
- Hardware
- from: Laptop
Medicine2024
Mahmood Lab (Mass General Brigham, Harvard) · USA
Image-plus-text models for pathology: search slides by an English description, classify without fine-tuning; TITAN describes a whole slide. Does not replace a doctor; decisions are made by a specialist.
- Text-query search across a slide archive for research
- Draft slide descriptions for research projects
- Tissue classification without labels at the start of a study
- Sizes
- about 160M – 300M
- Hardware
- from: Laptop
Search and RAGRUOllama2024
Snowflake · USA
Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.
- Search across documents and knowledge bases
- Picking passages for RAG
- Search across reports and internal data
- Sizes
- 22M – 568M
- Hardware
- from: Laptop
Deepfake detection2024
Meta · USA
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
- Marking speech synthesized by your service
- Finding your own mark in third-party publications
- Checking whether synthesis was mixed into a call recording
- Sizes
- a watermark generator and detector, 16-bit message
- Hardware
- from: Laptop
Search and RAG2024
TechWolf · Belgium
Finds mentions of skills in a vacancy or resume text and maps them to the company skill reference list. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting skills from a job description
- Matching candidate skills against requirements
- Building a competence map across departments
- Sizes
- 109M
- Hardware
- from: Laptop
Visual document search2024
Alibaba (Tongyi Lab) · China
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
- Finding a product by photo
- Search across a catalogue of images and cards
- Search across document pages as images
- Sizes
- 2B and 7B
- Hardware
- from: 1 GPU
Fact-checking and judges2024
Patronus AI · USA
A small judge: it scores against your criteria and highlights which part of the answer led to that score. The license is non-commercial. The judge itself makes mistakes and does not replace manual review.
- Scoring answers against your criteria with an explanation
- Understanding why a score was lowered
- Bulk review of assistant conversations
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Text to speechGGUF2024
Hugging Face · USA
Speech synthesis where the voice is set by a text description ("a calm female voice, clean recording"). English and 8 European languages, no Russian.
- Choosing a voice by description
- Voicing videos
- Voice service prototypes
- Sizes
- 880M – 2.2B
- Hardware
- from: Laptop
Music and sound2023–2024
Meta · USA
Generates instrumental music from a text description or a sample melody. One of the first open models of its kind.
- Draft music sketches
- Music for video prototypes
- Research
- Sizes
- 300M – 3.3B
- Hardware
- from: Laptop
Image + textGGUF2024
Google · USA
Google's vision model built on Gemma, designed as a base for fine-tuning on a narrow task: captions, object detection, reading text.
- Fine-tuning for your own recognition task
- Finding objects in photos
- Reading text in images
- Sizes
- 3B – 28B
- Hardware
- from: Laptop
Documents and OCR2024
StepFun · China
One of the first general-purpose new-generation OCR models: text, formulas, tables, sheet music and diagrams. Small and runs on low-end hardware, but already behind newer models.
- Recognising scanned invoices and contracts
- Converting tables into an editable format
- Recognising formulas and diagrams
- Sizes
- 580M
- Hardware
- from: Laptop
CodeOllama2024
INF Technology · China
Fully reproducible coding models: along with the weights, the data, its cleaning pipeline and the training recipe are open. Understand English and Chinese.
- Code generation and completion
- Training your own coding model from an open recipe
- A programming assistant on low-end hardware
- Sizes
- 1.5B – 8B
- Hardware
- from: Laptop
TextRU2024
MTS AI (MWS AI) · Russia
A lightweight Russian-language model from MTS AI for Russian texts: answers, summaries, drafts. A ready version for CPU without a GPU is available. The larger Cotype Pro is not released openly.
- Drafts of emails and descriptions in Russian
- Short document summaries
- Answers to common customer questions
- Sizes
- 1.5B
- Hardware
- from: Laptop
TranslationRU2024
deepvk (VK) · Russia
Compact Kazakh-Russian translators from VK. At 197M they translate as well as the 600M NLLB and run on a regular CPU.
- Translating requests from Kazakh to Russian
- Translating documents and instructions into Kazakh
- Bilingual customer support
- Sizes
- 197M
- Hardware
- from: Laptop
Finance2023–2024
The Fin AI / ChanceFocus · international project
One of the first open model families for financial text: reading statements, news and questions about numbers. Not investment advice: decisions are made by a specialist.
- Reading financial statements and press releases
- Answering questions about numeric data in documents
- Classifying financial texts
- Sizes
- 0.5B – 30B
- Hardware
- from: Laptop
Image generationGGUF2022–2024
Stability AI · UK
The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.
- Illustrations and banners for advertising
- Backgrounds and scenes for product cards
- Fine-tuning to a brand style
- Sizes
- 0.9B – 8B
- Hardware
- from: Laptop
Video2022–2024
hzwer (Zhewei Huang) and co-authors · China
Generates intermediate frames: turns 24–30 fps into 60 fps and more and makes smooth slow motion. Versions 4.24+ smooth out video from generative models well.
- Increasing video frame rate
- Smooth slow-motion video
- Smoothing clips from AI generators
- Sizes
- lightweight model (size not stated on the model card)
- Hardware
- from: Laptop
Cybersecurity2024
cybersectony · not disclosed
A very light classifier for emails and links showing signs of phishing. It errs in both directions, so borderline emails are still reviewed by a person.
- Flagging suspicious incoming emails
- Checking links from correspondence before opening them
- A first-level filter in a mail gateway
- Sizes
- about 66M
- Hardware
- from: Laptop
Visual document search2024
LightOn · France
A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.
- Refining search results over scans and PDFs
- Selecting pages before an AI assistant answers
- Sorting retrieved slides and reports
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Forecasting2024
Auton Lab, Carnegie Mellon University · USA
A foundation model for numeric series: one engine is used for forecasting, anomaly detection, filling gaps and classification.
- Forecasting demand and load
- Detecting anomalies in sensor readings and metrics
- Filling gaps in historical data
- Sizes
- about 40M – 385M
- Hardware
- from: Laptop
Forecasting2024
IBM Research · USA
Tiny forecasting models from IBM: they run on an ordinary CPU and sit next to the business system without a separate GPU server.
- Forecasting sales and warehouse stock
- Forecasting energy use and equipment load
- Fast forecasts right on the company server
- Sizes
- very small: TinyTimeMixers have about 1M parameters
- Hardware
- from: Laptop
CodeOllama2023–2024
DeepSeek · China
DeepSeek's coding model family: from small autocompletion models to the large MoE V2, which matched closed models in 2024. Later, coding moved into DeepSeek's general models.
- Code autocompletion and generation
- Translating code between programming languages
- Finding bugs and explaining other people's code
- Sizes
- 1.3B – 236B-A21B
- Hardware
- from: Laptop
Moderation and safety2024
iiiorg · not disclosed
A popular detector of 17 types of personal data in six European languages. No Russian and a non-commercial license: suitable for trials and research.
- Finding personal data in texts
- Comparing the quality of PII detectors
- Sizes
- 278M
- Hardware
- from: Laptop
CodeOllama2024
01.AI · China
Coding models from 01.AI at 1.5B and 9B with a 128K-token context and support for 52 programming languages. A separate line next to the text Yi models.
- Code autocompletion and generation
- Explaining and refactoring code
- A programming assistant without the cloud
- Sizes
- 1.5B – 9B
- Hardware
- from: Laptop
Tabular dataGGUF2024
RUCKBReasoning, Renmin University of China · China
A model for office work with tables: for a given question it returns either a direct answer or code to process the data in a table or document.
- Processing tables from Excel and documents from a text instruction
- Generating code for recalculations and selections
- Answering questions about data in reports
- Sizes
- 7B и 13B
- Hardware
- from: Laptop
Visual document search2024
University of Waterloo, Tevatron project · Canada
Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.
- Search across scans and PDFs without OCR
- Search across presentations and reports with complex layouts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Fact-checking and judges2024
Flow AI · not disclosed
A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.
- Checking AI assistant answers against the instruction
- Bulk scoring of exported conversations
- Quality control before rolling out changes
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Forecasting2024
The Time-MoE team · not disclosed
A forecasting model with a sparse architecture: only part of the network runs at each step, so it stays fast at a small size.
- Forecasting sales and stock levels
- Forecasting load on services and staff
- Planning purchases from history
- Sizes
- 50M and 200M
- Hardware
- from: Laptop
Fact-checking and judgesOllamaNot maintained2024
UT Austin and Bespoke Labs · USA
Checks whether each claim in an AI answer is supported by the source documents. The small versions are free; the larger 7B is in Ollama but non-commercial.
- Checking RAG bot answers against documents
- Finding unsupported claims in reports and summaries
- Automated quality control of AI answers
- Sizes
- 0.4B – 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Stability AI · UK
Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.
- A lightweight chatbot on an ordinary PC
- Code autocompletion in the editor
- A base for fine-tuning on your own task
- Sizes
- 1.6B – 12B
- Hardware
- from: Laptop
MedicineNot maintained2024
Paige · USA
Paige's pathology foundation model, trained on millions of digital slides. The first version is Apache 2.0, the second is for research only. Does not replace a doctor; decisions are made by a specialist.
- Slide patch features for research classifiers
- Selecting slides for re-review in research projects
- Comparison with other pathology models on your own archive
- Sizes
- about 632M
- Hardware
- from: Laptop
CodeOllamaNot maintained2024
Mistral AI · France
Mistral's coding model covering 80+ programming languages. The open weights of the main version cannot be used in production without a paid license; newer Codestral versions are API-only.
- Evaluation and testing before buying a license
- Code autocompletion (with a commercial license)
- Research on coding model quality
- Sizes
- 7B – 22B
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Lvmin Zhang (Stanford) and the community · USA
An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.
- Image from a sketch or outline
- Keeping pose and composition
- Interior visualization from a floor plan
- Sizes
- 0.4B – 1.3B
- Hardware
- from: Laptop
VideoNot maintained2023–2024
Shanghai AI Lab and CUHK · China
A module that brings Stable Diffusion image models to life, turning them into short animations. One of the first open video technologies.
- Short animations in brand style
- Animated covers and banners
- Animated stickers
- Sizes
- motion module on top of SD 1.5 / SDXL
- Hardware
- from: Laptop
AvatarsNot maintained2024
Kuaishou (Kling) · China
Animates a portrait from a reference video: an actor's facial expressions and head turns are transferred to the photo. Runs fast even on a weak GPU.
- Animating portraits
- Transferring an actor's expressions to a character
- Mascot animation
- Sizes
- under 1B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2023–2024
M42 Health · UAE
Clinical models from Abu Dhabi-based M42, built on Llama and tuned to answer medical questions. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff on clinical questions
- Draft discharge summaries and letters for a doctor to review
- Searching medical literature
- Sizes
- 8B – 70B
- Hardware
- from: Laptop
Text analysisRUNot maintained2020–2024
SberDevices (ai-forever) · Russia
Sber's Russian-language encoders trained on large Russian corpora. A base for classifiers, NER and semantic search in Russian.
- Classifying requests in Russian
- Extracting names, amounts and dates after fine-tuning
- Detecting review sentiment
- Sizes
- about 30M to 430M
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
Vectara · USA
A small model that checks whether an AI answer is grounded in the source text or made up. Runs on a CPU and works well as a filter in RAG systems.
- Checking knowledge base chatbot answers for fabrications
- Quality control of document summaries
- Comparing language models by their tendency to make errors
- Sizes
- 110M
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
Zhipu AI (Z.ai) and Tsinghua University · China
Coding models from the creators of GLM. CodeGeeX4-ALL-9B, based on GLM-4-9B, combines autocompletion, code chat, function calling and repository search in one model.
- Code autocompletion in the IDE
- A code chat assistant
- Answering questions about a repository
- Sizes
- 6B – 9B
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
BAAI (Beijing Academy of Artificial Intelligence) · China
Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.
- Refining search results before a chatbot answers
- Sorting knowledge base search results
- Selecting the most relevant clauses of contracts and policies
- Sizes
- 278M – 9B
- Hardware
- from: Laptop
Image + textNot maintained2024
Microsoft · USA
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Sizes
- 0.23B – 0.77B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
OpenChat (Tsinghua University) · China
Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.
- Chat assistant on an inexpensive server
- Drafts of emails and replies
- Help with simple code
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
01.AI · China
Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.
- Chat assistant on a single GPU
- Analysis of long documents
- Classification and data extraction from text
- Sizes
- 6B – 34B
- Hardware
- from: Laptop
Photo editingNot maintained2023–2024
Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China
All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.
- Removing and replacing objects in photos
- Extending the frame to a required format
- Inserting a product or detail from a text description
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Photo editingNot maintained2024
Lvmin Zhang (author of ControlNet) · USA
Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.
- Matching product lighting to a new background
- Studio lighting for portraits without a reshoot
- Consistent lighting style across a catalog
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2023–2024
Defog · USA
One of the first open models that turn a plain-language question into an SQL query against a database. Available in Ollama, but newer competitors are already stronger.
- Answering managers' questions from the sales database without an analyst
- Drafting SQL queries for reports
- An assistant inside a BI system
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
TextNot maintained2024
Equall · France
Language models for legal texts, fine-tuned on US and European legal corpora (based on Mistral and Mixtral). English only.
- Reviewing English-language contracts
- Spotting risks and non-standard terms
- Drafting legal memos
- Sizes
- 7B – 141B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2024
UC Santa Barbara and co-authors · USA
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
- Checking long articles and reports as a whole
- Filtering machine text in a publication flow
- Comparing detectors on your own data
- Sizes
- about 150M (Longformer-base)
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
RLHFlow · USA
An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data for model fine-tuning
- Scoring assistant answers across several attributes
- Sizes
- 8B
- Hardware
- from: 1 GPU
CodeOllamaNot maintained2023–2024
BigCode (Hugging Face and ServiceNow) · USA / France
One of the first open coding models, trained on an open set of source code with an option to exclude your own repository. Today it is more a base for fine-tuning than a leader.
- Code autocompletion in the editor
- Fine-tuning on the company's internal code
- Generating boilerplate code and tests
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Huawei Noah's Ark Lab and partners · China
A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.
- Illustrations for articles and social media
- Backgrounds for product cards
- Quick visual drafts
- Sizes
- 0.6B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2024
Avignon University and Nantes University · France
Mistral 7B fine-tuned on PubMed Central papers, plus several merges with the general model. Compact and easy to run. Does not replace a doctor; decisions are made by a specialist.
- Searching and summarising medical papers
- Draft reference materials for staff
- Explaining medical terminology
- Sizes
- 7B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2024
Saama AI Research · India
Llama 3 fine-tuned on medical and biological data. One of the first strong open medical models of 2024. Does not replace a doctor; decisions are made by a specialist.
- Extracting data from medical documents
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 8B, 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Hugging Face (H4) · USA
Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.
- Lightweight chat assistant
- Reference and starting point for your own fine-tuning
- Drafts of texts and replies
- Sizes
- 7B – 141B-A35B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
KAIST and LG AI Research (prometheus-eval) · South Korea
An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.
- Scoring chatbot answers on your own scale
- Comparing two answer options
- Quality checks before launching an AI service
- Sizes
- 7B – 8x7B
- Hardware
- from: Laptop
FacesNot maintained2023–2024
Tencent AI Lab (h94) · China
One of the first adapters that transfer a face from a photo into a generated image. The SD 1.5 versions run on low-end cards, but the weights are non-commercial.
- Portraits from a photo in different styles
- Image series with one character
- Avatar experiments
- Sizes
- adapters for SD 1.5 and SDXL
- Hardware
- from: Laptop
Text to speechNot maintained2024
MyShell and MIT · USA
Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.
- Voicing bot replies in foreign languages
- Voicing training materials
- Reading texts aloud on a server without a GPU
- Sizes
- small, runs in real time on a CPU
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2024
Chat2DB · China
A text-to-SQL model from the open Chat2DB database client: it supports different SQL dialects, with an English and Chinese model card.
- Turning a question into SQL inside a database client
- Drafting queries for different database engines
- Hints for developers working with a schema
- Sizes
- 7B
- Hardware
- from: Laptop
Text analysisNot maintained2024
University of Southern Denmark · Denmark
Turns a free-form occupation description into a standard HISCO code in 13 languages. Built for historical archives, but also useful for cleaning up job title reference lists. A human makes the decision about a candidate; automatic screening without review must not be used.
- Mapping mixed occupation names onto a single code
- Processing archives of HR and statistical data
- Preparing data for reporting
- Sizes
- based on CANINE-s, size not stated on the model card
- Hardware
- from: Laptop
TextRUNot maintained2023–2024
Sber (ai-forever) · Russia
Sber's Russian text-to-text model, successor to ruT5 (2021). Small and fast: fine-tuned for summarizing, paraphrasing and fixing errors in Russian text; ready-made SAGE spell-checking versions exist.
- Fixing spelling mistakes and typos in Russian text
- Short summaries and paraphrasing
- Normalizing requests and inquiries before processing
- Sizes
- 95M – 1.7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
CybersecurityGGUFNot maintained2023–2024
ZySec AI · India
A small open assistant for security professionals: questions about standards, reviewing threats and vulnerabilities, drafting internal documents.
- Answering questions about security policies and standards
- First-pass review of threat reports
- Drafting internal protection guidelines
- Sizes
- 2.8B и 7B
- Hardware
- from: Laptop
Search and RAGRUNot maintained2022–2024
Microsoft · USA
Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.
- Search across a knowledge base and documents
- Finding answers for a chatbot (RAG)
- Finding similar requests and duplicates
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
ForecastingNot maintained2024
ServiceNow, Mila and partners · Canada
One of the first open out-of-the-box forecasting models. Tiny, gives a probabilistic forecast, now behind Chronos and TimesFM.
- Probabilistic sales forecast
- Quick forecasting pilots
- Baseline model for comparison
- Sizes
- 2.4M
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2024
ChatDB · USA
A text-to-SQL model built on DeepSeek-Coder, aimed at complex questions spanning several tables and conditions.
- Complex queries joining several tables
- Answering database questions without an analyst
- Drafting SQL for reports and exports
- Sizes
- 7B
- Hardware
- from: Laptop
Documents and OCRNot maintained2024
Microsoft · USA
One model for every document task: reading, answering questions about a page, extracting fields, classification. In HR it is used to parse resumes and attached scans. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from a resume and its attachments
- Answering questions about document content
- Classifying incoming documents
- Sizes
- 742M
- Hardware
- from: Laptop
Search and RAGRUOllamaNot maintained2024
BAAI · China
A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.
- Search across a document base
- RAG for a chatbot
- Finding similar requests and duplicates
- Sizes
- 568M
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
Meta · USA
A version of Llama 2 further trained on code, with variants for Python and for chat. Outdated, but many ready-made fine-tuned versions and tools exist.
- Code autocompletion and explanation
- Generating Python scripts
- Base model for fine-tuning on your own stack
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
WizardLM (Microsoft and Peking University) · USA / China
Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.
- Complex multi-step instructions
- Help for developers
- Solving math problems
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2024
MotherDuck and Numbers Station · USA
A model for turning questions into SQL, built for the embedded analytics database DuckDB. Available in Ollama, convenient for working with CSV and Parquet locally.
- Plain-language questions about CSV and Parquet exports
- DuckDB queries inside analytics scripts
- Quick analytics on a laptop without a server
- Sizes
- 7B
- Hardware
- from: Laptop
Satellite and geoNot maintained2023–2024
Allen Institute for AI (Ai2) · USA
Pretrained models from the Satlas project for Sentinel-2, Landsat and high-resolution aerial imagery. The predecessor of OlmoEarth, still used in TorchGeo.
- Detecting objects in imagery: solar farms, wind turbines, ships
- Mapping roads and buildings from aerial photos
- A starting point for fine-tuning your own geo model
- Sizes
- Swin-v2 and ResNet backbones (Base)
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
University of Waterloo, Castorini group · Canada
Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.
- Reordering a long list of search results
- Selecting sources for an AI assistant answer
- Research comparisons of retrieval approaches
- Sizes
- 7B – 13B
- Hardware
- from: 1 GPU
Photo editingNot maintained2023
Alibaba DAMO Academy · China
Colorizes black-and-white photos in natural colors. A lightweight model with a commercial-friendly license; a compact tiny version is available.
- Colorizing archival photos
- Color versions of historical photos for publications
- Family photo restoration service
- Sizes
- DDColor-T (tiny) and DDColor-L
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2023
Resemble AI · USA
A speech enhancement model: removes noise and restores lost frequencies so a muffled recording sounds studio-quality. Good for preparing a voice for voiceover.
- Restoring old and phone recordings
- Cleaning a voice before voiceover and cloning
- Improving audio in videos and podcasts
- Sizes
- under 1B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Intel · USA
A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.
- A simple chat assistant
- Experiments with running on Intel hardware
- A base for fine-tuning
- Sizes
- 7B
- Hardware
- from: Laptop
FinanceGGUFNot maintained2023
AdaptLLM · not disclosed
Finance-tuned versions of Llama 2: reading industry texts, reports and questions about terminology. Not investment advice: decisions are made by a specialist.
- Reading financial news and reports
- Answering questions about financial terminology
- A base for fine-tuning to your own financial task
- Sizes
- 7B и 13B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023
Hello-SimpleAI · China
One of the first open AI-text classifiers, trained on the HC3 corpus of paired human and ChatGPT answers. It errs in both directions: its output is a reason to talk to the author, not proof.
- First-pass check of student work
- Filtering templated applications and reviews
- Flagging suspicious texts for manual review
- Sizes
- about 125M (RoBERTa-base)
- Hardware
- from: Laptop
RerankersNot maintained2023
NetEase Youdao · China
An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.
- Search across a knowledge base and reference materials
- Reordering retrieved passages
- Picking answers for a support chatbot
- Sizes
- about 280M
- Hardware
- from: Laptop
Text to speechNot maintained2023
Columbia University · USA
A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.
- Voicing texts in English
- A base for fine-tuning your own voice
- Voice service prototypes
- Sizes
- about 150M
- Hardware
- from: Laptop
Voice assistantsRUNot maintained2023
Meta · USA
Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.
- Speech-to-speech translation
- Translating and transcribing recordings
- Streaming translation
- Sizes
- 281M – 2.3B
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2023
Google · USA
Google's translator for more than 400 languages under a permissive license. Russian is supported. A good substitute for NLLB when commercial use is needed.
- Translating documents and emails
- Translating catalogs and product descriptions
- Translating into CIS and Asian languages
- Sizes
- 3B – 10B
- Hardware
- from: Laptop
Documents and OCRNot maintained2022–2023
Microsoft · USA
Small models that find tables on PDF and scanned pages and restore their structure: rows, columns, headers. The text inside is read by a separate OCR.
- Finding tables in reports, statements and invoices
- Restoring rows and columns for export to Excel
- Preparing tabular data for analysis and RAG
- Sizes
- 29M
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Microsoft Research · USA
Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.
- Research on reasoning methods
- Comparison with modern small models
- Training specialists
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
Tabular dataNot maintained2023
OSU NLP Group, Ohio State University · USA
A general-purpose model for tables: filling gaps, finding rows, matching columns and answering questions about the data.
- Answering questions about tables inside documents
- Matching columns across different tables
- Finding and completing records in reference books
- Sizes
- 7B
- Hardware
- from: Laptop
Text analysisNot maintained2022–2023
Mike Zhang, Rob van der Goot, Barbara Plank (IT University of Copenhagen and LMU Munich) · Denmark
A research line of models for labour market texts: trained on job postings and the European ESCO occupation taxonomy, they pull skills and requirements out of vacancies. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting skills and requirements from vacancy text
- Mapping skills to the single ESCO reference list
- Classifying vacancies and job titles
- Sizes
- 110M – 560M
- Hardware
- from: Laptop
Text to speechRUNot maintained2023
Coqui · Germany
A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.
- Voice cloning from a sample
- Multilingual voiceover
- Research and prototypes
- Sizes
- about 470M
- Hardware
- from: Laptop
Music and soundNot maintained2023
LAION · Germany
CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.
- Search sounds and music by description
- Automatic tags for an audio library
- Recognizing sound types (siren, breaking glass, voice)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Text to SQLNot maintained2023
RUCKBReasoning, Renmin University of China · China
An early line of open text-to-SQL models starting at 1B, including variants fine-tuned for specific database schemas.
- Turning an employee question into an SQL query
- Drafting warehouse queries for a report
- Embedding into a BI dashboard as a helper
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
Taehoon Kim (POSTECH) · South Korea
A salient object detection model and the ready-made transparent-background tool built on it: removes backgrounds from photos, video and webcam with one command.
- Batch background removal from photos
- Replacing the background with a color or blur
- Background removal in video
- Sizes
- small (based on Swin-B)
- Hardware
- from: Laptop
Documents and OCRNot maintained2023
Meta · USA
An early model that converts scientific PDFs into text with formulas. Now outdated and outperformed by almost all modern OCR models.
- Converting scientific papers from PDF into text with formulas
- Digitising technical documentation
- Sizes
- 250M – 350M
- Hardware
- from: Laptop
Deepfake detectionNot maintained2022–2023
EURECOM · France
A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.
- Spotting synthetic speech in calls
- Checking voice messages and recordings
- Fine-tuning for your own data and codecs
- Sizes
- about 0.3B (wav2vec 2.0 XLS-R encoder)
- Hardware
- from: Laptop
TextRUGGUFNot maintained2022–2023
Sber (ai-forever) · Russia
Sber's multilingual model covering 61 languages, including languages of the peoples of Russia and the CIS. Separate fine-tunes exist for Buryat, Yakut, Tatar, Bashkir, Kazakh and others, rare for open models.
- Texts in languages of the peoples of Russia and the CIS
- Base for fine-tuning on a less common language
- Drafts and templates in several languages
- Sizes
- 1.3B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
LMSYS (Berkeley and partners) · USA
One of the first open chat models (2023): LLaMA fine-tuned on user conversations with ChatGPT. A historical milestone; today it is weaker than any modern model of the same size.
- Experiments and team training
- Simple chat assistant for tests
- Comparison with newer models
- Sizes
- 7B – 33B
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, and others) · China
Transformer-based photo upscaling, more accurate than SwinIR on fine details. Versions for real noisy photos and a lightweight HAT-S.
- Upscaling product and interior photos
- Preparing images for print
- Sharpening archival photos
- Sizes
- 9M – 40M
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2022–2023
Hendrik Schröter (University of Erlangen) · Germany
Lightweight real-time speech noise suppression that works even on a regular CPU and low-power devices. Removes hum, street and office noise while keeping the voice.
- Cleaning calls and voice messages of noise
- Preparing recordings before speech recognition
- Noise suppression for video calls
- Sizes
- about 2M
- Hardware
- from: Laptop
CybersecurityNot maintained2022–2023
Ehsan Aghaei and co-authors · USA
A compact language encoder trained on cybersecurity texts: tagging threat reports, finding entities and classification.
- Tagging threat reports and vulnerability bulletins
- Extracting entities from security texts
- Classifying and searching an internal incident base
- Sizes
- около 125M
- Hardware
- from: Laptop
AvatarsNot maintained2023
Xi'an Jiaotong University and Tencent AI Lab · China
An older lightweight talking-head model: one photo plus audio becomes a video. Runs on weak hardware, but quality is noticeably below newer models.
- Talking photo for greetings
- Simple voiced avatars
- Sizes
- under 1B
- Hardware
- from: Laptop
Text to speechRUGGUFNot maintained2023
Suno · USA
One of the first open models to voice text with intonation, laughter and pauses. Supports about ten languages, including Russian. Now outdated.
- Draft voiceovers for videos
- Voice service prototypes
- Sound effects in speech
- Sizes
- about 300M – 1B
- Hardware
- from: Laptop
Weather and climateNot maintained2023
Huawei Cloud · China
One of the first weather neural networks, published in Nature and added to ECMWF charts. The weights are open for research only; commercial use is prohibited.
- Research weather forecasts a week ahead
- Comparison with other weather models on your own data
- Training courses on weather neural networks
- Sizes
- 4 models of ~1.1 GB each (1, 3, 6 and 24-hour steps)
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2023
Numbers Station · USA
One of the first open text-to-SQL lines, including very small versions from 350M that run on an ordinary PC.
- Turning a question into SQL from a table description
- Hints while writing queries
- A local analyst helper with no data leaving the company
- Sizes
- 350M – 7B
- Hardware
- from: Laptop
3DNot maintained2022–2023
OpenAI · USA
Early open OpenAI models that create a 3D object from text or an image in seconds. Quality is basic, but they are fast and easy to run.
- Rough 3D mock-ups from a description
- Quick object prototypes for games and AR
- Training and research pilots in 3D
- Sizes
- 40M – 1B
- Hardware
- from: Laptop
Text analysisRUNot maintained2022–2023
David Dale (cointegrated) and the community · Russia
Ready-made tiny rubert-tiny models for Russian text: detect rudeness and insults, sentiment and emotions. They run on a CPU in milliseconds.
- Filtering insults in Russian chats and comments
- Labeling reviews as positive, neutral or negative
- Spotting irritated customers in requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
Photo editingNot maintained2022–2023
S-Lab, Nanyang Technological University · Singapore
Popular face restoration for old and blurry photos; also works on video. Has face inpainting and colorization modes. Non-commercial license.
- Restoring faces in old photos
- Enhancing faces in low-quality video
- Family archive restoration pilots
- Sizes
- small, up to 0.1B
- Hardware
- from: Laptop
TextNot maintained2022–2023
Google · USA
Compact input-output models trained to follow instructions. Still used as a cheap base for classification, extraction and short answers.
- Classification of requests and documents
- Extracting fields from text
- Short answers and summaries
- Sizes
- 80M – 20B
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2022–2023
Meta · USA
A translator for 200 languages, including rare and minor ones. Russian is supported. Strong language coverage, but the license prohibits commercial use.
- Translating texts between 200 languages
- Translating into rare languages where no other models exist
- Comparing quality when choosing a translator
- Sizes
- 600M to 3.3B (plus 54B MoE)
- Hardware
- from: Laptop
Image + textNot maintained2023
Google · USA
Reads a document or a screenshot as an image and answers with structure: text, fields, answers to questions. In HR it is fine-tuned for resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting data from resumes and forms supplied as images
- Questions about the content of a scan
- Parsing tables and diagrams in documents
- Sizes
- 282M – 1.3B
- Hardware
- from: Laptop
TextNot maintained2022–2023
EleutherAI · USA
Fully open models from the non-profit lab EleutherAI: GPT-NeoX-20B and the Pythia series with published intermediate training checkpoints.
- Base model for fine-tuning
- Research into model behavior
- Simple text generation and completion
- Sizes
- 70M – 20B
- Hardware
- from: Laptop
TextNot maintained2022
Meta · USA
An early open Meta series matching GPT-3 in size. Outdated; useful for research and comparison.
- Research experiments
- Training specialists
- Comparison with modern models
- Sizes
- 125M – 66B (175B on request)
- Hardware
- from: Laptop
TextGGUFNot maintained2022
BigScience (Hugging Face and the community) · France
One of the first large open models, trained by a community of hundreds of researchers in 46 languages. Today it is interesting mostly as a historical milestone.
- Text generation and translation in many languages
- Experiments and team training
- Base model for fine-tuning on a narrow task
- Sizes
- 560M – 176B
- Hardware
- from: Laptop
Music and soundNot maintained2022
MIT · USA
A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.
- Sound event recognition
- Tagging an audio archive
- Detecting alarm sounds
- Sizes
- about 87M
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
SCUT DLVC Lab, South China University of Technology · China
A light model that takes both the text and the position of blocks on the page into account: trained in one language and transferable to others. Good for tagging fields in resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Tagging fields in resumes and forms
- Extracting data from forms and templates
- Parsing documents in several languages
- Sizes
- about 130M for the English version and about 280M for the multilingual one
- Hardware
- from: Laptop
Documents and OCRNot maintained2021–2022
Microsoft · USA
Recognizes a single line of text, including handwriting. The official weights are English only, but the model is often fine-tuned for other languages; there are community Russian versions.
- Recognizing handwritten lines in questionnaires and forms
- Recognizing printed lines after text detection on the page
- A base for fine-tuning to your own handwriting or font
- Sizes
- 62M – 608M
- Hardware
- from: Laptop
Text analysisNot maintained2020–2022
Microsoft · USA
Classic document understanding models: they take into account the text, its position on the page and the image. They are fine-tuned to extract fields from forms and receipts. Only the first version is free for commercial use.
- Extracting fields from questionnaires, forms and receipts after fine-tuning
- Classifying document types
- Answering questions about a scanned page
- Sizes
- about 110M – 370M
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
NAVER CLOVA · South Korea
Reads a scanned document and returns a filled-in field structure straight away, with no separate OCR step. In HR it is fine-tuned for parsing resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from forms and resumes
- Parsing scans of certificates and diplomas
- Detecting the type of an incoming document
- Sizes
- about 200M
- Hardware
- from: Laptop
RerankersRUNot maintained2022
UKP Lab and the Sentence Transformers community · Germany
The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.
- Reordering knowledge base search results
- Selecting passages before a chatbot answers
- Finding duplicates among tickets and product cards
- Sizes
- about 4M – 120M
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021–2022
Tencent ARC Lab · China
The classic for upscaling photos 2–4x while cleaning noise and compression artifacts. Lightweight, runs even on a CPU. Versions for drawings and anime.
- Upscaling old and small product photos
- Cleaning images of compression artifacts
- Preparing images for print
- Sizes
- about 17M
- Hardware
- from: Laptop
Computer visionNot maintained2021–2022
OpenAI · USA
The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.
- Image search by text query
- Automatic tags for a catalog
- Finding similar images
- Sizes
- about 0.15B – 0.6B
- Hardware
- from: Laptop
Computer visionNot maintained2022
Microsoft · USA
From an image it works out what kind of document it is: resume, diploma, certificate, contract. Helps sort candidate file bundles by type. A human makes the decision about a candidate; automatic screening without review must not be used.
- Sorting incoming candidate documents by type
- Checking that a document package is complete
- Finding the right scan in an archive
- Sizes
- base and large versions
- Hardware
- from: Laptop
Photo editingNot maintained2021–2022
Tencent ARC Lab · China
Proven face restoration for old and compressed photos, with a commercial-friendly license. Often paired with Real-ESRGAN; the most used versions are 1.3 and 1.4.
- Restoring faces in old photos
- Enhancing avatars and profile photos
- Restoration in a photo shop or online service
- Sizes
- small, up to 0.1B
- Hardware
- from: Laptop
Text analysisRUNot maintained2019–2022
Meta · USA
A classic multilingual encoder for 100 languages, including Russian. The base of many sentiment, NER and embedding models, including BGE-M3.
- Detecting review sentiment in different languages
- Extracting names and organizations after fine-tuning
- Classifying requests
- Sizes
- 270M – 10.7B
- Hardware
- from: Laptop
Computer visionRUNot maintained2022
Sber AI and SberDevices (ai-forever) · Russia
A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.
- Product search by photo and by Russian description
- Sorting images into categories without labeling
- Checking that a photo matches its caption
- Sizes
- 150M – 430M
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
Microsoft · USA
A time-tested encoder behind many classifiers and NER models (including GLiNER). The multilingual mDeBERTa-v3 understands Russian.
- Classifying review sentiment
- Entity extraction after fine-tuning
- Checking whether a conclusion follows from a text
- Sizes
- 70M – 435M
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
David Dale (cointegrated) · Russia
A very small Russian-English BERT that runs fast on a regular CPU. Ready-made fine-tuned versions exist for sentiment, toxicity and emotions.
- Detecting review sentiment
- Filtering rude chat messages
- Fast classification of requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
Deepfake detectionNot maintained2021
NAVER Clova AI Research and EURECOM · South Korea
The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.
- Voice check during phone authentication
- Filtering replays and synthesis in a voice menu
- A baseline when comparing voice detectors
- Sizes
- weight files of 0.4 and 1.3 MB
- Hardware
- from: Laptop
Photo editingGGUFNot maintained2021
Samsung AI Center Moscow (with Skoltech) · Russia
Removes unwanted objects, text and watermarks from photos with clean background fill. Lightweight and fast; still the standard for this task.
- Removing price tags, people and clutter from photos
- Cleaning interior and real estate photos
- Removing text and dates from archival photos
- Sizes
- about 51M
- Hardware
- from: Laptop
Photo editingNot maintained2021
ETH Zurich · Switzerland
A transformer model for upscaling, denoising and removing JPEG artifacts from photos. Lightweight and proven; often embedded in other systems.
- Photo upscaling
- Image denoising
- Removing compression artifacts
- Sizes
- about 12M
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2020–2021
Meta · USA
An early Meta translator that translates directly between 100 languages, without English in the middle. Russian is supported. Old, but light and permissively licensed.
- Translation between any pair of 100 languages
- Quick draft translation on modest hardware
- Base for fine-tuning to your subject area
- Sizes
- 418M – 12B
- Hardware
- from: Laptop
FinanceNot maintained2020
Prosus · Netherlands
A classic model that determines the tone of financial news: positive, negative or neutral. English only, runs fast on a CPU.
- Scoring the tone of company news
- Labeling reports and press releases
- Signals for analytics dashboards
- Sizes
- 110M
- Hardware
- from: Laptop
AvatarsNot maintained2020
IIIT Hyderabad · India
The classic lip-to-audio sync model, still popular in hobbyist setups. Lip movements are accurate but the face looks blurry; the license is non-commercial.
- Quick dubbing tests
- Comparison with newer lip-sync models
- Educational and research projects
- Sizes
- small model, 96-pixel face
- Hardware
- from: Laptop
Deepfake detectionNot maintained2020
MiniVision Technology · China
Practically the only fully open weight set for single-frame face liveness: it tells a live person from a photo, a screen or a mask. It errs in both directions - a person must be able to appeal a rejection.
- Liveness check when signing in by selfie
- Protecting an access system from a photo on a phone
- A check during remote customer identification
- Sizes
- two models of about 1.8 MB each
- Hardware
- from: Laptop