TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextGGUF2025–2026
Xiaomi · China
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
- Logic and calculation tasks
- Agents with tools
- Help for developers
- Sizes
- 7B – 1,02T-A42B
- Hardware
- from: Laptop
Text2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact text models that run directly on a device: laptop, phone or mini PC. The 1B and 2B MiniCPM5 models focus on tool calling and long context.
- A local chat assistant without the cloud
- Data extraction and text classification
- Tool calling and simple agents on low-end hardware
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Text2024–2026
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
- Reasoning, maths and technical questions
- Analysing long documents
- Agents and writing code
- Sizes
- 0.9B – 375B-A23B
- Hardware
- from: Laptop
TextGGUF2025–2026
Renmin University of China (GSAI) and Ant Group (inclusionAI) · China
Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.
- Fast generation of code and text
- Agent scenarios with long context
- Research into alternatives to standard LLMs
- Sizes
- 8B – 100B (MoE)
- Hardware
- from: 1 GPU
TextRU2026
SberDevices (ai-forever) · Russia
A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.
- Experiments with faster generation
- Fine-tuning small models for your own tasks
- Research
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
Search and RAGRU2024–2026
Sber (SberDevices) · Russia
Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.
- Search across Russian-language documents
- RAG for chatbots in Russian
- Classifying requests and reviews
- Sizes
- 480M – 10B-A1.8B
- Hardware
- from: Laptop
Speech to textGGUF2024–2026
Moonshine AI (Useful Sensors) · USA
Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.
- Voice control of devices
- Offline recognition on a phone
- Live subtitles
- Sizes
- 27M – 245M
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
Text2024–2026
Tencent · China
Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.
- Corporate assistant
- Translation and multilingual texts
- Agents with tools
- Sizes
- 0.5B – 770B-A49B
- Hardware
- from: Laptop
Text2025–2026
Ant Group (inclusionAI) · China
An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.
- Corporate assistant
- Agents for office processes
- Financial analytics (Fin version available)
- Sizes
- 7.9B-A1.3B – 1T
- Hardware
- from: Laptop
TextOllama2024–2026
NVIDIA · USA
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
- Agents with tool calling
- Reasoning and calculation tasks
- Answers based on long documents
- Sizes
- 4B – 550B-A55B
- Hardware
- from: Laptop
TextOllama2024–2026
IBM · USA
IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.
- Answers based on internal documents (RAG)
- Tool calling and agent work
- Data extraction and classification
- Sizes
- 350M – 34B
- Hardware
- from: Laptop
TextOllama2026
Meta Superintelligence Labs · USA
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
- Agents with tool calling
- Analysis of screenshots, charts and documents
- Multilingual assistant
- Sizes
- 30B
- Hardware
- from: 1 GPU
Text analysis2024–2026
Urchade Zaratiana and Fastino AI · France / USA
Finds the entities you need in text without training: just list what to look for (name, amount, date). GLiNER2 also classifies text. Multilingual versions understand Russian.
- Extracting names, amounts and dates from emails and contracts
- Parsing requests into CRM fields
- Classifying requests by topic
- Sizes
- about 50M to 500M
- Hardware
- from: Laptop
Computer-use agents2025–2026
XLANG Lab (University of Hong Kong) · China
Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.
- Working in desktop software without an API
- Moving data between systems
- Running user scenarios for tests
- Sizes
- 7B – about 400B (MoE)
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2025–2026
Ant Group (inclusionAI) · China
An Ant Group family for finding elements on screen and completing tasks in phone and computer interfaces. UI-Venus-2 was specifically trained to refuse dangerous actions.
- Automating actions in mobile apps
- Filling in forms in web interfaces
- UI autotests
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
CodeOllama2026
DeepReinforce · not disclosed
Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.
- A developer agent in the terminal
- Fixing bugs from a task description
- Understanding and extending a large repository
- Sizes
- 9B – 397B
- Hardware
- from: 1 GPU
TextRUOllama2023–2026
Mistral AI · France
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
- Fast chat responses
- Data extraction from text
- Translation and multilingual work
- Sizes
- 3B – 675B
- Hardware
- from: Laptop
Code2025–2026
Kwaipilot (Kuaishou) · China
Kuaishou models for agentic development, trained to solve real tasks in repositories. KAT-Coder-V2.5-Dev (35B, 3B active) is the open version of their closed flagship.
- An agent that fixes tasks in the repository
- Code generation and refactoring
- Automating routine development tasks
- Sizes
- 32B – 72B, 35B-A3B
- Hardware
- from: 1 GPU
Search and RAGRU2025–2026
NVIDIA · USA
NVIDIA embeddings for search and RAG. Nemotron-3-Embed, released in 2026, is under the permissive OpenMDW license and works in many languages.
- Search across corporate documents
- RAG for chatbots and assistants
- Search across images and pages (VL versions)
- Sizes
- 1B – 8B
- Hardware
- from: Laptop
TextGGUF2025–2026
Moonshot AI · China
Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.
- Multi-step agents: search, data collection, reports
- In-depth document analysis
- Help for developers
- Sizes
- 16B-A3B – 2.8T-A104B
- Hardware
- from: 1 GPU
TextOllama2024–2026
LG AI Research · South Korea
Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.
- Corporate assistant
- Working with Korean and English texts
- Analysis of documents and images (4.5)
- Sizes
- 1.2B – 750B-A37B
- Hardware
- from: Laptop
TextOllama2023–2026
Upstage · South Korea
Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.
- Working with office documents
- Agents for routine tasks
- Help for developers
- Sizes
- 10.7B – 250B-A15B
- Hardware
- from: 1 GPU
TextGGUF2026
Thinking Machines Lab · USA
Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.
- Flagship-level corporate assistant
- Analysis of documents, images and audio
- Programming help
- Sizes
- 276B-A12B, 975B-A41B
- Hardware
- from: Cluster
CodeGGUF2026
Poolside · USA
Models for agentic programming: they edit code in a repository on their own. The small XS runs on a Mac with 36 GB of memory; S 2.1 has a 1M-token context.
- Coding agent for in-house development
- Bug fixing and code improvements
- Working with large codebases
- Sizes
- 33B-A3B – 225B-A23B
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2025–2026
Microsoft · USA
Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.
- Filling in web forms and applications
- Collecting data from web portals without an API
- Checking websites against scenarios
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
RerankersGGUF2024–2026
Jina AI · Germany
Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.
- Refining search results before a chatbot answers
- Sorting retrieved PDF pages and slides
- Catalog and knowledge base search
- Sizes
- 33M – 2.4B
- Hardware
- from: Laptop
Rerankers2026
Tencent · China
A pair of small Tencent models based on Qwen3 that pick the right skill for an AI agent for a given request: the embedding model finds candidates, the reranker chooses the best one.
- Choosing a tool or skill for an AI agent
- Routing requests between bot scenarios
- Search across a catalog of internal tools
- Sizes
- 0.6B
- Hardware
- from: Laptop
CodeGGUF2025–2026
JetBrains · Czech Republic
JetBrains models for fast code autocompletion. Mellum2 (12B, 2.5B active) is already a full assistant: it writes and edits code, calls tools and reasons.
- Fast code autocompletion on your own server
- A developer assistant that does not send code to the cloud
- Fine-tuning on the company's code
- Sizes
- 4B – 12B-A2.5B
- Hardware
- from: Laptop
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
MiniMax · China
Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.
- Analysis of large document archives in a single request
- Agents with tools
- Help for developers
- Sizes
- 230B-A10B – 456B-A46B
- Hardware
- from: Cluster
Search and RAGRUGGUF2023–2026
Jina AI · Germany
Strong multilingual embeddings with long context; v5-omni understands text, images and audio. Recent versions are open for non-commercial use only.
- Search across documents in many languages
- Search across images and scans
- Classification and clustering
- Sizes
- 33M – 3.8B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Cognitive Computations (Eric Hartford) · USA
Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.
- Assistant that does not refuse legal but sensitive topics
- Internal tools under a strict system prompt
- Role-play and creative scenarios
- Sizes
- 0.5B – 405B
- Hardware
- from: Laptop
TextGGUF2025–2026
StepFun · China
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
- High-load agents
- Analysis of documents with diagrams and screenshots
- Help for developers
- Sizes
- 8B – 321B
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2025–2026
H Company · France
A French model family for controlling a browser and computer: precisely finds the right element on screen and handles multi-step tasks. The latest Holo3 and 3.1 are open under Apache 2.0.
- Working in web portals and legacy software without an API
- Filling in forms and applications
- Testing interfaces against scenarios
- Sizes
- 0.8B – 235B-A22B
- Hardware
- from: Laptop
Rerankers2025–2026
NVIDIA · USA
A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.
- Reordering passages before an AI assistant answers
- Sorting retrieved scan and PDF pages
- Search across internal policies and instructions
- Sizes
- 1B
- Hardware
- from: Laptop
Forecasting2024–2026
THUML, Tsinghua University · China
A compact forecasting foundation model from the Tsinghua lab: trained on a large set of diverse series and fine-tunable on your own data.
- Forecasting demand and load
- Forecasting sensor readings on the shop floor
- Fine-tuning forecasts on your own history
- Sizes
- 84M (timer-base)
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2026
IBM · USA
Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.
- Search across corporate documents
- RAG on a regular server without a GPU
- Reranking results
- Sizes
- 30M – 311M
- Hardware
- from: Laptop
ForecastingGGUF2025–2026
Datadog · USA
A Datadog forecasting model trained on server and application metrics. Especially strong for IT monitoring: load, latency, errors.
- Server load forecasting
- Anomaly detection in metrics
- Capacity planning
- Sizes
- 4M – 2.5B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
Arcee AI · USA
An American family of MoE models trained from scratch: Nano, Mini and Large. Trinity-Large-Thinking (398B) reasons before answering.
- Agents with tool calling
- Reasoning tasks
- Corporate assistant on your own servers
- Sizes
- 6B – 398B-A13B
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2026
IBM · USA
IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.
- Checking bot requests and replies
- Finding made-up facts in knowledge-base answers
- Checking your own rules written as text
- Sizes
- 38M – 8B
- Hardware
- from: Laptop
Computer-use agents2024–2026
Show Lab (National University of Singapore) · Singapore
A lightweight model for working with interfaces: finds buttons and fields by description and performs actions on the web and on a phone. ShowUI-π can drag with the mouse.
- Clicking and filling in forms from a task description
- Web UI autotests
- An assistant on a low-end computer without the cloud
- Sizes
- 2B (ShowUI), about 500M (ShowUI-π)
- Hardware
- from: Laptop
Text to speech2025–2026
Kyutai · France
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
- Streaming speech transcription for voice bots
- Voicing replies with minimal delay
- Speech synthesis on a server without a GPU (Pocket TTS)
- Sizes
- 100M (Pocket TTS) – 2.6B
- Hardware
- from: Laptop
TextGGUF2025–2026
ServiceNow · USA
ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.
- A reasoning assistant for internal services
- Tool calling and enterprise agents
- Analysing screenshots and documents with images
- Sizes
- 5B – 15B
- Hardware
- from: Laptop
CodeOllama2025–2026
Essential AI · USA
An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.
- Writing and fixing code
- A developer agent on a single GPU
- Solving technical and scientific problems
- Sizes
- 8B
- Hardware
- from: Laptop
Search and RAGGGUF2025–2026
Octen · USA / Singapore
Qwen3-Embedding models fine-tuned by the startup Octen for search in legal, financial and medical texts. As of January 2026 the 8B version topped the RTEB leaderboard.
- Search across contracts and case law
- Search across financial reports
- Search across long documents up to 32K tokens
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2025–2026
Princeton University · USA
Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.
- Formal verification of mathematical workings
- Verifying code correctness
- Training
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Search and RAGGGUF2026
Perplexity · USA
Embeddings from the Perplexity search service. Some versions take into account the context of the whole document, not just a single fragment.
- Search across large document collections
- RAG that accounts for document context
- Website and catalog search
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
Voice: speakers and soundRU2026
FireRedTeam (Xiaohongshu) · China
A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.
- Cutting recordings before speech recognition
- Separating speech from music and singing in broadcasts and videos
- Speech detection in voice bots
- Sizes
- compact, exact size not stated
- Hardware
- from: Laptop
Search and RAGRU2026
Microsoft · USA
Microsoft's 2026 multilingual embeddings with context up to 32K tokens; Russian is on the language list. The 270M and 0.6B versions run on a regular server, 27B is the most accurate.
- Multilingual knowledge base search
- Picking passages for RAG
- Search across long documents
- Sizes
- 270M – 27B
- Hardware
- from: Laptop
CodeOllama2024–2026
Alibaba (Qwen team) · China
The broadest open coding family: from 0.5B for autocompletion to 480B for agents. Qwen3-Coder-Next (80B, 3B active) works as a developer agent on a single GPU.
- Code autocompletion in the editor
- An agent that edits code in the repository on its own
- Writing and refining scripts, SQL and integrations
- Sizes
- 0.5B – 480B-A35B
- Hardware
- from: Laptop
CodeGGUF2026
IQuest Research · China
A family of coding models with standard and reasoning versions, including a Loop variant that runs through its layers a second time. Sizes from 7B to 40B.
- Writing and refining code
- Solving tasks with step-by-step reasoning
- Agentic work with a repository
- Sizes
- 7B – 40B
- Hardware
- from: Laptop
Code2026
Ai2 (Allen Institute for AI) · USA
Fully open developer agents from Ai2: weights, data and training recipe are all public. Designed so a company can cheaply fine-tune the agent on its own repository.
- An agent for fixing issues in code
- Fine-tuning the agent on an internal repository
- Automating small edits and tests
- Sizes
- 8B – 32B
- Hardware
- from: Laptop
Voice assistants2024–2026
Kyutai · France
A voice assistant that listens and speaks at the same time, with no delay for recognition and synthesis. Hibiki does simultaneous speech-to-speech translation between several European languages.
- Real-time voice conversation partner
- Simultaneous speech translation
- Zero-latency voice interfaces
- Sizes
- 2B – 7B
- Hardware
- from: Laptop
Voice assistantsGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.
- Voice assistant on your own server
- Analyzing videos and documents
- Voice answers about a camera image
- Sizes
- 8B – 9B
- Hardware
- from: Laptop
Computer-use agents2025–2026
Alibaba (Tongyi Lab, X-PLUG) · China
Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.
- Automating actions in mobile apps
- Working in desktop software without an API
- Testing apps against scenarios
- Sizes
- 2B – 32B
- Hardware
- from: Laptop
Search and RAGRUOllama2025–2026
Alibaba (Qwen) · China
Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.
- Knowledge base search for RAG
- Reranking results before answering
- Search across scans, slides and screenshots
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Search and RAG2026
Voyage AI (MongoDB) · USA
The only open model in the Voyage 4 line: its vectors are compatible with the paid larger versions, so you can start locally and move to the API later.
- Document search on your own server
- RAG for small knowledge bases
- Finding similar texts
- Sizes
- about 340M
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
TextGGUF2024–2026
Prime Intellect · USA
Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.
- Reasoning and math tasks
- Programming help
- Agents with tool calling
- Sizes
- 10B – 106B-A12B
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2026
Meituan · China
Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.
- Working in office and legacy software without an API
- Moving data between systems
- Running test scenarios
- Sizes
- 8B – 32B
- Hardware
- from: 1 GPU
Cybersecurity2025–2026
Cisco (Foundation AI) · USA
Cisco models for information security based on Llama 3.1 8B: analysis of vulnerabilities, threats and incidents. Can be deployed inside your own perimeter.
- Analyzing vulnerability and threat reports
- Helping SOC analysts during incidents
- Mapping threats to MITRE ATT&CK
- Sizes
- 8B
- Hardware
- from: Laptop
Voice: speakers and soundRU2025–2026
Daily (Pipecat) · USA
Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.
- Voice bot does not interrupt the customer during pauses
- Fast reply when the customer has really finished
- An add-on to a standard speech detector in voice assistants
- Sizes
- 8M (v3) – 580M (v1)
- Hardware
- from: Laptop
CodeRUOllama2025
Mistral AI (with All Hands AI) · France
Mistral models for agentic development: they read the repository, edit files and run commands on their own. The 24B version fits on a single GPU.
- A developer agent that fixes tickets from the tracker
- Extending internal systems from a description
- Automating routine code edits
- Sizes
- 24B – 123B
- Hardware
- from: 1 GPU
Computer-use agents2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.
- Automating actions in mobile apps
- Working in web interfaces without an API
- Testing apps against scenarios
- Sizes
- 9B – 18B
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2025
Alibaba (Tongyi-MAI) · China
Compact Alibaba models for working in smartphone and computer interfaces: they find elements and complete multi-step tasks. The small size allows running on an ordinary GPU.
- Automating actions in mobile apps
- Working in software without an API
- UI autotests
- Sizes
- 2B – 8B
- Hardware
- from: Laptop
Moderation and safety2025
ServiceNow · USA
A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.
- Screening chatbot requests for attacks and jailbreaks
- Filtering harmful model answers
- Monitoring the actions of AI agents that use tools
- Sizes
- 8B
- Hardware
- from: Laptop
Tabular dataGGUF2024–2025
Zhejiang University · China
A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.
- Answering questions about tables and data exports
- Automated data analysis with generated code
- A helper for BI and internal reporting
- Sizes
- 7B – 72B
- Hardware
- from: 1 GPU
TextOllama2023–2025
Nous Research · USA
Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.
- Agents that call functions and APIs
- Data extraction in strict JSON format
- Assistant with flexible role and tone settings
- Sizes
- 3B – 405B
- Hardware
- from: Laptop
Voice: speakers and soundRU2020–2025
Silero · Russia
The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.
- Cutting calls and recordings before speech recognition
- Detecting when the customer is speaking in a voice bot
- Filtering out silence and noise to save on transcription
- Sizes
- about 2 MB
- Hardware
- from: Laptop
TextOllama2025
Deep Cogito · USA
Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.
- A chat assistant with a reasoning mode
- Writing code and calling tools
- Answering complex questions about documents
- Sizes
- 3B – 671B
- Hardware
- from: Laptop
Rerankers2025
ZeroEntropy · USA
Rerankers built on Qwen3. The model card lists the target domains — finance, law, code, medicine, science; the stated language is English.
- Refining results before an AI assistant answers
- Sorting search results across contracts and reports
- Search across technical and scientific documentation
- Sizes
- zerank-2 — 4B (based on Qwen3-4B), plus a smaller "small" version
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Mixedbread · Germany
Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.
- Search across a knowledge base
- Reranking results before a bot answers
- Product catalog search
- Sizes
- 17M – 1.5B
- Hardware
- from: Laptop
Visual document search2025
Illuin Technology, EPFL, CentraleSupélec · France
A compact (250M) model for searching document pages as images. According to the authors, it matches models 10 times larger and runs without a GPU.
- Search across scans and PDFs on a modest server
- Indexing document archives
- Search across slides and manuals
- Sizes
- 250M
- Hardware
- from: Laptop
Search and RAGOllama2025
Google · USA
A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.
- On-device document search
- RAG without sending data outside
- Text classification
- Sizes
- 300M
- Hardware
- from: Laptop
Search and RAGOllama2023–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.
- Search across English-language documents
- Picking passages for chatbot answers (RAG)
- Code search (bge-code)
- Sizes
- 24M – 9B
- Hardware
- from: Laptop
Text to SQLGGUF2024–2025
Prem AI · UK
A text-to-SQL model of just 1B parameters, designed to run locally so the database never leaves for external services.
- Local translation of questions into SQL with no internet access
- Query hints on modest hardware
- Embedding into internal analytics tools
- Sizes
- 1B
- Hardware
- from: Laptop
TextOllama2025
OpenAI · USA
OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.
- AI agent that calls internal systems
- Answers based on internal policies
- Drafts of emails and reports
- Sizes
- 20B, 120B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Moonshot AI and Project Numina · China, France
Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.
- Formal verification of mathematical workings
- Translating a problem from plain language into Lean
- Training and olympiad preparation
- Sizes
- 0.6B – 72B
- Hardware
- from: Laptop
TextGGUF2025
ByteDance · China
An open ByteDance 36B model with up to 512K tokens of context and an adjustable thinking budget. Fits on a single powerful GPU.
- Analysis of long documents
- Agents with tools
- Corporate assistant
- Sizes
- 36B
- Hardware
- from: 1 GPU
Text2024–2025
xAI · USA
xAI publishes the weights of previous Grok generations. The models are very large and need a GPU cluster, so in practice they are rarely run.
- Research on large models
- Assistant on your own infrastructure
- Text generation and analysis
- Sizes
- 314B (Grok-1), Grok-2 is larger
- Hardware
- from: Cluster
Voice: speakers and sound2025
Agora (TEN project) · USA / China
A lightweight speech detector for real-time voice assistants: it notices the start and end of a phrase faster than Silero VAD. Runs on servers, phones and in the browser.
- Zero-lag speech detection in a voice bot
- Fast assistant response at the end of a phrase
- Use in mobile apps and the browser
- Sizes
- very small, the library is smaller than Silero VAD
- Hardware
- from: Laptop
Math and reasoningOllama2025
Agentica (Berkeley, Sky Computing Lab) and Together AI · USA
Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.
- Solving maths problems with step-by-step working
- Generating and checking code
- An agent for fixing bugs in a repository
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Visual document search2024–2025
Illuin Technology (ViDoRe team) · France
Searches PDFs and scans as images: pages do not need to be OCR'd first, the model finds the right one for a question directly, including tables and charts. Trained on English.
- Search across scans, presentations and PDFs
- RAG over documents with tables and charts
- Search across technical documentation
- Sizes
- 256M – 3B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
OpenCompass (Shanghai AI Laboratory) · China
A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring model answers against set criteria
- Checking an answer against a reference solution
- Comparing several models on your own data
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoning2025
NVIDIA · USA
NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Working through programming problems
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
Skywork (Kunlun Tech) · China
Reward models: they score how good a language model's answer is for the user. Used for fine-tuning your own models and picking the best of several answers.
- Choosing the best of several bot answers
- Scoring answer quality during model fine-tuning
- Comparing models before rollout
- Sizes
- 0.6B – 27B
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Clouditera · China
A Chinese open family for cybersecurity: reviewing vulnerabilities, analysing logs and traffic, explaining commands and scripts.
- Reviewing vulnerabilities and drafting fix recommendations
- Analysing logs and reconstructing an attack chain
- Explaining suspicious commands and scripts
- Sizes
- 1.5B – 14B
- Hardware
- from: Laptop
CybersecurityGGUF2025
Trendyol · Turkey
Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.
- Reviewing alerts and first-pass incident assessment
- Explaining suspicious activity in reports
- Helping the on-duty shift of a monitoring centre
- Sizes
- 32B и 70B
- Hardware
- from: 1 GPU
Text to SQLGGUF2025
IDEA Research · China
A text-to-SQL model trained with reinforcement learning: it works through the schema and the conditions step by step before producing a query.
- Database queries for questions with several conditions
- Reviewing and fixing other people SQL queries
- An analyst helper inside a BI system
- Sizes
- 3B – 14B
- Hardware
- from: Laptop
TextOllama2025
DeepSeek · China
A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.
- Complex calculations and logic checks
- Analysis of contracts and internal policies
- Help for developers
- Sizes
- 1,5B – 671B
- Hardware
- from: Laptop
CodeGGUF2025
ByteDance Seed · China
A compact 8B coding model from ByteDance in base, instruct and reasoning versions. Its training data was selected by the model itself, with almost no hand-written rules.
- Code autocompletion and generation
- Solving algorithmic problems
- A base for fine-tuning on your own stack
- Sizes
- 8B
- Hardware
- from: Laptop
Image generationGGUF2025
ByteDance Seed · China
A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.
- Photo editing in a conversation
- Answering questions about an image
- Image generation with explanations
- Sizes
- 14B-A7B
- Hardware
- from: 1 GPU
Text analysisRU2023–2025
deepvk (VK) · Russia
Russian encoders from the VK team: RuModernBERT reads long texts, USER produces vectors for search, GeRaCl classifies texts by topic without training.
- Classifying requests without labeled data
- Knowledge base search in Russian
- Analyzing long contracts
- Sizes
- 35M – 360M
- Hardware
- from: Laptop
Text to SQL2025
Snowflake · USA
A Snowflake model for turning questions into SQL, trained with reinforcement learning by checking query results. The open 7B version is based on Qwen2.5-Coder.
- Plain-language questions to a data warehouse
- Generating SQL for reports and dashboards
- Checking and fixing analysts' queries
- Sizes
- 7B
- Hardware
- from: Laptop
CodeRU2025
MTS AI (MWS AI) · Russia
A small coding assistant from MTS AI that understands requests in Russian. Runs locally, with plugins for VS Code and JetBrains.
- Code suggestions and completion in the editor
- Code explanations in Russian
- Drafts of tests and documentation
- Sizes
- 1.5B
- Hardware
- from: Laptop
Rerankers2023–2025
Stanford NLP, later Answer.AI and LightOn · USA and France
A different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.
- Search across a knowledge base of long documents
- Reordering retrieved passages
- Search across policies and technical documentation
- Sizes
- about 33M – 150M
- Hardware
- from: Laptop
Fact-checking and judgesRU2024–2025
NVIDIA · USA
Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data to fine-tune your own model
- Scoring assistant answers in Russian and other languages
- Sizes
- 49B, 70B and 340B
- Hardware
- from: Cluster
Forecasting2025
THUML, Tsinghua University · China
A forecasting model that returns a set of possible scenarios rather than a single line — useful when you need a range for demand or load, not one number.
- Forecasting demand with a range of values
- Planning stock while accounting for spread
- Forecasting load on services and staff
- Sizes
- 128M (sundial-base)
- Hardware
- from: Laptop
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek models for formal proofs in Lean 4: the proof is checked by a program, not a person. A narrow tool for mathematicians and engineers.
- Formal verification of mathematical workings
- Verifying algorithm correctness
- Training and olympiad preparation
- Sizes
- 7B – 671B
- Hardware
- from: Laptop
Moderation and safety2024–2025
Meta · USA
Tiny classifiers that catch attempts to hack a bot: prompt injections and rule bypassing. The 86M version is multilingual, 22M is English only.
- Protecting a bot from prompt injections
- Checking emails and documents that reach an AI agent
- Fast filter in front of a large model
- Sizes
- 22M – 86M
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
ByteDance Seed · China
A model that looks at a screenshot and controls the mouse and keyboard itself: clicks, fills in fields, navigates menus. The first generation and 1.5-7B are open; UI-TARS-2 weights were not released.
- Working in legacy software without an API
- Filling in forms and moving data between systems
- UI autotests from plain-language scenarios
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Text to SQL2025
Alibaba · China
Alibaba models for turning questions into SQL, based on Qwen2.5-Coder. They work with different SQL dialects; a small 3B version suits modest hardware.
- Plain-language database questions
- Queries for different databases (PostgreSQL, MySQL, SQLite)
- Automating routine reports
- Sizes
- 3B – 32B
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Kindo · USA
One of the best-known open families for security and DevSecOps work: reviewing code for weaknesses, test scenarios, explaining attacks.
- Finding weak spots in code and configurations
- Reviewing incidents and explaining attack techniques
- Drafting scripts and procedures for the security team
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Image generationGGUF2024–2025
NVIDIA · USA
NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.
- Bulk image generation
- High-resolution visuals
- Real-time generation inside apps
- Sizes
- 0.6B – 4.8B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Nomic AI · USA
Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.
- Search across documents and a knowledge base
- Code search
- Search across scans and PDFs without text recognition
- Sizes
- 137M – 7B
- Hardware
- from: Laptop
Text to speechRUGGUF2023–2025
Rhasspy / Open Home Foundation · USA
Very fast speech synthesis that runs even on a Raspberry Pi. Ready-made voices in 35+ languages, including several Russian ones.
- Voicing notifications and bot replies
- Voice for offline devices
- Voice menus
- Sizes
- about 5M – 30M
- Hardware
- from: Laptop
Text to speechGGUF2025
Canopy Labs · USA
Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.
- Real-time voice for an assistant
- Emotional voiceover
- Voice cloning
- Sizes
- 3B
- Hardware
- from: Laptop
Text to speechGGUF2025
Sesame · USA
A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.
- Voice for a conversational assistant
- Voicing dialogues
- Voice product prototypes
- Sizes
- 1B
- Hardware
- from: Laptop
Text to SQL2025
Renmin University of China (RUC) · China
Models for turning questions into SQL, trained on millions of synthetic query examples across different databases. Three sizes for different hardware.
- Database questions without knowing SQL
- Generating queries for reports
- A base for fine-tuning on your own database schema
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer-use agents2024–2025
Salesforce · USA
Salesforce models for function calling and agents: they pick the right tool and fill in its parameters. Strong on benchmarks, but the license is non-commercial.
- Calling APIs and internal services on user request
- Multi-step agents with several tools
- Comparing approaches before choosing a commercial model
- Sizes
- 1B – 8x22B
- Hardware
- from: Laptop
Computer vision2023–2025
Google · USA
Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.
- Image search by text query
- Automatic catalog labeling and tagging
- Filtering prohibited content
- Sizes
- about 0.2B to 2B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
hexgrad (independent developer) · not disclosed
A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.
- Voicing articles and notifications
- Voice for apps without a GPU
- Bulk text voiceover
- Sizes
- 82M
- Hardware
- from: Laptop
Computer-use agents2024–2025
Microsoft · USA
Breaks a screenshot down into buttons, fields and icons with labels so a regular language model can understand and control the screen. It does not click itself; it serves as the agent's eyes.
- Mapping legacy software screens for automation
- Preparing an agent to work in an interface
- Checking that the required elements are on screen
- Sizes
- under 1B (detector + captioning)
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
Microsoft Research · USA
An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.
- Pilots in interface control
- Research projects spanning screens and robotics
- Analyzing screenshots with an action plan
- Sizes
- 8B
- Hardware
- from: 1 GPU
Search and RAGRU2023–2025
Alibaba · China
Alibaba embeddings and rerankers for search: from tiny to 7B based on Qwen2. There is a multilingual mGTE version with long context.
- Semantic search across documents
- Reranking search results
- Clustering and classifying texts
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Text analysisGGUF2024–2025
Answer.AI and LightOn · USA / France
A modern replacement for classic BERT: faster, reads up to 8 thousand tokens at once. A base for your own classifiers. Trained on English and code; for Russian there is RuModernBERT.
- Classifying requests and documents
- Finding relevant passages in long texts
- Base for your own classifier after fine-tuning
- Sizes
- 150M – 395M
- Hardware
- from: Laptop
Search and RAGRUOllama2019–2025
UKP Lab (TU Darmstadt), later Hugging Face · Germany
The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.
- Search across a knowledge base and FAQ
- Finding similar tickets and duplicates
- Grouping reviews and requests by topic
- Sizes
- about 20M – 470M
- Hardware
- from: Laptop
Text analysisRUOllama2024–2025
Jina AI · Germany
Small models that turn raw web page HTML into clean Markdown or JSON. Handy for preparing websites for a knowledge base. Non-commercial license only.
- Cleaning website pages for a knowledge base
- Extracting data from pages into JSON
- Preparing texts for RAG
- Sizes
- 0.5B – 1.5B
- Hardware
- from: Laptop
Fact-checking and judges2025
Atla · UK
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring chatbot answers against your own criteria
- Comparing two versions of a prompt or model
- Filtering out weak answers before they reach a person
- Sizes
- 8B
- Hardware
- from: 1 GPU
Search and RAGRUOllama2024
Snowflake · USA
Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.
- Search across documents and knowledge bases
- Picking passages for RAG
- Search across reports and internal data
- Sizes
- 22M – 568M
- Hardware
- from: Laptop
Visual document search2024
Alibaba (Tongyi Lab) · China
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
- Finding a product by photo
- Search across a catalogue of images and cards
- Search across document pages as images
- Sizes
- 2B and 7B
- Hardware
- from: 1 GPU
Fact-checking and judges2024
Patronus AI · USA
A small judge: it scores against your criteria and highlights which part of the answer led to that score. The license is non-commercial. The judge itself makes mistakes and does not replace manual review.
- Scoring answers against your criteria with an explanation
- Understanding why a score was lowered
- Bulk review of assistant conversations
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
CodeOllama2024
INF Technology · China
Fully reproducible coding models: along with the weights, the data, its cleaning pipeline and the training recipe are open. Understand English and Chinese.
- Code generation and completion
- Training your own coding model from an open recipe
- A programming assistant on low-end hardware
- Sizes
- 1.5B – 8B
- Hardware
- from: Laptop
TextOllama2024
Nexusflow · USA
Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.
- Research on agents and function calling
- Comparison with commercial models
- Experiments with a chat assistant
- Sizes
- 70B – 72B
- Hardware
- from: 1 GPU
Visual document search2024
LightOn · France
A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.
- Refining search results over scans and PDFs
- Selecting pages before an AI assistant answers
- Sorting retrieved slides and reports
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Forecasting2024
Auton Lab, Carnegie Mellon University · USA
A foundation model for numeric series: one engine is used for forecasting, anomaly detection, filling gaps and classification.
- Forecasting demand and load
- Detecting anomalies in sensor readings and metrics
- Filling gaps in historical data
- Sizes
- about 40M – 385M
- Hardware
- from: Laptop
Forecasting2024
IBM Research · USA
Tiny forecasting models from IBM: they run on an ordinary CPU and sit next to the business system without a separate GPU server.
- Forecasting sales and warehouse stock
- Forecasting energy use and equipment load
- Fast forecasts right on the company server
- Sizes
- very small: TinyTimeMixers have about 1M parameters
- Hardware
- from: Laptop
CodeOllama2023–2024
DeepSeek · China
DeepSeek's coding model family: from small autocompletion models to the large MoE V2, which matched closed models in 2024. Later, coding moved into DeepSeek's general models.
- Code autocompletion and generation
- Translating code between programming languages
- Finding bugs and explaining other people's code
- Sizes
- 1.3B – 236B-A21B
- Hardware
- from: Laptop
CodeOllama2024
01.AI · China
Coding models from 01.AI at 1.5B and 9B with a 128K-token context and support for 52 programming languages. A separate line next to the text Yi models.
- Code autocompletion and generation
- Explaining and refactoring code
- A programming assistant without the cloud
- Sizes
- 1.5B – 9B
- Hardware
- from: Laptop
Fact-checking and judges2024
Flow AI · not disclosed
A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.
- Checking AI assistant answers against the instruction
- Bulk scoring of exported conversations
- Quality control before rolling out changes
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Forecasting2024
The Time-MoE team · not disclosed
A forecasting model with a sparse architecture: only part of the network runs at each step, so it stays fast at a small size.
- Forecasting sales and stock levels
- Forecasting load on services and staff
- Planning purchases from history
- Sizes
- 50M and 200M
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Stability AI · UK
Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.
- A lightweight chatbot on an ordinary PC
- Code autocompletion in the editor
- A base for fine-tuning on your own task
- Sizes
- 1.6B – 12B
- Hardware
- from: Laptop
CodeOllamaNot maintained2024
Mistral AI · France
Mistral's coding model covering 80+ programming languages. The open weights of the main version cannot be used in production without a paid license; newer Codestral versions are API-only.
- Evaluation and testing before buying a license
- Code autocompletion (with a commercial license)
- Research on coding model quality
- Sizes
- 7B – 22B
- Hardware
- from: Laptop
Text analysisRUNot maintained2020–2024
SberDevices (ai-forever) · Russia
Sber's Russian-language encoders trained on large Russian corpora. A base for classifiers, NER and semantic search in Russian.
- Classifying requests in Russian
- Extracting names, amounts and dates after fine-tuning
- Detecting review sentiment
- Sizes
- about 30M to 430M
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
Zhipu AI (Z.ai) and Tsinghua University · China
Coding models from the creators of GLM. CodeGeeX4-ALL-9B, based on GLM-4-9B, combines autocompletion, code chat, function calling and repository search in one model.
- Code autocompletion in the IDE
- A code chat assistant
- Answering questions about a repository
- Sizes
- 6B – 9B
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
BAAI (Beijing Academy of Artificial Intelligence) · China
Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.
- Refining search results before a chatbot answers
- Sorting knowledge base search results
- Selecting the most relevant clauses of contracts and policies
- Sizes
- 278M – 9B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
Patronus AI · USA
Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.
- Finding invented facts in AI assistant answers
- Checking that answers rest on the attached documents
- Filtering out answers before they go to a customer
- Sizes
- 8B and 70B
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
OpenChat (Tsinghua University) · China
Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.
- Chat assistant on an inexpensive server
- Drafts of emails and replies
- Help with simple code
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2023–2024
Defog · USA
One of the first open models that turn a plain-language question into an SQL query against a database. Available in Ollama, but newer competitors are already stronger.
- Answering managers' questions from the sales database without an analyst
- Drafting SQL queries for reports
- An assistant inside a BI system
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
RLHFlow · USA
An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data for model fine-tuning
- Scoring assistant answers across several attributes
- Sizes
- 8B
- Hardware
- from: 1 GPU
CodeOllamaNot maintained2023–2024
BigCode (Hugging Face and ServiceNow) · USA / France
One of the first open coding models, trained on an open set of source code with an option to exclude your own repository. Today it is more a base for fine-tuning than a leader.
- Code autocompletion in the editor
- Fine-tuning on the company's internal code
- Generating boilerplate code and tests
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
KAIST and LG AI Research (prometheus-eval) · South Korea
An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.
- Scoring chatbot answers on your own scale
- Comparing two answer options
- Quality checks before launching an AI service
- Sizes
- 7B – 8x7B
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2024
Chat2DB · China
A text-to-SQL model from the open Chat2DB database client: it supports different SQL dialects, with an English and Chinese model card.
- Turning a question into SQL inside a database client
- Drafting queries for different database engines
- Hints for developers working with a schema
- Sizes
- 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
CybersecurityGGUFNot maintained2023–2024
ZySec AI · India
A small open assistant for security professionals: questions about standards, reviewing threats and vulnerabilities, drafting internal documents.
- Answering questions about security policies and standards
- First-pass review of threat reports
- Drafting internal protection guidelines
- Sizes
- 2.8B и 7B
- Hardware
- from: Laptop
Search and RAGRUNot maintained2022–2024
Microsoft · USA
Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.
- Search across a knowledge base and documents
- Finding answers for a chatbot (RAG)
- Finding similar requests and duplicates
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2024
ChatDB · USA
A text-to-SQL model built on DeepSeek-Coder, aimed at complex questions spanning several tables and conditions.
- Complex queries joining several tables
- Answering database questions without an analyst
- Drafting SQL for reports and exports
- Sizes
- 7B
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
Meta · USA
A version of Llama 2 further trained on code, with variants for Python and for chat. Outdated, but many ready-made fine-tuned versions and tools exist.
- Code autocompletion and explanation
- Generating Python scripts
- Base model for fine-tuning on your own stack
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
WizardLM (Microsoft and Peking University) · USA / China
Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.
- Complex multi-step instructions
- Help for developers
- Solving math problems
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2024
MotherDuck and Numbers Station · USA
A model for turning questions into SQL, built for the embedded analytics database DuckDB. Available in Ollama, convenient for working with CSV and Parquet locally.
- Plain-language questions about CSV and Parquet exports
- DuckDB queries inside analytics scripts
- Quick analytics on a laptop without a server
- Sizes
- 7B
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
University of Waterloo, Castorini group · Canada
Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.
- Reordering a long list of search results
- Selecting sources for an AI assistant answer
- Research comparisons of retrieval approaches
- Sizes
- 7B – 13B
- Hardware
- from: 1 GPU
RerankersNot maintained2023
NetEase Youdao · China
An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.
- Search across a knowledge base and reference materials
- Reordering retrieved passages
- Picking answers for a support chatbot
- Sizes
- about 280M
- Hardware
- from: Laptop
Text to speechNot maintained2023
Columbia University · USA
A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.
- Voicing texts in English
- A base for fine-tuning your own voice
- Voice service prototypes
- Sizes
- about 150M
- Hardware
- from: Laptop
Text to SQLNot maintained2023
RUCKBReasoning, Renmin University of China · China
An early line of open text-to-SQL models starting at 1B, including variants fine-tuned for specific database schemas.
- Turning an employee question into an SQL query
- Drafting warehouse queries for a report
- Embedding into a BI dashboard as a helper
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
Text to SQLGGUFNot maintained2023
Numbers Station · USA
One of the first open text-to-SQL lines, including very small versions from 350M that run on an ordinary PC.
- Turning a question into SQL from a table description
- Hints while writing queries
- A local analyst helper with no data leaving the company
- Sizes
- 350M – 7B
- Hardware
- from: Laptop
RerankersRUNot maintained2022
UKP Lab and the Sentence Transformers community · Germany
The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.
- Reordering knowledge base search results
- Selecting passages before a chatbot answers
- Finding duplicates among tickets and product cards
- Sizes
- about 4M – 120M
- Hardware
- from: Laptop
Text analysisRUNot maintained2019–2022
Meta · USA
A classic multilingual encoder for 100 languages, including Russian. The base of many sentiment, NER and embedding models, including BGE-M3.
- Detecting review sentiment in different languages
- Extracting names and organizations after fine-tuning
- Classifying requests
- Sizes
- 270M – 10.7B
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
Microsoft · USA
A time-tested encoder behind many classifiers and NER models (including GLiNER). The multilingual mDeBERTa-v3 understands Russian.
- Classifying review sentiment
- Entity extraction after fine-tuning
- Checking whether a conclusion follows from a text
- Sizes
- 70M – 435M
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
David Dale (cointegrated) · Russia
A very small Russian-English BERT that runs fast on a regular CPU. Ready-made fine-tuned versions exist for sentiment, toxicity and emotions.
- Detecting review sentiment
- Filtering rude chat messages
- Fast classification of requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop