AI models for software development

For development teams, open models help write and review code, generate tests and documentation, search the codebase and automate routine work with agents. Local deployment keeps source code away from external services. Check support for your languages and frameworks, context length, the license and GPU requirements.

165 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
TextOllama2023–2026

DeepSeek

DeepSeek · China

DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.

  • Employee assistant on your own server
  • Analysis of long contracts and reports
  • Agents that work with tools and APIs
Sizes
7B – 1.6T-A49B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2026

InternLM / Intern-S

Shanghai AI Laboratory · China

Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.

  • Research assistant: papers, formulas, data
  • Analysis of scientific and technical documents
  • Corporate chat on small models
Sizes
1.8B – about 1T
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

MiMo

Xiaomi · China

Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.

  • Logic and calculation tasks
  • Agents with tools
  • Help for developers
Sizes
7B – 1,02T-A42B
Hardware
from: Laptop
Commercial use allowedDetails
Text2024–2026

MiniCPM

OpenBMB (ModelBest and Tsinghua University) · China

Compact text models that run directly on a device: laptop, phone or mini PC. The 1B and 2B MiniCPM5 models focus on tool calling and long context.

  • A local chat assistant without the cloud
  • Data extraction and text classification
  • Tool calling and simple agents on low-end hardware
Sizes
0.5B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Text2024–2026

K2 (K2-Think, K2-V2, K2-Horizon)

MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE

Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.

  • Reasoning, maths and technical questions
  • Analysing long documents
  • Agents and writing code
Sizes
0.9B – 375B-A23B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

LLaDA

Renmin University of China (GSAI) and Ant Group (inclusionAI) · China

Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.

  • Fast generation of code and text
  • Agent scenarios with long context
  • Research into alternatives to standard LLMs
Sizes
8B – 100B (MoE)
Hardware
from: 1 GPU
Commercial use allowedDetails
TextRU2026

Zarya (ai-forever)

SberDevices (ai-forever) · Russia

A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.

  • Experiments with faster generation
  • Fine-tuning small models for your own tasks
  • Research
Sizes
0.6B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2023–2026

Qwen

Alibaba · China

A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.

  • Chatbot and knowledge-base assistant
  • Replies to emails and customer requests
  • Document parsing and classification
Sizes
0,6B – 2,4T-A95B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRU2024–2026

FRIDA / Giga-Embeddings

Sber (SberDevices) · Russia

Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.

  • Search across Russian-language documents
  • RAG for chatbots in Russian
  • Classifying requests and reviews
Sizes
480M – 10B-A1.8B
Hardware
from: Laptop
Commercial use allowedDetails
Speech to textGGUF2024–2026

Moonshine

Moonshine AI (Useful Sensors) · USA

Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.

  • Voice control of devices
  • Offline recognition on a phone
  • Live subtitles
Sizes
27M – 245M
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2026

GLM (ChatGLM)

Zhipu AI (Z.ai) · China

One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.

  • Corporate chat assistant
  • Agents for routine office tasks
  • Help for developers
Sizes
1.5B – 744B-A40B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text2024–2026

Hunyuan / Hy

Tencent · China

Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.

  • Corporate assistant
  • Translation and multilingual texts
  • Agents with tools
Sizes
0.5B – 770B-A49B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text2025–2026

Ling / Ring

Ant Group (inclusionAI) · China

An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.

  • Corporate assistant
  • Agents for office processes
  • Financial analytics (Fin version available)
Sizes
7.9B-A1.3B – 1T
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2024–2026

NVIDIA Nemotron

NVIDIA · USA

NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).

  • Agents with tool calling
  • Reasoning and calculation tasks
  • Answers based on long documents
Sizes
4B – 550B-A55B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024–2026

IBM Granite

IBM · USA

IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.

  • Answers based on internal documents (RAG)
  • Tool calling and agent work
  • Data extraction and classification
Sizes
350M – 34B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2026

Muse Glimmer

Meta Superintelligence Labs · USA

An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.

  • Agents with tool calling
  • Analysis of screenshots, charts and documents
  • Multilingual assistant
Sizes
30B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text analysis2024–2026

GLiNER

Urchade Zaratiana and Fastino AI · France / USA

Finds the entities you need in text without training: just list what to look for (name, amount, date). GLiNER2 also classifies text. Multilingual versions understand Russian.

  • Extracting names, amounts and dates from emails and contracts
  • Parsing requests into CRM fields
  • Classifying requests by topic
Sizes
about 50M to 500M
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2025–2026

OpenCUA / Qwen-CUA

XLANG Lab (University of Hong Kong) · China

Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.

  • Working in desktop software without an API
  • Moving data between systems
  • Running user scenarios for tests
Sizes
7B – about 400B (MoE)
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2025–2026

UI-Venus

Ant Group (inclusionAI) · China

An Ant Group family for finding elements on screen and completing tasks in phone and computer interfaces. UI-Venus-2 was specifically trained to refuse dangerous actions.

  • Automating actions in mobile apps
  • Filling in forms in web interfaces
  • UI autotests
Sizes
2B – 72B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllama2026

Ornith

DeepReinforce · not disclosed

Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.

  • A developer agent in the terminal
  • Fixing bugs from a task description
  • Understanding and extending a large repository
Sizes
9B – 397B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextRUOllama2023–2026

Mistral

Mistral AI · France

European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).

  • Fast chat responses
  • Data extraction from text
  • Translation and multilingual work
Sizes
3B – 675B
Hardware
from: Laptop
Commercial use with conditionsDetails
Code2025–2026

KAT-Coder / KAT-Dev

Kwaipilot (Kuaishou) · China

Kuaishou models for agentic development, trained to solve real tasks in repositories. KAT-Coder-V2.5-Dev (35B, 3B active) is the open version of their closed flagship.

  • An agent that fixes tasks in the repository
  • Code generation and refactoring
  • Automating routine development tasks
Sizes
32B – 72B, 35B-A3B
Hardware
from: 1 GPU
Commercial use allowedDetails
Search and RAGRU2025–2026

NVIDIA Nemotron Embed

NVIDIA · USA

NVIDIA embeddings for search and RAG. Nemotron-3-Embed, released in 2026, is under the permissive OpenMDW license and works in many languages.

  • Search across corporate documents
  • RAG for chatbots and assistants
  • Search across images and pages (VL versions)
Sizes
1B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

Kimi

Moonshot AI · China

Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.

  • Multi-step agents: search, data collection, reports
  • In-depth document analysis
  • Help for developers
Sizes
16B-A3B – 2.8T-A104B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextOllama2024–2026

EXAONE

LG AI Research · South Korea

Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.

  • Corporate assistant
  • Working with Korean and English texts
  • Analysis of documents and images (4.5)
Sizes
1.2B – 750B-A37B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2026

Solar

Upstage · South Korea

Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.

  • Working with office documents
  • Agents for routine tasks
  • Help for developers
Sizes
10.7B – 250B-A15B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextGGUF2026

Inkling

Thinking Machines Lab · USA

Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.

  • Flagship-level corporate assistant
  • Analysis of documents, images and audio
  • Programming help
Sizes
276B-A12B, 975B-A41B
Hardware
from: Cluster
Commercial use allowedDetails
CodeGGUF2026

Laguna

Poolside · USA

Models for agentic programming: they edit code in a repository on their own. The small XS runs on a Mac with 36 GB of memory; S 2.1 has a 1M-token context.

  • Coding agent for in-house development
  • Bug fixing and code improvements
  • Working with large codebases
Sizes
33B-A3B – 225B-A23B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2025–2026

Fara

Microsoft · USA

Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.

  • Filling in web forms and applications
  • Collecting data from web portals without an API
  • Checking websites against scenarios
Sizes
4B – 27B
Hardware
from: Laptop
Commercial use allowedDetails
RerankersGGUF2024–2026

Jina Reranker

Jina AI · Germany

Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.

  • Refining search results before a chatbot answers
  • Sorting retrieved PDF pages and slides
  • Catalog and knowledge base search
Sizes
33M – 2.4B
Hardware
from: Laptop
Commercial use with conditionsDetails
Rerankers2026

R3 (R3-Embedding и R3-Rerank)

Tencent · China

A pair of small Tencent models based on Qwen3 that pick the right skill for an AI agent for a given request: the embedding model finds candidates, the reranker chooses the best one.

  • Choosing a tool or skill for an AI agent
  • Routing requests between bot scenarios
  • Search across a catalog of internal tools
Sizes
0.6B
Hardware
from: Laptop
Commercial use allowedDetails
CodeGGUF2025–2026

Mellum

JetBrains · Czech Republic

JetBrains models for fast code autocompletion. Mellum2 (12B, 2.5B active) is already a full assistant: it writes and edits code, calls tools and reasons.

  • Fast code autocompletion on your own server
  • A developer assistant that does not send code to the cloud
  • Fine-tuning on the company's code
Sizes
4B – 12B-A2.5B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningOllama2025–2026

OpenThinker

Open Thoughts (Stanford, Berkeley and other universities) · USA

Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Checking the logic of internal policies
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUGGUF2025–2026

MiniMax

MiniMax · China

Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.

  • Analysis of large document archives in a single request
  • Agents with tools
  • Help for developers
Sizes
230B-A10B – 456B-A46B
Hardware
from: Cluster
Commercial use with conditionsDetails
Search and RAGRUGGUF2023–2026

Jina Embeddings

Jina AI · Germany

Strong multilingual embeddings with long context; v5-omni understands text, images and audio. Recent versions are open for non-commercial use only.

  • Search across documents in many languages
  • Search across images and scans
  • Classification and clustering
Sizes
33M – 3.8B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2023–2026

Dolphin

Cognitive Computations (Eric Hartford) · USA

Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.

  • Assistant that does not refuse legal but sensitive topics
  • Internal tools under a strict system prompt
  • Role-play and creative scenarios
Sizes
0.5B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

Step

StepFun · China

StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.

  • High-load agents
  • Analysis of documents with diagrams and screenshots
  • Help for developers
Sizes
8B – 321B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2025–2026

Holo

H Company · France

A French model family for controlling a browser and computer: precisely finds the right element on screen and handles multi-step tasks. The latest Holo3 and 3.1 are open under Apache 2.0.

  • Working in web portals and legacy software without an API
  • Filling in forms and applications
  • Testing interfaces against scenarios
Sizes
0.8B – 235B-A22B
Hardware
from: Laptop
Commercial use with conditionsDetails
Rerankers2025–2026

Llama Nemotron Rerank

NVIDIA · USA

A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.

  • Reordering passages before an AI assistant answers
  • Sorting retrieved scan and PDF pages
  • Search across internal policies and instructions
Sizes
1B
Hardware
from: Laptop
Commercial use with conditionsDetails
Forecasting2024–2026

Timer (THUML)

THUML, Tsinghua University · China

A compact forecasting foundation model from the Tsinghua lab: trained on a large set of diverse series and fine-tunable on your own data.

  • Forecasting demand and load
  • Forecasting sensor readings on the shop floor
  • Fine-tuning forecasts on your own history
Sizes
84M (timer-base)
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2024–2026

Granite Embedding

IBM · USA

Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.

  • Search across corporate documents
  • RAG on a regular server without a GPU
  • Reranking results
Sizes
30M – 311M
Hardware
from: Laptop
Commercial use allowedDetails
ForecastingGGUF2025–2026

Toto

Datadog · USA

A Datadog forecasting model trained on server and application metrics. Especially strong for IT monitoring: load, latency, errors.

  • Server load forecasting
  • Anomaly detection in metrics
  • Capacity planning
Sizes
4M – 2.5B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUGGUF2025–2026

Arcee Trinity

Arcee AI · USA

An American family of MoE models trained from scratch: Nano, Mini and Large. Trinity-Large-Thinking (398B) reasons before answering.

  • Agents with tool calling
  • Reasoning tasks
  • Corporate assistant on your own servers
Sizes
6B – 398B-A13B
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safetyOllama2024–2026

Granite Guardian

IBM · USA

IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.

  • Checking bot requests and replies
  • Finding made-up facts in knowledge-base answers
  • Checking your own rules written as text
Sizes
38M – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2024–2026

ShowUI

Show Lab (National University of Singapore) · Singapore

A lightweight model for working with interfaces: finds buttons and fields by description and performs actions on the web and on a phone. ShowUI-π can drag with the mouse.

  • Clicking and filling in forms from a task description
  • Web UI autotests
  • An assistant on a low-end computer without the cloud
Sizes
2B (ShowUI), about 500M (ShowUI-π)
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speech2025–2026

Kyutai STT, TTS и Pocket TTS

Kyutai · France

Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.

  • Streaming speech transcription for voice bots
  • Voicing replies with minimal delay
  • Speech synthesis on a server without a GPU (Pocket TTS)
Sizes
100M (Pocket TTS) – 2.6B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

Apriel

ServiceNow · USA

ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.

  • A reasoning assistant for internal services
  • Tool calling and enterprise agents
  • Analysing screenshots and documents with images
Sizes
5B – 15B
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2025–2026

Rnj-1

Essential AI · USA

An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.

  • Writing and fixing code
  • A developer agent on a single GPU
  • Solving technical and scientific problems
Sizes
8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGGGUF2025–2026

Octen Embedding

Octen · USA / Singapore

Qwen3-Embedding models fine-tuned by the startup Octen for search in legal, financial and medical texts. As of January 2026 the 8B version topped the RTEB leaderboard.

  • Search across contracts and case law
  • Search across financial reports
  • Search across long documents up to 32K tokens
Sizes
0.6B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judgesRU2025–2026

POLLUX Judge (ai-forever)

SberDevices (ai-forever) · Russia

Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.

  • Automatic quality checks of Russian chatbot answers
  • Comparing several models before choosing one
  • Checking answers after fine-tuning
Sizes
4B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningGGUF2025–2026

Goedel-Prover

Princeton University · USA

Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.

  • Formal verification of mathematical workings
  • Verifying code correctness
  • Training
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGGGUF2026

pplx-embed

Perplexity · USA

Embeddings from the Perplexity search service. Some versions take into account the context of the whole document, not just a single fragment.

  • Search across large document collections
  • RAG that accounts for document context
  • Website and catalog search
Sizes
0.6B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
Voice: speakers and soundRU2026

FireRedVAD

FireRedTeam (Xiaohongshu) · China

A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.

  • Cutting recordings before speech recognition
  • Separating speech from music and singing in broadcasts and videos
  • Speech detection in voice bots
Sizes
compact, exact size not stated
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRU2026

Harrier (harrier-oss)

Microsoft · USA

Microsoft's 2026 multilingual embeddings with context up to 32K tokens; Russian is on the language list. The 270M and 0.6B versions run on a regular server, 27B is the most accurate.

  • Multilingual knowledge base search
  • Picking passages for RAG
  • Search across long documents
Sizes
270M – 27B
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2024–2026

Qwen Coder

Alibaba (Qwen team) · China

The broadest open coding family: from 0.5B for autocompletion to 480B for agents. Qwen3-Coder-Next (80B, 3B active) works as a developer agent on a single GPU.

  • Code autocompletion in the editor
  • An agent that edits code in the repository on its own
  • Writing and refining scripts, SQL and integrations
Sizes
0.5B – 480B-A35B
Hardware
from: Laptop
Commercial use allowedDetails
CodeGGUF2026

IQuest-Coder

IQuest Research · China

A family of coding models with standard and reasoning versions, including a Loop variant that runs through its layers a second time. Sizes from 7B to 40B.

  • Writing and refining code
  • Solving tasks with step-by-step reasoning
  • Agentic work with a repository
Sizes
7B – 40B
Hardware
from: Laptop
Commercial use with conditionsDetails
Code2026

SERA (Ai2 Open Coding Agents)

Ai2 (Allen Institute for AI) · USA

Fully open developer agents from Ai2: weights, data and training recipe are all public. Designed so a company can cheaply fine-tune the agent on its own repository.

  • An agent for fixing issues in code
  • Fine-tuning the agent on an internal repository
  • Automating small edits and tests
Sizes
8B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2024–2026

Moshi / Hibiki

Kyutai · France

A voice assistant that listens and speaks at the same time, with no delay for recognition and synthesis. Hibiki does simultaneous speech-to-speech translation between several European languages.

  • Real-time voice conversation partner
  • Simultaneous speech translation
  • Zero-latency voice interfaces
Sizes
2B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice assistantsGGUF2025–2026

MiniCPM-o

OpenBMB (ModelBest, Tsinghua University) · China

A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.

  • Voice assistant on your own server
  • Analyzing videos and documents
  • Voice answers about a camera image
Sizes
8B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2025–2026

GUI-Owl / Mobile-Agent

Alibaba (Tongyi Lab, X-PLUG) · China

Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.

  • Automating actions in mobile apps
  • Working in desktop software without an API
  • Testing apps against scenarios
Sizes
2B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2025–2026

Qwen3 Embedding / Reranker

Alibaba (Qwen) · China

Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.

  • Knowledge base search for RAG
  • Reranking results before answering
  • Search across scans, slides and screenshots
Sizes
0.6B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAG2026

Voyage 4 nano

Voyage AI (MongoDB) · USA

The only open model in the Voyage 4 line: its vectors are compatible with the paid larger versions, so you can start locally and move to the API later.

  • Document search on your own server
  • RAG for small knowledge bases
  • Finding similar texts
Sizes
about 340M
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2023–2026

Phi

Microsoft · USA

Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.

  • Assistant on a laptop or your own server
  • Reasoning and calculation tasks
  • Analysis of images and diagrams (vision versions)
Sizes
1.3B – 42B-A6.6B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2024–2026

INTELLECT

Prime Intellect · USA

Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.

  • Reasoning and math tasks
  • Programming help
  • Agents with tool calling
Sizes
10B – 106B-A12B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2026

EvoCUA

Meituan · China

Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.

  • Working in office and legacy software without an API
  • Moving data between systems
  • Running test scenarios
Sizes
8B – 32B
Hardware
from: 1 GPU
Commercial use allowedDetails
Cybersecurity2025–2026

Foundation-Sec

Cisco (Foundation AI) · USA

Cisco models for information security based on Llama 3.1 8B: analysis of vulnerabilities, threats and incidents. Can be deployed inside your own perimeter.

  • Analyzing vulnerability and threat reports
  • Helping SOC analysts during incidents
  • Mapping threats to MITRE ATT&CK
Sizes
8B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and soundRU2025–2026

Smart Turn

Daily (Pipecat) · USA

Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.

  • Voice bot does not interrupt the customer during pauses
  • Fast reply when the customer has really finished
  • An add-on to a standard speech detector in voice assistants
Sizes
8M (v3) – 580M (v1)
Hardware
from: Laptop
Commercial use allowedDetails
CodeRUOllama2025

Devstral

Mistral AI (with All Hands AI) · France

Mistral models for agentic development: they read the repository, edit files and run commands on their own. The 24B version fits on a single GPU.

  • A developer agent that fixes tickets from the tracker
  • Extending internal systems from a description
  • Automating routine code edits
Sizes
24B – 123B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer-use agents2023–2025

CogAgent / AutoGLM

Zhipu AI (Z.ai) and Tsinghua University · China

One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.

  • Automating actions in mobile apps
  • Working in web interfaces without an API
  • Testing apps against scenarios
Sizes
9B – 18B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer-use agentsGGUF2025

MAI-UI

Alibaba (Tongyi-MAI) · China

Compact Alibaba models for working in smartphone and computer interfaces: they find elements and complete multi-step tasks. The small size allows running on an ordinary GPU.

  • Automating actions in mobile apps
  • Working in software without an API
  • UI autotests
Sizes
2B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safety2025

AprielGuard

ServiceNow · USA

A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.

  • Screening chatbot requests for attacks and jailbreaks
  • Filtering harmful model answers
  • Monitoring the actions of AI agents that use tools
Sizes
8B
Hardware
from: Laptop
Commercial use allowedDetails
Tabular dataGGUF2024–2025

TableGPT2 / TableGPT-R1

Zhejiang University · China

A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.

  • Answering questions about tables and data exports
  • Automated data analysis with generated code
  • A helper for BI and internal reporting
Sizes
7B – 72B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2023–2025

Hermes

Nous Research · USA

Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.

  • Agents that call functions and APIs
  • Data extraction in strict JSON format
  • Assistant with flexible role and tone settings
Sizes
3B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice: speakers and soundRU2020–2025

Silero VAD

Silero · Russia

The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.

  • Cutting calls and recordings before speech recognition
  • Detecting when the customer is speaking in a voice bot
  • Filtering out silence and noise to save on transcription
Sizes
about 2 MB
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2025

Cogito

Deep Cogito · USA

Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.

  • A chat assistant with a reasoning mode
  • Writing code and calling tools
  • Answering complex questions about documents
Sizes
3B – 671B
Hardware
from: Laptop
Commercial use with conditionsDetails
Rerankers2025

zerank (ZeroEntropy)

ZeroEntropy · USA

Rerankers built on Qwen3. The model card lists the target domains — finance, law, code, medicine, science; the stated language is English.

  • Refining results before an AI assistant answers
  • Sorting search results across contracts and reports
  • Search across technical and scientific documentation
Sizes
zerank-2 — 4B (based on Qwen3-4B), plus a smaller "small" version
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2024–2025

mxbai (Mixedbread Embed и Rerank)

Mixedbread · Germany

Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.

  • Search across a knowledge base
  • Reranking results before a bot answers
  • Product catalog search
Sizes
17M – 1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2025

ModernVBERT / ColModernVBERT

Illuin Technology, EPFL, CentraleSupélec · France

A compact (250M) model for searching document pages as images. According to the authors, it matches models 10 times larger and runs without a GPU.

  • Search across scans and PDFs on a modest server
  • Indexing document archives
  • Search across slides and manuals
Sizes
250M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGOllama2025

EmbeddingGemma

Google · USA

A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.

  • On-device document search
  • RAG without sending data outside
  • Text classification
Sizes
300M
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGOllama2023–2025

BGE (BAAI General Embedding)

BAAI (Beijing Academy of Artificial Intelligence) · China

Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.

  • Search across English-language documents
  • Picking passages for chatbot answers (RAG)
  • Code search (bge-code)
Sizes
24M – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQLGGUF2024–2025

Prem-1B-SQL

Prem AI · UK

A text-to-SQL model of just 1B parameters, designed to run locally so the database never leaves for external services.

  • Local translation of questions into SQL with no internet access
  • Query hints on modest hardware
  • Embedding into internal analytics tools
Sizes
1B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2025

gpt-oss

OpenAI · USA

OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.

  • AI agent that calls internal systems
  • Answers based on internal policies
  • Drafts of emails and reports
Sizes
20B, 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
Math and reasoningGGUF2025

Kimina-Prover

Moonshot AI and Project Numina · China, France

Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.

  • Formal verification of mathematical workings
  • Translating a problem from plain language into Lean
  • Training and olympiad preparation
Sizes
0.6B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025

Seed-OSS

ByteDance · China

An open ByteDance 36B model with up to 512K tokens of context and an adjustable thinking budget. Fits on a single powerful GPU.

  • Analysis of long documents
  • Agents with tools
  • Corporate assistant
Sizes
36B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text2024–2025

Grok (открытые веса)

xAI · USA

xAI publishes the weights of previous Grok generations. The models are very large and need a GPU cluster, so in practice they are rarely run.

  • Research on large models
  • Assistant on your own infrastructure
  • Text generation and analysis
Sizes
314B (Grok-1), Grok-2 is larger
Hardware
from: Cluster
Commercial use with conditionsDetails
Voice: speakers and sound2025

TEN VAD

Agora (TEN project) · USA / China

A lightweight speech detector for real-time voice assistants: it notices the start and end of a phrase faster than Silero VAD. Runs on servers, phones and in the browser.

  • Zero-lag speech detection in a voice bot
  • Fast assistant response at the end of a phrase
  • Use in mobile apps and the browser
Sizes
very small, the library is smaller than Silero VAD
Hardware
from: Laptop
Commercial use with conditionsDetails
Math and reasoningOllama2025

DeepScaleR, DeepCoder, DeepSWE

Agentica (Berkeley, Sky Computing Lab) and Together AI · USA

Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.

  • Solving maths problems with step-by-step working
  • Generating and checking code
  • An agent for fixing bugs in a repository
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2024–2025

ColPali / ColQwen

Illuin Technology (ViDoRe team) · France

Searches PDFs and scans as images: pages do not need to be OCR'd first, the model finds the right one for a question directly, including tables and charts. Trained on English.

  • Search across scans, presentations and PDFs
  • RAG over documents with tables and charts
  • Search across technical documentation
Sizes
256M – 3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judges2024–2025

CompassJudger / CompassVerifier

OpenCompass (Shanghai AI Laboratory) · China

A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.

  • Scoring model answers against set criteria
  • Checking an answer against a reference solution
  • Comparing several models on your own data
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoning2025

AceMath / AceReason

NVIDIA · USA

NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Working through programming problems
Sizes
1.5B – 72B
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judges2024–2025

Skywork-Reward

Skywork (Kunlun Tech) · China

Reward models: they score how good a language model's answer is for the user. Used for fine-tuning your own models and picking the best of several answers.

  • Choosing the best of several bot answers
  • Scoring answer quality during model fine-tuning
  • Comparing models before rollout
Sizes
0.6B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
CybersecurityGGUF2023–2025

SecGPT

Clouditera · China

A Chinese open family for cybersecurity: reviewing vulnerabilities, analysing logs and traffic, explaining commands and scripts.

  • Reviewing vulnerabilities and drafting fix recommendations
  • Analysing logs and reconstructing an attack chain
  • Explaining suspicious commands and scripts
Sizes
1.5B – 14B
Hardware
from: Laptop
Commercial use allowedDetails
CybersecurityGGUF2025

Trendyol Cybersecurity LLM

Trendyol · Turkey

Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.

  • Reviewing alerts and first-pass incident assessment
  • Explaining suspicious activity in reports
  • Helping the on-duty shift of a monitoring centre
Sizes
32B и 70B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text to SQLGGUF2025

SQL-R1

IDEA Research · China

A text-to-SQL model trained with reinforcement learning: it works through the schema and the conditions step by step before producing a query.

  • Database queries for questions with several conditions
  • Reviewing and fixing other people SQL queries
  • An analyst helper inside a BI system
Sizes
3B – 14B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2025

DeepSeek-R1

DeepSeek · China

A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.

  • Complex calculations and logic checks
  • Analysis of contracts and internal policies
  • Help for developers
Sizes
1,5B – 671B
Hardware
from: Laptop
Commercial use allowedDetails
CodeGGUF2025

Seed-Coder

ByteDance Seed · China

A compact 8B coding model from ByteDance in base, instruct and reasoning versions. Its training data was selected by the model itself, with almost no hand-written rules.

  • Code autocompletion and generation
  • Solving algorithmic problems
  • A base for fine-tuning on your own stack
Sizes
8B
Hardware
from: Laptop
Commercial use allowedDetails
Image generationGGUF2025

BAGEL

ByteDance Seed · China

A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.

  • Photo editing in a conversation
  • Answering questions about an image
  • Image generation with explanations
Sizes
14B-A7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text analysisRU2023–2025

RuModernBERT и USER (deepvk)

deepvk (VK) · Russia

Russian encoders from the VK team: RuModernBERT reads long texts, USER produces vectors for search, GeRaCl classifies texts by topic without training.

  • Classifying requests without labeled data
  • Knowledge base search in Russian
  • Analyzing long contracts
Sizes
35M – 360M
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQL2025

Arctic-Text2SQL-R1

Snowflake · USA

A Snowflake model for turning questions into SQL, trained with reinforcement learning by checking query results. The open 7B version is based on Qwen2.5-Coder.

  • Plain-language questions to a data warehouse
  • Generating SQL for reports and dashboards
  • Checking and fixing analysts' queries
Sizes
7B
Hardware
from: Laptop
Commercial use allowedDetails
CodeRU2025

Kodify-Nano (МТС AI)

MTS AI (MWS AI) · Russia

A small coding assistant from MTS AI that understands requests in Russian. Runs locally, with plugins for VS Code and JetBrains.

  • Code suggestions and completion in the editor
  • Code explanations in Russian
  • Drafts of tests and documentation
Sizes
1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Rerankers2023–2025

ColBERT (поиск с поздним взаимодействием)

Stanford NLP, later Answer.AI and LightOn · USA and France

A different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.

  • Search across a knowledge base of long documents
  • Reordering retrieved passages
  • Search across policies and technical documentation
Sizes
about 33M – 150M
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judgesRU2024–2025

Nemotron Reward / GenRM

NVIDIA · USA

Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.

  • Choosing the best of several candidate answers
  • Preparing data to fine-tune your own model
  • Scoring assistant answers in Russian and other languages
Sizes
49B, 70B and 340B
Hardware
from: Cluster
Commercial use with conditionsDetails
Forecasting2025

Sundial

THUML, Tsinghua University · China

A forecasting model that returns a set of possible scenarios rather than a single line — useful when you need a range for demand or load, not one number.

  • Forecasting demand with a range of values
  • Planning stock while accounting for spread
  • Forecasting load on services and staff
Sizes
128M (sundial-base)
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2025

Llama

Meta · USA

The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.

  • Assistant for employees
  • Summaries of meetings and documents
  • Base for industry-specific fine-tuning
Sizes
1B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Math and reasoningGGUF2024–2025

DeepSeek-Prover

DeepSeek · China

DeepSeek models for formal proofs in Lean 4: the proof is checked by a program, not a person. A narrow tool for mathematicians and engineers.

  • Formal verification of mathematical workings
  • Verifying algorithm correctness
  • Training and olympiad preparation
Sizes
7B – 671B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safety2024–2025

Prompt Guard

Meta · USA

Tiny classifiers that catch attempts to hack a bot: prompt injections and rule bypassing. The 86M version is multilingual, 22M is English only.

  • Protecting a bot from prompt injections
  • Checking emails and documents that reach an AI agent
  • Fast filter in front of a large model
Sizes
22M – 86M
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer-use agentsGGUF2025

UI-TARS

ByteDance Seed · China

A model that looks at a screenshot and controls the mouse and keyboard itself: clicks, fills in fields, navigates menus. The first generation and 1.5-7B are open; UI-TARS-2 weights were not released.

  • Working in legacy software without an API
  • Filling in forms and moving data between systems
  • UI autotests from plain-language scenarios
Sizes
2B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQL2025

XiYanSQL-QwenCoder

Alibaba · China

Alibaba models for turning questions into SQL, based on Qwen2.5-Coder. They work with different SQL dialects; a small 3B version suits modest hardware.

  • Plain-language database questions
  • Queries for different databases (PostgreSQL, MySQL, SQLite)
  • Automating routine reports
Sizes
3B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
CybersecurityGGUF2023–2025

WhiteRabbitNeo / DeepHat

Kindo · USA

One of the best-known open families for security and DevSecOps work: reviewing code for weaknesses, test scenarios, explaining attacks.

  • Finding weak spots in code and configurations
  • Reviewing incidents and explaining attack techniques
  • Drafting scripts and procedures for the security team
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image generationGGUF2024–2025

SANA

NVIDIA · USA

NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.

  • Bulk image generation
  • High-resolution visuals
  • Real-time generation inside apps
Sizes
0.6B – 4.8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2024–2025

Nomic Embed

Nomic AI · USA

Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.

  • Search across documents and a knowledge base
  • Code search
  • Search across scans and PDFs without text recognition
Sizes
137M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechRUGGUF2023–2025

Piper

Rhasspy / Open Home Foundation · USA

Very fast speech synthesis that runs even on a Raspberry Pi. Ready-made voices in 35+ languages, including several Russian ones.

  • Voicing notifications and bot replies
  • Voice for offline devices
  • Voice menus
Sizes
about 5M – 30M
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speechGGUF2025

Orpheus TTS

Canopy Labs · USA

Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.

  • Real-time voice for an assistant
  • Emotional voiceover
  • Voice cloning
Sizes
3B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechGGUF2025

Sesame CSM

Sesame · USA

A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.

  • Voice for a conversational assistant
  • Voicing dialogues
  • Voice product prototypes
Sizes
1B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQL2025

OmniSQL

Renmin University of China (RUC) · China

Models for turning questions into SQL, trained on millions of synthetic query examples across different databases. Three sizes for different hardware.

  • Database questions without knowing SQL
  • Generating queries for reports
  • A base for fine-tuning on your own database schema
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2024–2025

xLAM

Salesforce · USA

Salesforce models for function calling and agents: they pick the right tool and fill in its parameters. Strong on benchmarks, but the license is non-commercial.

  • Calling APIs and internal services on user request
  • Multi-step agents with several tools
  • Comparing approaches before choosing a commercial model
Sizes
1B – 8x22B
Hardware
from: Laptop
Non-commercial onlyDetails
Computer vision2023–2025

SigLIP (наследник CLIP)

Google · USA

Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.

  • Image search by text query
  • Automatic catalog labeling and tagging
  • Filtering prohibited content
Sizes
about 0.2B to 2B
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechGGUF2024–2025

Kokoro

hexgrad (independent developer) · not disclosed

A tiny speech synthesis model (82M) that sounds on par with large ones. Runs on a regular CPU; English and a few other languages, no Russian.

  • Voicing articles and notifications
  • Voice for apps without a GPU
  • Bulk text voiceover
Sizes
82M
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2024–2025

OmniParser

Microsoft · USA

Breaks a screenshot down into buttons, fields and icons with labels so a regular language model can understand and control the screen. It does not click itself; it serves as the agent's eyes.

  • Mapping legacy software screens for automation
  • Preparing an agent to work in an interface
  • Checking that the required elements are on screen
Sizes
under 1B (detector + captioning)
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer-use agentsGGUF2025

Magma

Microsoft Research · USA

An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.

  • Pilots in interface control
  • Research projects spanning screens and robotics
  • Analyzing screenshots with an action plan
Sizes
8B
Hardware
from: 1 GPU
Commercial use allowedDetails
Search and RAGRU2023–2025

GTE

Alibaba · China

Alibaba embeddings and rerankers for search: from tiny to 7B based on Qwen2. There is a multilingual mGTE version with long context.

  • Semantic search across documents
  • Reranking search results
  • Clustering and classifying texts
Sizes
33M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisGGUF2024–2025

ModernBERT

Answer.AI and LightOn · USA / France

A modern replacement for classic BERT: faster, reads up to 8 thousand tokens at once. A base for your own classifiers. Trained on English and code; for Russian there is RuModernBERT.

  • Classifying requests and documents
  • Finding relevant passages in long texts
  • Base for your own classifier after fine-tuning
Sizes
150M – 395M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2019–2025

Sentence Transformers (all-MiniLM, paraphrase-multilingual)

UKP Lab (TU Darmstadt), later Hugging Face · Germany

The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.

  • Search across a knowledge base and FAQ
  • Finding similar tickets and duplicates
  • Grouping reviews and requests by topic
Sizes
about 20M – 470M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRUOllama2024–2025

ReaderLM (Jina)

Jina AI · Germany

Small models that turn raw web page HTML into clean Markdown or JSON. Handy for preparing websites for a knowledge base. Non-commercial license only.

  • Cleaning website pages for a knowledge base
  • Extracting data from pages into JSON
  • Preparing texts for RAG
Sizes
0.5B – 1.5B
Hardware
from: Laptop
Non-commercial onlyDetails
Fact-checking and judges2025

Atla Selene

Atla · UK

An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.

  • Scoring chatbot answers against your own criteria
  • Comparing two versions of a prompt or model
  • Filtering out weak answers before they reach a person
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Search and RAGRUOllama2024

Snowflake Arctic Embed

Snowflake · USA

Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.

  • Search across documents and knowledge bases
  • Picking passages for RAG
  • Search across reports and internal data
Sizes
22M – 568M
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2024

GME (General Multimodal Embedding)

Alibaba (Tongyi Lab) · China

One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.

  • Finding a product by photo
  • Search across a catalogue of images and cards
  • Search across document pages as images
Sizes
2B and 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Fact-checking and judges2024

GLIDER (Patronus)

Patronus AI · USA

A small judge: it scores against your criteria and highlights which part of the answer led to that score. The license is non-commercial. The judge itself makes mistakes and does not replace manual review.

  • Scoring answers against your criteria with an explanation
  • Understanding why a score was lowered
  • Bulk review of assistant conversations
Sizes
3.8B (based on Phi-3.5-mini)
Hardware
from: Laptop
Non-commercial onlyDetails
CodeOllama2024

OpenCoder

INF Technology · China

Fully reproducible coding models: along with the weights, the data, its cleaning pipeline and the training recipe are open. Understand English and Chinese.

  • Code generation and completion
  • Training your own coding model from an open recipe
  • A programming assistant on low-end hardware
Sizes
1.5B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024

Athene

Nexusflow · USA

Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.

  • Research on agents and function calling
  • Comparison with commercial models
  • Experiments with a chat assistant
Sizes
70B – 72B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Visual document search2024

MonoQwen2-VL (LightOn)

LightOn · France

A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.

  • Refining search results over scans and PDFs
  • Selecting pages before an AI assistant answers
  • Sorting retrieved slides and reports
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails
Forecasting2024

MOMENT

Auton Lab, Carnegie Mellon University · USA

A foundation model for numeric series: one engine is used for forecasting, anomaly detection, filling gaps and classification.

  • Forecasting demand and load
  • Detecting anomalies in sensor readings and metrics
  • Filling gaps in historical data
Sizes
about 40M – 385M
Hardware
from: Laptop
Commercial use allowedDetails
Forecasting2024

Granite Time Series (TinyTimeMixers, PatchTST)

IBM Research · USA

Tiny forecasting models from IBM: they run on an ordinary CPU and sit next to the business system without a separate GPU server.

  • Forecasting sales and warehouse stock
  • Forecasting energy use and equipment load
  • Fast forecasts right on the company server
Sizes
very small: TinyTimeMixers have about 1M parameters
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2023–2024

DeepSeek-Coder

DeepSeek · China

DeepSeek's coding model family: from small autocompletion models to the large MoE V2, which matched closed models in 2024. Later, coding moved into DeepSeek's general models.

  • Code autocompletion and generation
  • Translating code between programming languages
  • Finding bugs and explaining other people's code
Sizes
1.3B – 236B-A21B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllama2024

Yi-Coder

01.AI · China

Coding models from 01.AI at 1.5B and 9B with a 128K-token context and support for 52 programming languages. A separate line next to the text Yi models.

  • Code autocompletion and generation
  • Explaining and refactoring code
  • A programming assistant without the cloud
Sizes
1.5B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judges2024

Flow Judge

Flow AI · not disclosed

A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.

  • Checking AI assistant answers against the instruction
  • Bulk scoring of exported conversations
  • Quality control before rolling out changes
Sizes
3.8B (based on Phi-3.5-mini)
Hardware
from: Laptop
Commercial use allowedDetails
Forecasting2024

Time-MoE

The Time-MoE team · not disclosed

A forecasting model with a sparse architecture: only part of the network runs at each step, so it stays fast at a small size.

  • Forecasting sales and stock levels
  • Forecasting load on services and staff
  • Planning purchases from history
Sizes
50M and 200M
Hardware
from: Laptop
Commercial use allowedDetails
TextOllamaNot maintained2023–2024

StableLM и Stable Code

Stability AI · UK

Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.

  • A lightweight chatbot on an ordinary PC
  • Code autocompletion in the editor
  • A base for fine-tuning on your own task
Sizes
1.6B – 12B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllamaNot maintained2024

Codestral

Mistral AI · France

Mistral's coding model covering 80+ programming languages. The open weights of the main version cannot be used in production without a paid license; newer Codestral versions are API-only.

  • Evaluation and testing before buying a license
  • Code autocompletion (with a commercial license)
  • Research on coding model quality
Sizes
7B – 22B
Hardware
from: Laptop
Non-commercial onlyDetails
Text analysisRUNot maintained2020–2024

ruBERT, ruRoBERTa, ruELECTRA (ai-forever)

SberDevices (ai-forever) · Russia

Sber's Russian-language encoders trained on large Russian corpora. A base for classifiers, NER and semantic search in Russian.

  • Classifying requests in Russian
  • Extracting names, amounts and dates after fine-tuning
  • Detecting review sentiment
Sizes
about 30M to 430M
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllamaNot maintained2023–2024

CodeGeeX

Zhipu AI (Z.ai) and Tsinghua University · China

Coding models from the creators of GLM. CodeGeeX4-ALL-9B, based on GLM-4-9B, combines autocompletion, code chat, function calling and repository search in one model.

  • Code autocompletion in the IDE
  • A code chat assistant
  • Answering questions about a repository
Sizes
6B – 9B
Hardware
from: Laptop
Commercial use with conditionsDetails
RerankersGGUFNot maintained2023–2024

BGE Reranker

BAAI (Beijing Academy of Artificial Intelligence) · China

Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.

  • Refining search results before a chatbot answers
  • Sorting knowledge base search results
  • Selecting the most relevant clauses of contracts and policies
Sizes
278M – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judgesNot maintained2024

Patronus Lynx

Patronus AI · USA

Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.

  • Finding invented facts in AI assistant answers
  • Checking that answers rest on the attached documents
  • Filtering out answers before they go to a customer
Sizes
8B and 70B
Hardware
from: 1 GPU
Non-commercial onlyDetails
TextOllamaNot maintained2023–2024

OpenChat

OpenChat (Tsinghua University) · China

Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.

  • Chat assistant on an inexpensive server
  • Drafts of emails and replies
  • Help with simple code
Sizes
7B – 13B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to SQLOllamaNot maintained2023–2024

SQLCoder

Defog · USA

One of the first open models that turn a plain-language question into an SQL query against a database. Available in Ollama, but newer competitors are already stronger.

  • Answering managers' questions from the sales database without an analyst
  • Drafting SQL queries for reports
  • An assistant inside a BI system
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judgesNot maintained2024

ArmoRM (RLHFlow)

RLHFlow · USA

An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.

  • Choosing the best of several candidate answers
  • Preparing data for model fine-tuning
  • Scoring assistant answers across several attributes
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
CodeOllamaNot maintained2023–2024

StarCoder

BigCode (Hugging Face and ServiceNow) · USA / France

One of the first open coding models, trained on an open set of source code with an option to exclude your own repository. Today it is more a base for fine-tuning than a leader.

  • Code autocompletion in the editor
  • Fine-tuning on the company's internal code
  • Generating boilerplate code and tests
Sizes
1B – 15B
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judgesNot maintained2023–2024

Prometheus 2

KAIST and LG AI Research (prometheus-eval) · South Korea

An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.

  • Scoring chatbot answers on your own scale
  • Comparing two answer options
  • Quality checks before launching an AI service
Sizes
7B – 8x7B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQLGGUFNot maintained2024

Chat2DB-SQL

Chat2DB · China

A text-to-SQL model from the open Chat2DB database client: it supports different SQL dialects, with an English and Chinese model card.

  • Turning a question into SQL inside a database client
  • Drafting queries for different database engines
  • Hints for developers working with a schema
Sizes
7B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllamaNot maintained2023–2024

TinyLlama

TinyLlama (SUTD researchers) · Singapore

A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.

  • Simple chatbots on low-end hardware
  • Experiments and team training
  • A base for fine-tuning on a narrow task
Sizes
1.1B
Hardware
from: Laptop
Commercial use allowedDetails
CybersecurityGGUFNot maintained2023–2024

ZySec

ZySec AI · India

A small open assistant for security professionals: questions about standards, reviewing threats and vulnerabilities, drafting internal documents.

  • Answering questions about security policies and standards
  • First-pass review of threat reports
  • Drafting internal protection guidelines
Sizes
2.8B и 7B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUNot maintained2022–2024

E5 / multilingual-e5

Microsoft · USA

Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.

  • Search across a knowledge base and documents
  • Finding answers for a chatbot (RAG)
  • Finding similar requests and duplicates
Sizes
33M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQLGGUFNot maintained2024

Natural-SQL

ChatDB · USA

A text-to-SQL model built on DeepSeek-Coder, aimed at complex questions spanning several tables and conditions.

  • Complex queries joining several tables
  • Answering database questions without an analyst
  • Drafting SQL for reports and exports
Sizes
7B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllamaNot maintained2023–2024

Code Llama

Meta · USA

A version of Llama 2 further trained on code, with variants for Python and for chat. Outdated, but many ready-made fine-tuned versions and tools exist.

  • Code autocompletion and explanation
  • Generating Python scripts
  • Base model for fine-tuning on your own stack
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

WizardLM, WizardCoder, WizardMath

WizardLM (Microsoft and Peking University) · USA / China

Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.

  • Complex multi-step instructions
  • Help for developers
  • Solving math problems
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to SQLOllamaNot maintained2024

DuckDB-NSQL

MotherDuck and Numbers Station · USA

A model for turning questions into SQL, built for the embedded analytics database DuckDB. Available in Ollama, convenient for working with CSV and Parquet locally.

  • Plain-language questions about CSV and Parquet exports
  • DuckDB queries inside analytics scripts
  • Quick analytics on a laptop without a server
Sizes
7B
Hardware
from: Laptop
Commercial use with conditionsDetails
RerankersGGUFNot maintained2023–2024

RankVicuna / RankLLaMA / RankZephyr (Castorini)

University of Waterloo, Castorini group · Canada

Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.

  • Reordering a long list of search results
  • Selecting sources for an AI assistant answer
  • Research comparisons of retrieval approaches
Sizes
7B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
RerankersNot maintained2023

BCEmbedding (Youdao)

NetEase Youdao · China

An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.

  • Search across a knowledge base and reference materials
  • Reordering retrieved passages
  • Picking answers for a support chatbot
Sizes
about 280M
Hardware
from: Laptop
Commercial use allowedDetails
Text to speechNot maintained2023

StyleTTS 2

Columbia University · USA

A lightweight English speech synthesis model with natural intonation. Many other models, such as Kokoro, are built on it.

  • Voicing texts in English
  • A base for fine-tuning your own voice
  • Voice service prototypes
Sizes
about 150M
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to SQLNot maintained2023

CodeS

RUCKBReasoning, Renmin University of China · China

An early line of open text-to-SQL models starting at 1B, including variants fine-tuned for specific database schemas.

  • Turning an employee question into an SQL query
  • Drafting warehouse queries for a report
  • Embedding into a BI dashboard as a helper
Sizes
1B – 15B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQLGGUFNot maintained2023

NSQL

Numbers Station · USA

One of the first open text-to-SQL lines, including very small versions from 350M that run on an ordinary PC.

  • Turning a question into SQL from a table description
  • Hints while writing queries
  • A local analyst helper with no data leaving the company
Sizes
350M – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
RerankersRUNot maintained2022

Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO)

UKP Lab and the Sentence Transformers community · Germany

The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.

  • Reordering knowledge base search results
  • Selecting passages before a chatbot answers
  • Finding duplicates among tickets and product cards
Sizes
about 4M – 120M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRUNot maintained2019–2022

XLM-RoBERTa

Meta · USA

A classic multilingual encoder for 100 languages, including Russian. The base of many sentiment, NER and embedding models, including BGE-M3.

  • Detecting review sentiment in different languages
  • Extracting names and organizations after fine-tuning
  • Classifying requests
Sizes
270M – 10.7B
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRUNot maintained2021

DeBERTa-v3 и mDeBERTa-v3

Microsoft · USA

A time-tested encoder behind many classifiers and NER models (including GLiNER). The multilingual mDeBERTa-v3 understands Russian.

  • Classifying review sentiment
  • Entity extraction after fine-tuning
  • Checking whether a conclusion follows from a text
Sizes
70M – 435M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRUNot maintained2021

rubert-tiny

David Dale (cointegrated) · Russia

A very small Russian-English BERT that runs fast on a regular CPU. Ready-made fine-tuned versions exist for sentiment, toxicity and emotions.

  • Detecting review sentiment
  • Filtering rude chat messages
  • Fast classification of requests
Sizes
12M – 29M
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment