Models available in Ollama

This collection lists models available in the Ollama library that run with a single command on a computer or server. It is the simplest way to try local AI and connect it to your services through an API. Pick the version size to match your memory, and check the model's own license, since Ollama does not change it.

78 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
TextOllama2023–2026

DeepSeek

DeepSeek · China

DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.

  • Employee assistant on your own server
  • Analysis of long contracts and reports
  • Agents that work with tools and APIs
Sizes
7B – 1.6T-A49B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2026

InternLM / Intern-S

Shanghai AI Laboratory · China

Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.

  • Research assistant: papers, formulas, data
  • Analysis of scientific and technical documents
  • Corporate chat on small models
Sizes
1.8B – about 1T
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2024–2026

Aya

Cohere Labs · Canada

Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.

  • Translation and correspondence in less common languages
  • Multilingual chat assistant
  • Analysis of images with text (Vision)
Sizes
3.3B – 35B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2023–2026

Qwen

Alibaba · China

A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.

  • Chatbot and knowledge-base assistant
  • Replies to emails and customer requests
  • Document parsing and classification
Sizes
0,6B – 2,4T-A95B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2026

GLM (ChatGLM)

Zhipu AI (Z.ai) · China

One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.

  • Corporate chat assistant
  • Agents for routine office tasks
  • Help for developers
Sizes
1.5B – 744B-A40B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2024–2026

Cohere Command

Cohere · Canada

Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.

  • Knowledge-base answers with source citations
  • Agents that work with internal systems
  • Translation and correspondence in different languages
Sizes
2.5B – 218B-A25B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024–2026

NVIDIA Nemotron

NVIDIA · USA

NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).

  • Agents with tool calling
  • Reasoning and calculation tasks
  • Answers based on long documents
Sizes
4B – 550B-A55B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024–2026

IBM Granite

IBM · USA

IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.

  • Answers based on internal documents (RAG)
  • Tool calling and agent work
  • Data extraction and classification
Sizes
350M – 34B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2025–2026

Liquid LFM

Liquid AI · USA

Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.

  • Offline assistant on a laptop or phone
  • Data extraction from documents
  • Tool calling in apps
Sizes
230M – 24B-A2B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2026

Muse Glimmer

Meta Superintelligence Labs · USA

An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.

  • Agents with tool calling
  • Analysis of screenshots, charts and documents
  • Multilingual assistant
Sizes
30B
Hardware
from: 1 GPU
Commercial use allowedDetails
CodeOllama2026

Ornith

DeepReinforce · not disclosed

Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.

  • A developer agent in the terminal
  • Fixing bugs from a task description
  • Understanding and extending a large repository
Sizes
9B – 397B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextRUOllama2023–2026

Mistral

Mistral AI · France

European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).

  • Fast chat responses
  • Data extraction from text
  • Translation and multilingual work
Sizes
3B – 675B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024–2026

EXAONE

LG AI Research · South Korea

Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.

  • Corporate assistant
  • Working with Korean and English texts
  • Analysis of documents and images (4.5)
Sizes
1.2B – 750B-A37B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2026

Solar

Upstage · South Korea

Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.

  • Working with office documents
  • Agents for routine tasks
  • Help for developers
Sizes
10.7B – 250B-A15B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextOllama2024–2026

Gemma

Google · USA

Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.

  • Offline assistant on a laptop
  • Reading photos of documents and receipts
  • Customer request classification
Sizes
270M – 31B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningOllama2025–2026

OpenThinker

Open Thoughts (Stanford, Berkeley and other universities) · USA

Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Checking the logic of internal policies
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2024–2026

Moondream

Moondream (M87 Labs) · USA

A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.

  • Finding and counting objects in photos
  • Checking photos from field reports
  • Captions and tags for a catalogue
Sizes
2B – 9B-A2B
Hardware
from: Laptop
Commercial use with conditionsDetails
MedicineOllama2023–2026

Meditron

EPFL · Switzerland

Open medical models from Swiss EPFL, fine-tuned on clinical guidelines on top of various base models. Does not replace a doctor; decisions are made by a specialist.

  • Answering staff questions based on clinical guidelines
  • Draft discharge summaries for a doctor to review
  • Searching medical literature
Sizes
2B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2023–2026

Dolphin

Cognitive Computations (Eric Hartford) · USA

Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.

  • Assistant that does not refuse legal but sensitive topics
  • Internal tools under a strict system prompt
  • Role-play and creative scenarios
Sizes
0.5B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textOllama2023–2026

LLaVA

LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China

The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.

  • Answering questions about photos and screenshots
  • Describing products from a photo
  • Frame-by-frame video analysis
Sizes
0.5B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2024–2026

MiniCPM-V

OpenBMB (ModelBest and Tsinghua University) · China

Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.

  • On-device text recognition in photos
  • Processing receipts and documents without sending them to the cloud
  • Describing photos and video
Sizes
1.3B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2024–2026

Granite Embedding

IBM · USA

Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.

  • Search across corporate documents
  • RAG on a regular server without a GPU
  • Reranking results
Sizes
30M – 311M
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safetyOllama2024–2026

Granite Guardian

IBM · USA

IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.

  • Checking bot requests and replies
  • Finding made-up facts in knowledge-base answers
  • Checking your own rules written as text
Sizes
38M – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2025–2026

Granite Vision

IBM · USA

Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.

  • Extracting fields from forms and invoices
  • Turning charts and tables into data
  • Answering questions about documents
Sizes
2B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisOllama2024–2026

NuExtract

NuMind · France

Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.

  • Extracting company details, amounts and dates from invoices and contracts into JSON
  • Parsing receipts, waybills and forms against a set template
  • Converting scans to Markdown for search
Sizes
0.5B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2025–2026

Rnj-1

Essential AI · USA

An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.

  • Writing and fixing code
  • A developer agent on a single GPU
  • Solving technical and scientific problems
Sizes
8B
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2024–2026

Qwen Coder

Alibaba (Qwen team) · China

The broadest open coding family: from 0.5B for autocompletion to 480B for agents. Qwen3-Coder-Next (80B, 3B active) works as a developer agent on a single GPU.

  • Code autocompletion in the editor
  • An agent that edits code in the repository on its own
  • Writing and refining scripts, SQL and integrations
Sizes
0.5B – 480B-A35B
Hardware
from: Laptop
Commercial use allowedDetails
MedicineOllama2025–2026

MedGemma

Google · USA

Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.

  • Draft discharge summaries and reports for a doctor to review
  • Hints for doctors when reviewing images
  • Searching and summarising medical literature
Sizes
4B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRUOllama2025–2026

Qwen3 Embedding / Reranker

Alibaba (Qwen) · China

Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.

  • Knowledge base search for RAG
  • Reranking results before answering
  • Search across scans, slides and screenshots
Sizes
0.6B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2023–2026

Falcon

Technology Innovation Institute (TII) · UAE

A family from Abu Dhabi: from the early Falcon 40B and 180B to hybrid Falcon-H1 and tiny Falcon-H1-Tiny models of 90–600M parameters for devices.

  • Assistant and answers based on documents
  • Running on low-end hardware and devices
  • Tool calling in simple agents
Sizes
90M – 180B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2023–2026

Phi

Microsoft · USA

Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.

  • Assistant on a laptop or your own server
  • Reasoning and calculation tasks
  • Analysis of images and diagrams (vision versions)
Sizes
1.3B – 42B-A6.6B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2024–2026

OLMo

Allen Institute for AI (Ai2) · USA

Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.

  • Assistant and answers based on documents
  • Reasoning tasks (Think versions)
  • Fine-tuning on your data with a clear model history
Sizes
1B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
TranslationRUOllama2026

TranslateGemma

Google · USA

Translators based on Gemma 3 for 55 languages that can also translate text in images. Russian is supported. The 4B version fits on a laptop.

  • Translating documents and correspondence
  • Translating text from screenshots and photos
  • Localizing websites and apps
Sizes
4B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCROllama2025–2026

DeepSeek-OCR

DeepSeek · China

An OCR model that compresses a page into a small number of visual tokens, so it processes large volumes quickly. Version 2 better understands reading order.

  • Bulk recognition of scanned invoices and contracts
  • Table recognition
  • Converting PDFs to Markdown for search and RAG
Sizes
about 3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRUOllama2026

GLM-OCR

Zhipu AI (Z.ai) · China

A lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.

  • Recognising invoices, contracts and delivery notes, including in Russian
  • Recognising tables and formulas
  • Extracting fields to JSON
Sizes
0.9B
Hardware
from: Laptop
Commercial use allowedDetails
CodeRUOllama2025

Devstral

Mistral AI (with All Hands AI) · France

Mistral models for agentic development: they read the repository, edit files and run commands on their own. The 24B version fits on a single GPU.

  • A developer agent that fixes tickets from the tracker
  • Extending internal systems from a description
  • Automating routine code edits
Sizes
24B – 123B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextOllama2023–2025

Hermes

Nous Research · USA

Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.

  • Agents that call functions and APIs
  • Data extraction in strict JSON format
  • Assistant with flexible role and tone settings
Sizes
3B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2025

Cogito

Deep Cogito · USA

Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.

  • A chat assistant with a reasoning mode
  • Writing code and calling tools
  • Answering complex questions about documents
Sizes
3B – 671B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safetyOllama2025

gpt-oss-safeguard

OpenAI · USA

Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.

  • Moderation by internal company rules
  • Labeling disputed messages with an explanation
  • Checking reviews and listings before publishing
Sizes
20B – 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textOllama2023–2025

Qwen-VL

Alibaba (Qwen team) · China

One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.

  • Extracting data from scanned invoices and delivery notes
  • Analysing photos of products and shelves
  • Analysing video and camera footage
Sizes
2B – 235B-A22B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2024–2025

mxbai (Mixedbread Embed и Rerank)

Mixedbread · Germany

Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.

  • Search across a knowledge base
  • Reranking results before a bot answers
  • Product catalog search
Sizes
17M – 1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGOllama2025

EmbeddingGemma

Google · USA

A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.

  • On-device document search
  • RAG without sending data outside
  • Text classification
Sizes
300M
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGOllama2023–2025

BGE (BAAI General Embedding)

BAAI (Beijing Academy of Artificial Intelligence) · China

Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.

  • Search across English-language documents
  • Picking passages for chatbot answers (RAG)
  • Code search (bge-code)
Sizes
24M – 9B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2025

gpt-oss

OpenAI · USA

OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.

  • AI agent that calls internal systems
  • Answers based on internal policies
  • Drafts of emails and reports
Sizes
20B, 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextRUOllama2024–2025

SmolLM

Hugging Face · USA

Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.

  • Simple on-device assistant
  • Classification and routing of requests
  • Base for fine-tuning on a narrow task
Sizes
135M – 3B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningOllama2025

DeepScaleR, DeepCoder, DeepSWE

Agentica (Berkeley, Sky Computing Lab) and Together AI · USA

Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.

  • Solving maths problems with step-by-step working
  • Generating and checking code
  • An agent for fixing bugs in a repository
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2025

DeepSeek-R1

DeepSeek · China

A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.

  • Complex calculations and logic checks
  • Analysis of contracts and internal policies
  • Help for developers
Sizes
1,5B – 671B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2025

Llama

Meta · USA

The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.

  • Assistant for employees
  • Summaries of meetings and documents
  • Base for industry-specific fine-tuning
Sizes
1B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safetyOllama2023–2025

Llama Guard

Meta · USA

Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.

  • Checking user questions to the bot
  • Checking bot replies before sending
  • Reporting which rule category was violated
Sizes
1B – 12B
Hardware
from: Laptop
Commercial use with conditionsDetails
Math and reasoningOllama2024–2025

QwQ

Qwen (Alibaba) · China

Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Checking the logic of contracts and internal policies
Sizes
32B
Hardware
from: 1 GPU
Commercial use allowedDetails
Search and RAGRUOllama2024–2025

Nomic Embed

Nomic AI · USA

Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.

  • Search across documents and a knowledge base
  • Code search
  • Search across scans and PDFs without text recognition
Sizes
137M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safetyOllama2024–2025

ShieldGemma

Google · USA

Gemma-based filters: they check text for dangerous and offensive content, and ShieldGemma 2 checks images. Focused on English.

  • Moderating user messages
  • Checking bot replies
  • Checking generated images before publishing
Sizes
2B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2025

Tulu

Ai2 · USA

Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.

  • Employee assistant on your own server
  • Math and precise instruction following
  • Reference recipe for your own fine-tuning
Sizes
7B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Math and reasoningOllama2024–2025

Qwen2.5-Math

Qwen (Alibaba) · China

Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.

  • Calculations and formula checks
  • Checking calculations in estimates and reports
  • Working through problems step by step
Sizes
1.5B – 72B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRUOllama2019–2025

Sentence Transformers (all-MiniLM, paraphrase-multilingual)

UKP Lab (TU Darmstadt), later Hugging Face · Germany

The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.

  • Search across a knowledge base and FAQ
  • Finding similar tickets and duplicates
  • Grouping reviews and requests by topic
Sizes
about 20M – 470M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRUOllama2024–2025

ReaderLM (Jina)

Jina AI · Germany

Small models that turn raw web page HTML into clean Markdown or JSON. Handy for preparing websites for a knowledge base. Non-commercial license only.

  • Cleaning website pages for a knowledge base
  • Extracting data from pages into JSON
  • Preparing texts for RAG
Sizes
0.5B – 1.5B
Hardware
from: Laptop
Non-commercial onlyDetails
Search and RAGRUOllama2024

Snowflake Arctic Embed

Snowflake · USA

Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.

  • Search across documents and knowledge bases
  • Picking passages for RAG
  • Search across reports and internal data
Sizes
22M – 568M
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2024

OpenCoder

INF Technology · China

Fully reproducible coding models: along with the weights, the data, its cleaning pipeline and the training recipe are open. Understand English and Chinese.

  • Code generation and completion
  • Training your own coding model from an open recipe
  • A programming assistant on low-end hardware
Sizes
1.5B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024

Athene

Nexusflow · USA

Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.

  • Research on agents and function calling
  • Comparison with commercial models
  • Experiments with a chat assistant
Sizes
70B – 72B
Hardware
from: 1 GPU
Non-commercial onlyDetails
CodeOllama2023–2024

DeepSeek-Coder

DeepSeek · China

DeepSeek's coding model family: from small autocompletion models to the large MoE V2, which matched closed models in 2024. Later, coding moved into DeepSeek's general models.

  • Code autocompletion and generation
  • Translating code between programming languages
  • Finding bugs and explaining other people's code
Sizes
1.3B – 236B-A21B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllama2024

Yi-Coder

01.AI · China

Coding models from 01.AI at 1.5B and 9B with a 128K-token context and support for 52 programming languages. A separate line next to the text Yi models.

  • Code autocompletion and generation
  • Explaining and refactoring code
  • A programming assistant without the cloud
Sizes
1.5B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judgesOllamaNot maintained2024

MiniCheck

UT Austin and Bespoke Labs · USA

Checks whether each claim in an AI answer is supported by the source documents. The small versions are free; the larger 7B is in Ollama but non-commercial.

  • Checking RAG bot answers against documents
  • Finding unsupported claims in reports and summaries
  • Automated quality control of AI answers
Sizes
0.4B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

StableLM и Stable Code

Stability AI · UK

Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.

  • A lightweight chatbot on an ordinary PC
  • Code autocompletion in the editor
  • A base for fine-tuning on your own task
Sizes
1.6B – 12B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllamaNot maintained2024

Codestral

Mistral AI · France

Mistral's coding model covering 80+ programming languages. The open weights of the main version cannot be used in production without a paid license; newer Codestral versions are API-only.

  • Evaluation and testing before buying a license
  • Code autocompletion (with a commercial license)
  • Research on coding model quality
Sizes
7B – 22B
Hardware
from: Laptop
Non-commercial onlyDetails
CodeOllamaNot maintained2023–2024

CodeGeeX

Zhipu AI (Z.ai) and Tsinghua University · China

Coding models from the creators of GLM. CodeGeeX4-ALL-9B, based on GLM-4-9B, combines autocompletion, code chat, function calling and repository search in one model.

  • Code autocompletion in the IDE
  • A code chat assistant
  • Answering questions about a repository
Sizes
6B – 9B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

OpenChat

OpenChat (Tsinghua University) · China

Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.

  • Chat assistant on an inexpensive server
  • Drafts of emails and replies
  • Help with simple code
Sizes
7B – 13B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

Yi

01.AI · China

Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.

  • Chat assistant on a single GPU
  • Analysis of long documents
  • Classification and data extraction from text
Sizes
6B – 34B
Hardware
from: Laptop
Commercial use allowedDetails
Text to SQLOllamaNot maintained2023–2024

SQLCoder

Defog · USA

One of the first open models that turn a plain-language question into an SQL query against a database. Available in Ollama, but newer competitors are already stronger.

  • Answering managers' questions from the sales database without an analyst
  • Drafting SQL queries for reports
  • An assistant inside a BI system
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
CodeOllamaNot maintained2023–2024

StarCoder

BigCode (Hugging Face and ServiceNow) · USA / France

One of the first open coding models, trained on an open set of source code with an option to exclude your own repository. Today it is more a base for fine-tuning than a leader.

  • Code autocompletion in the editor
  • Fine-tuning on the company's internal code
  • Generating boilerplate code and tests
Sizes
1B – 15B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

Zephyr

Hugging Face (H4) · USA

Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.

  • Lightweight chat assistant
  • Reference and starting point for your own fine-tuning
  • Drafts of texts and replies
Sizes
7B – 141B-A35B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllamaNot maintained2023–2024

TinyLlama

TinyLlama (SUTD researchers) · Singapore

A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.

  • Simple chatbots on low-end hardware
  • Experiments and team training
  • A base for fine-tuning on a narrow task
Sizes
1.1B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllamaNot maintained2024

BGE-M3

BAAI · China

A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.

  • Search across a document base
  • RAG for a chatbot
  • Finding similar requests and duplicates
Sizes
568M
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllamaNot maintained2023–2024

Code Llama

Meta · USA

A version of Llama 2 further trained on code, with variants for Python and for chat. Outdated, but many ready-made fine-tuned versions and tools exist.

  • Code autocompletion and explanation
  • Generating Python scripts
  • Base model for fine-tuning on your own stack
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

WizardLM, WizardCoder, WizardMath

WizardLM (Microsoft and Peking University) · USA / China

Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.

  • Complex multi-step instructions
  • Help for developers
  • Solving math problems
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to SQLOllamaNot maintained2024

DuckDB-NSQL

MotherDuck and Numbers Station · USA

A model for turning questions into SQL, built for the embedded analytics database DuckDB. Available in Ollama, convenient for working with CSV and Parquet locally.

  • Plain-language questions about CSV and Parquet exports
  • DuckDB queries inside analytics scripts
  • Quick analytics on a laptop without a server
Sizes
7B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023

NeuralChat

Intel · USA

A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.

  • A simple chat assistant
  • Experiments with running on Intel hardware
  • A base for fine-tuning
Sizes
7B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllamaNot maintained2023

Orca 2

Microsoft Research · USA

Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.

  • Research on reasoning methods
  • Comparison with modern small models
  • Training specialists
Sizes
7B – 13B
Hardware
from: Laptop
Non-commercial onlyDetails
TextOllamaNot maintained2023

Vicuna

LMSYS (Berkeley and partners) · USA

One of the first open chat models (2023): LLaMA fine-tuned on user conversations with ChatGPT. A historical milestone; today it is weaker than any modern model of the same size.

  • Experiments and team training
  • Simple chat assistant for tests
  • Comparison with newer models
Sizes
7B – 33B
Hardware
from: Laptop
Commercial use with conditionsDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment