TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere · Canada
Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.
- Knowledge-base answers with source citations
- Agents that work with internal systems
- Translation and correspondence in different languages
- Sizes
- 2.5B – 218B-A25B
- Hardware
- from: Laptop
TextOllama2024–2026
NVIDIA · USA
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
- Agents with tool calling
- Reasoning and calculation tasks
- Answers based on long documents
- Sizes
- 4B – 550B-A55B
- Hardware
- from: Laptop
TextOllama2024–2026
IBM · USA
IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.
- Answers based on internal documents (RAG)
- Tool calling and agent work
- Data extraction and classification
- Sizes
- 350M – 34B
- Hardware
- from: Laptop
TextRUOllama2025–2026
Liquid AI · USA
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
- Offline assistant on a laptop or phone
- Data extraction from documents
- Tool calling in apps
- Sizes
- 230M – 24B-A2B
- Hardware
- from: Laptop
TextOllama2026
Meta Superintelligence Labs · USA
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
- Agents with tool calling
- Analysis of screenshots, charts and documents
- Multilingual assistant
- Sizes
- 30B
- Hardware
- from: 1 GPU
CodeOllama2026
DeepReinforce · not disclosed
Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.
- A developer agent in the terminal
- Fixing bugs from a task description
- Understanding and extending a large repository
- Sizes
- 9B – 397B
- Hardware
- from: 1 GPU
TextRUOllama2023–2026
Mistral AI · France
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
- Fast chat responses
- Data extraction from text
- Translation and multilingual work
- Sizes
- 3B – 675B
- Hardware
- from: Laptop
TextOllama2024–2026
LG AI Research · South Korea
Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.
- Corporate assistant
- Working with Korean and English texts
- Analysis of documents and images (4.5)
- Sizes
- 1.2B – 750B-A37B
- Hardware
- from: Laptop
TextOllama2023–2026
Upstage · South Korea
Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.
- Working with office documents
- Agents for routine tasks
- Help for developers
- Sizes
- 10.7B – 250B-A15B
- Hardware
- from: 1 GPU
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Image + textOllama2024–2026
Moondream (M87 Labs) · USA
A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.
- Finding and counting objects in photos
- Checking photos from field reports
- Captions and tags for a catalogue
- Sizes
- 2B – 9B-A2B
- Hardware
- from: Laptop
MedicineOllama2023–2026
EPFL · Switzerland
Open medical models from Swiss EPFL, fine-tuned on clinical guidelines on top of various base models. Does not replace a doctor; decisions are made by a specialist.
- Answering staff questions based on clinical guidelines
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 2B – 70B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Cognitive Computations (Eric Hartford) · USA
Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.
- Assistant that does not refuse legal but sensitive topics
- Internal tools under a strict system prompt
- Role-play and creative scenarios
- Sizes
- 0.5B – 405B
- Hardware
- from: Laptop
Image + textOllama2023–2026
LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China
The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.
- Answering questions about photos and screenshots
- Describing products from a photo
- Frame-by-frame video analysis
- Sizes
- 0.5B – 72B
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2026
IBM · USA
Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.
- Search across corporate documents
- RAG on a regular server without a GPU
- Reranking results
- Sizes
- 30M – 311M
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2026
IBM · USA
IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.
- Checking bot requests and replies
- Finding made-up facts in knowledge-base answers
- Checking your own rules written as text
- Sizes
- 38M – 8B
- Hardware
- from: Laptop
Image + textOllama2025–2026
IBM · USA
Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.
- Extracting fields from forms and invoices
- Turning charts and tables into data
- Answering questions about documents
- Sizes
- 2B – 4B
- Hardware
- from: Laptop
Text analysisOllama2024–2026
NuMind · France
Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.
- Extracting company details, amounts and dates from invoices and contracts into JSON
- Parsing receipts, waybills and forms against a set template
- Converting scans to Markdown for search
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
CodeOllama2025–2026
Essential AI · USA
An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.
- Writing and fixing code
- A developer agent on a single GPU
- Solving technical and scientific problems
- Sizes
- 8B
- Hardware
- from: Laptop
CodeOllama2024–2026
Alibaba (Qwen team) · China
The broadest open coding family: from 0.5B for autocompletion to 480B for agents. Qwen3-Coder-Next (80B, 3B active) works as a developer agent on a single GPU.
- Code autocompletion in the editor
- An agent that edits code in the repository on its own
- Writing and refining scripts, SQL and integrations
- Sizes
- 0.5B – 480B-A35B
- Hardware
- from: Laptop
MedicineOllama2025–2026
Google · USA
Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.
- Draft discharge summaries and reports for a doctor to review
- Hints for doctors when reviewing images
- Searching and summarising medical literature
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Search and RAGRUOllama2025–2026
Alibaba (Qwen) · China
Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.
- Knowledge base search for RAG
- Reranking results before answering
- Search across scans, slides and screenshots
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Technology Innovation Institute (TII) · UAE
A family from Abu Dhabi: from the early Falcon 40B and 180B to hybrid Falcon-H1 and tiny Falcon-H1-Tiny models of 90–600M parameters for devices.
- Assistant and answers based on documents
- Running on low-end hardware and devices
- Tool calling in simple agents
- Sizes
- 90M – 180B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Allen Institute for AI (Ai2) · USA
Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.
- Assistant and answers based on documents
- Reasoning tasks (Think versions)
- Fine-tuning on your data with a clear model history
- Sizes
- 1B – 32B
- Hardware
- from: Laptop
TranslationRUOllama2026
Google · USA
Translators based on Gemma 3 for 55 languages that can also translate text in images. Russian is supported. The 4B version fits on a laptop.
- Translating documents and correspondence
- Translating text from screenshots and photos
- Localizing websites and apps
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Documents and OCROllama2025–2026
DeepSeek · China
An OCR model that compresses a page into a small number of visual tokens, so it processes large volumes quickly. Version 2 better understands reading order.
- Bulk recognition of scanned invoices and contracts
- Table recognition
- Converting PDFs to Markdown for search and RAG
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUOllama2026
Zhipu AI (Z.ai) · China
A lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables and formulas
- Extracting fields to JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
CodeRUOllama2025
Mistral AI (with All Hands AI) · France
Mistral models for agentic development: they read the repository, edit files and run commands on their own. The 24B version fits on a single GPU.
- A developer agent that fixes tickets from the tracker
- Extending internal systems from a description
- Automating routine code edits
- Sizes
- 24B – 123B
- Hardware
- from: 1 GPU
TextOllama2023–2025
Nous Research · USA
Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.
- Agents that call functions and APIs
- Data extraction in strict JSON format
- Assistant with flexible role and tone settings
- Sizes
- 3B – 405B
- Hardware
- from: Laptop
TextOllama2025
Deep Cogito · USA
Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.
- A chat assistant with a reasoning mode
- Writing code and calling tools
- Answering complex questions about documents
- Sizes
- 3B – 671B
- Hardware
- from: Laptop
Moderation and safetyOllama2025
OpenAI · USA
Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.
- Moderation by internal company rules
- Labeling disputed messages with an explanation
- Checking reviews and listings before publishing
- Sizes
- 20B – 120B
- Hardware
- from: 1 GPU
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Mixedbread · Germany
Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.
- Search across a knowledge base
- Reranking results before a bot answers
- Product catalog search
- Sizes
- 17M – 1.5B
- Hardware
- from: Laptop
Search and RAGOllama2025
Google · USA
A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.
- On-device document search
- RAG without sending data outside
- Text classification
- Sizes
- 300M
- Hardware
- from: Laptop
Search and RAGOllama2023–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.
- Search across English-language documents
- Picking passages for chatbot answers (RAG)
- Code search (bge-code)
- Sizes
- 24M – 9B
- Hardware
- from: Laptop
TextOllama2025
OpenAI · USA
OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.
- AI agent that calls internal systems
- Answers based on internal policies
- Drafts of emails and reports
- Sizes
- 20B, 120B
- Hardware
- from: 1 GPU
TextRUOllama2024–2025
Hugging Face · USA
Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.
- Simple on-device assistant
- Classification and routing of requests
- Base for fine-tuning on a narrow task
- Sizes
- 135M – 3B
- Hardware
- from: Laptop
Math and reasoningOllama2025
Agentica (Berkeley, Sky Computing Lab) and Together AI · USA
Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.
- Solving maths problems with step-by-step working
- Generating and checking code
- An agent for fixing bugs in a repository
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
TextOllama2025
DeepSeek · China
A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.
- Complex calculations and logic checks
- Analysis of contracts and internal policies
- Help for developers
- Sizes
- 1,5B – 671B
- Hardware
- from: Laptop
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
Moderation and safetyOllama2023–2025
Meta · USA
Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.
- Checking user questions to the bot
- Checking bot replies before sending
- Reporting which rule category was violated
- Sizes
- 1B – 12B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of contracts and internal policies
- Sizes
- 32B
- Hardware
- from: 1 GPU
Search and RAGRUOllama2024–2025
Nomic AI · USA
Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.
- Search across documents and a knowledge base
- Code search
- Search across scans and PDFs without text recognition
- Sizes
- 137M – 7B
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2025
Google · USA
Gemma-based filters: they check text for dangerous and offensive content, and ShieldGemma 2 checks images. Focused on English.
- Moderating user messages
- Checking bot replies
- Checking generated images before publishing
- Sizes
- 2B – 27B
- Hardware
- from: Laptop
TextOllama2023–2025
Ai2 · USA
Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.
- Employee assistant on your own server
- Math and precise instruction following
- Reference recipe for your own fine-tuning
- Sizes
- 7B – 405B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.
- Calculations and formula checks
- Checking calculations in estimates and reports
- Working through problems step by step
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Search and RAGRUOllama2019–2025
UKP Lab (TU Darmstadt), later Hugging Face · Germany
The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.
- Search across a knowledge base and FAQ
- Finding similar tickets and duplicates
- Grouping reviews and requests by topic
- Sizes
- about 20M – 470M
- Hardware
- from: Laptop
Text analysisRUOllama2024–2025
Jina AI · Germany
Small models that turn raw web page HTML into clean Markdown or JSON. Handy for preparing websites for a knowledge base. Non-commercial license only.
- Cleaning website pages for a knowledge base
- Extracting data from pages into JSON
- Preparing texts for RAG
- Sizes
- 0.5B – 1.5B
- Hardware
- from: Laptop
Search and RAGRUOllama2024
Snowflake · USA
Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.
- Search across documents and knowledge bases
- Picking passages for RAG
- Search across reports and internal data
- Sizes
- 22M – 568M
- Hardware
- from: Laptop
CodeOllama2024
INF Technology · China
Fully reproducible coding models: along with the weights, the data, its cleaning pipeline and the training recipe are open. Understand English and Chinese.
- Code generation and completion
- Training your own coding model from an open recipe
- A programming assistant on low-end hardware
- Sizes
- 1.5B – 8B
- Hardware
- from: Laptop
TextOllama2024
Nexusflow · USA
Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.
- Research on agents and function calling
- Comparison with commercial models
- Experiments with a chat assistant
- Sizes
- 70B – 72B
- Hardware
- from: 1 GPU
CodeOllama2023–2024
DeepSeek · China
DeepSeek's coding model family: from small autocompletion models to the large MoE V2, which matched closed models in 2024. Later, coding moved into DeepSeek's general models.
- Code autocompletion and generation
- Translating code between programming languages
- Finding bugs and explaining other people's code
- Sizes
- 1.3B – 236B-A21B
- Hardware
- from: Laptop
CodeOllama2024
01.AI · China
Coding models from 01.AI at 1.5B and 9B with a 128K-token context and support for 52 programming languages. A separate line next to the text Yi models.
- Code autocompletion and generation
- Explaining and refactoring code
- A programming assistant without the cloud
- Sizes
- 1.5B – 9B
- Hardware
- from: Laptop
Fact-checking and judgesOllamaNot maintained2024
UT Austin and Bespoke Labs · USA
Checks whether each claim in an AI answer is supported by the source documents. The small versions are free; the larger 7B is in Ollama but non-commercial.
- Checking RAG bot answers against documents
- Finding unsupported claims in reports and summaries
- Automated quality control of AI answers
- Sizes
- 0.4B – 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Stability AI · UK
Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.
- A lightweight chatbot on an ordinary PC
- Code autocompletion in the editor
- A base for fine-tuning on your own task
- Sizes
- 1.6B – 12B
- Hardware
- from: Laptop
CodeOllamaNot maintained2024
Mistral AI · France
Mistral's coding model covering 80+ programming languages. The open weights of the main version cannot be used in production without a paid license; newer Codestral versions are API-only.
- Evaluation and testing before buying a license
- Code autocompletion (with a commercial license)
- Research on coding model quality
- Sizes
- 7B – 22B
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
Zhipu AI (Z.ai) and Tsinghua University · China
Coding models from the creators of GLM. CodeGeeX4-ALL-9B, based on GLM-4-9B, combines autocompletion, code chat, function calling and repository search in one model.
- Code autocompletion in the IDE
- A code chat assistant
- Answering questions about a repository
- Sizes
- 6B – 9B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
OpenChat (Tsinghua University) · China
Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.
- Chat assistant on an inexpensive server
- Drafts of emails and replies
- Help with simple code
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
01.AI · China
Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.
- Chat assistant on a single GPU
- Analysis of long documents
- Classification and data extraction from text
- Sizes
- 6B – 34B
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2023–2024
Defog · USA
One of the first open models that turn a plain-language question into an SQL query against a database. Available in Ollama, but newer competitors are already stronger.
- Answering managers' questions from the sales database without an analyst
- Drafting SQL queries for reports
- An assistant inside a BI system
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
BigCode (Hugging Face and ServiceNow) · USA / France
One of the first open coding models, trained on an open set of source code with an option to exclude your own repository. Today it is more a base for fine-tuning than a leader.
- Code autocompletion in the editor
- Fine-tuning on the company's internal code
- Generating boilerplate code and tests
- Sizes
- 1B – 15B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Hugging Face (H4) · USA
Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.
- Lightweight chat assistant
- Reference and starting point for your own fine-tuning
- Drafts of texts and replies
- Sizes
- 7B – 141B-A35B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
Search and RAGRUOllamaNot maintained2024
BAAI · China
A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.
- Search across a document base
- RAG for a chatbot
- Finding similar requests and duplicates
- Sizes
- 568M
- Hardware
- from: Laptop
CodeOllamaNot maintained2023–2024
Meta · USA
A version of Llama 2 further trained on code, with variants for Python and for chat. Outdated, but many ready-made fine-tuned versions and tools exist.
- Code autocompletion and explanation
- Generating Python scripts
- Base model for fine-tuning on your own stack
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
WizardLM (Microsoft and Peking University) · USA / China
Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.
- Complex multi-step instructions
- Help for developers
- Solving math problems
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Text to SQLOllamaNot maintained2024
MotherDuck and Numbers Station · USA
A model for turning questions into SQL, built for the embedded analytics database DuckDB. Available in Ollama, convenient for working with CSV and Parquet locally.
- Plain-language questions about CSV and Parquet exports
- DuckDB queries inside analytics scripts
- Quick analytics on a laptop without a server
- Sizes
- 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Intel · USA
A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.
- A simple chat assistant
- Experiments with running on Intel hardware
- A base for fine-tuning
- Sizes
- 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Microsoft Research · USA
Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.
- Research on reasoning methods
- Comparison with modern small models
- Training specialists
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
LMSYS (Berkeley and partners) · USA
One of the first open chat models (2023): LLaMA fine-tuned on user conversations with ChatGPT. A historical milestone; today it is weaker than any modern model of the same size.
- Experiments and team training
- Simple chat assistant for tests
- Comparison with newer models
- Sizes
- 7B – 33B
- Hardware
- from: Laptop