MedicineGGUF2023–2026
FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China
A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Hints for doctors when reviewing images (Vision)
- Sizes
- 7B – 72B
- Hardware
- from: Laptop
TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextGGUF2025–2026
Xiaomi · China
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
- Logic and calculation tasks
- Agents with tools
- Help for developers
- Sizes
- 7B – 1,02T-A42B
- Hardware
- from: Laptop
TextRU2024–2026
Sber · Russia
Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.
- Russian-language employee assistant on your own server
- Customer replies and request handling in Russian
- Working with contracts and internal policies
- Sizes
- 10B-A1.8B – 702B-A36B
- Hardware
- from: Laptop
TextRU2022–2026
Yandex · Russia
Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.
- Russian-language assistant and chatbot
- Answers based on the company knowledge base
- Base for industry-specific fine-tuning
- Sizes
- 8B – 100B
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
Text2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact text models that run directly on a device: laptop, phone or mini PC. The 1B and 2B MiniCPM5 models focus on tool calling and long context.
- A local chat assistant without the cloud
- Data extraction and text classification
- Tool calling and simple agents on low-end hardware
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Text2024–2026
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
- Reasoning, maths and technical questions
- Analysing long documents
- Agents and writing code
- Sizes
- 0.9B – 375B-A23B
- Hardware
- from: Laptop
TextGGUF2025–2026
Renmin University of China (GSAI) and Ant Group (inclusionAI) · China
Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.
- Fast generation of code and text
- Agent scenarios with long context
- Research into alternatives to standard LLMs
- Sizes
- 8B – 100B (MoE)
- Hardware
- from: 1 GPU
TextRU2026
SberDevices (ai-forever) · Russia
A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.
- Experiments with faster generation
- Fine-tuning small models for your own tasks
- Research
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
Text2024–2026
Tencent · China
Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.
- Corporate assistant
- Translation and multilingual texts
- Agents with tools
- Sizes
- 0.5B – 770B-A49B
- Hardware
- from: Laptop
Text2025–2026
Ant Group (inclusionAI) · China
An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.
- Corporate assistant
- Agents for office processes
- Financial analytics (Fin version available)
- Sizes
- 7.9B-A1.3B – 1T
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere · Canada
Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.
- Knowledge-base answers with source citations
- Agents that work with internal systems
- Translation and correspondence in different languages
- Sizes
- 2.5B – 218B-A25B
- Hardware
- from: Laptop
TextOllama2024–2026
NVIDIA · USA
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
- Agents with tool calling
- Reasoning and calculation tasks
- Answers based on long documents
- Sizes
- 4B – 550B-A55B
- Hardware
- from: Laptop
TextOllama2024–2026
IBM · USA
IBM enterprise models with transparent training data and ISO 42001 certification. Granite 4 is a memory-efficient Mamba and Transformer hybrid.
- Answers based on internal documents (RAG)
- Tool calling and agent work
- Data extraction and classification
- Sizes
- 350M – 34B
- Hardware
- from: Laptop
TextRUOllama2025–2026
Liquid AI · USA
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
- Offline assistant on a laptop or phone
- Data extraction from documents
- Tool calling in apps
- Sizes
- 230M – 24B-A2B
- Hardware
- from: Laptop
TextOllama2026
Meta Superintelligence Labs · USA
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
- Agents with tool calling
- Analysis of screenshots, charts and documents
- Multilingual assistant
- Sizes
- 30B
- Hardware
- from: 1 GPU
TextRUOllama2023–2026
Mistral AI · France
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
- Fast chat responses
- Data extraction from text
- Translation and multilingual work
- Sizes
- 3B – 675B
- Hardware
- from: Laptop
TextGGUF2025–2026
Moonshot AI · China
Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.
- Multi-step agents: search, data collection, reports
- In-depth document analysis
- Help for developers
- Sizes
- 16B-A3B – 2.8T-A104B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Meituan · China
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
- Agents for orders and service processes
- Corporate assistant
- Analysis of long documents
- Sizes
- 1B – 1.6T-A48B
- Hardware
- from: Laptop
TextOllama2024–2026
LG AI Research · South Korea
Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.
- Corporate assistant
- Working with Korean and English texts
- Analysis of documents and images (4.5)
- Sizes
- 1.2B – 750B-A37B
- Hardware
- from: Laptop
TextOllama2023–2026
Upstage · South Korea
Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.
- Working with office documents
- Agents for routine tasks
- Help for developers
- Sizes
- 10.7B – 250B-A15B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Kakao · South Korea
Compact Korean-English models from Kakao. Kanana 2 30B-A3B is fast thanks to MoE; small 1–3B versions suit a regular PC.
- Support chatbot
- Customer request classification
- Lightweight assistant on your own PC
- Sizes
- 1.3B – 30B-A3B
- Hardware
- from: Laptop
TextRU2024–2026
T-Bank · Russia
T-Bank models fine-tuned from Qwen for Russian: they write and reason in Russian noticeably better than the original. T-Lite is 8B, T-Pro 32B on one GPU; T-Search is a multi-step search agent in Russian and English.
- Russian-language support chatbot
- Analysis of requests and documents in Russian
- Answers based on the company knowledge base
- Sizes
- 7B – 36B-A3B
- Hardware
- from: Laptop
TextGGUF2025–2026
Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland
Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.
- Multilingual assistant
- Answers based on documents
- Analysis of images and scans (v1.5)
- Sizes
- 0.5B – 70B
- Hardware
- from: Laptop
TextGGUF2026
Thinking Machines Lab · USA
Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.
- Flagship-level corporate assistant
- Analysis of documents, images and audio
- Programming help
- Sizes
- 276B-A12B, 975B-A41B
- Hardware
- from: Cluster
CodeGGUF2026
Poolside · USA
Models for agentic programming: they edit code in a repository on their own. The small XS runs on a Mac with 36 GB of memory; S 2.1 has a 1M-token context.
- Coding agent for in-house development
- Bug fixing and code improvements
- Working with large codebases
- Sizes
- 33B-A3B – 225B-A23B
- Hardware
- from: 1 GPU
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
MiniMax · China
Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.
- Analysis of large document archives in a single request
- Agents with tools
- Help for developers
- Sizes
- 230B-A10B – 456B-A46B
- Hardware
- from: Cluster
MedicineOllama2023–2026
EPFL · Switzerland
Open medical models from Swiss EPFL, fine-tuned on clinical guidelines on top of various base models. Does not replace a doctor; decisions are made by a specialist.
- Answering staff questions based on clinical guidelines
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 2B – 70B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Cognitive Computations (Eric Hartford) · USA
Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.
- Assistant that does not refuse legal but sensitive topics
- Internal tools under a strict system prompt
- Role-play and creative scenarios
- Sizes
- 0.5B – 405B
- Hardware
- from: Laptop
TextGGUF2025–2026
StepFun · China
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
- High-load agents
- Analysis of documents with diagrams and screenshots
- Help for developers
- Sizes
- 8B – 321B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
Arcee AI · USA
An American family of MoE models trained from scratch: Nano, Mini and Large. Trinity-Large-Thinking (398B) reasons before answering.
- Agents with tool calling
- Reasoning tasks
- Corporate assistant on your own servers
- Sizes
- 6B – 398B-A13B
- Hardware
- from: Laptop
TextGGUF2025–2026
ServiceNow · USA
ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.
- A reasoning assistant for internal services
- Tool calling and enterprise agents
- Analysing screenshots and documents with images
- Sizes
- 5B – 15B
- Hardware
- from: Laptop
Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
TextRU2024–2026
Ivan Bondarenko (bond005), Novosibirsk State University · Russia
Russian-language models for working with documents rather than chatting: knowledge-base answers, extraction of entities and facts from Russian text, long context.
- Answers to questions based on internal documents
- Extracting names, dates and amounts from contracts
- Short summaries of long Russian texts
- Sizes
- 1.5B – 7.6B
- Hardware
- from: Laptop
TextGGUF2024–2026
Sarvam AI · India
Indian models focused on 22 languages of India. Sarvam 30B and 105B (2026) are MoE models with strong reasoning and agent skills.
- Multilingual customer support
- Reasoning and calculation tasks
- Agents with tool calling
- Sizes
- 2B – 105B-A10B
- Hardware
- from: Laptop
Text2025–2026
Reka AI · USA
Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.
- Photo and video analysis (Edge)
- Object detection in images
- Reasoning tasks (Flash)
- Sizes
- 7B – 21B
- Hardware
- from: Laptop
Math and reasoningGGUF2026
LM Provers (CMU, Hugging Face, ETH Zurich, Project Numina) · USA, Switzerland, France
A small 4B model on Qwen3 that writes mathematical proofs in plain language almost at the level of large models. Runs on a laptop.
- Checking the logic of reasoning and workings
- Step-by-step explanations of solutions
- Training and olympiad preparation
- Sizes
- 4B
- Hardware
- from: Laptop
MedicineOllama2025–2026
Google · USA
Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.
- Draft discharge summaries and reports for a doctor to review
- Hints for doctors when reviewing images
- Searching and summarising medical literature
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Ant Healthcare (Ant Group) and Zhejiang Provincial Medical Information Center · China
A large medical MoE model based on Ling-flash-2.0: 100B parameters with 6B active, so it answers quickly. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff on clinical questions
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 100B-A6B
- Hardware
- from: 1 GPU
TextGGUF2023–2026
Baichuan Intelligence · China
First general-purpose Chinese models, then a medical line from 2025. Baichuan-M3 (based on Qwen3-235B) is trained to model a doctor's clinical reasoning.
- Reference assistant for doctors
- Preliminary patient intake questions
- Analysis of medical documents
- Sizes
- 7B – 235B-A22B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Technology Innovation Institute (TII) · UAE
A family from Abu Dhabi: from the early Falcon 40B and 180B to hybrid Falcon-H1 and tiny Falcon-H1-Tiny models of 90–600M parameters for devices.
- Assistant and answers based on documents
- Running on low-end hardware and devices
- Tool calling in simple agents
- Sizes
- 90M – 180B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
TextOllama2024–2026
Allen Institute for AI (Ai2) · USA
Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.
- Assistant and answers based on documents
- Reasoning tasks (Think versions)
- Fine-tuning on your data with a clear model history
- Sizes
- 1B – 32B
- Hardware
- from: Laptop
TextGGUF2024–2026
AI21 Labs · Israel
A hybrid of Transformer and Mamba with a window of up to 256K tokens: handles long documents faster than conventional models. Jamba2 focuses on accurate, source-based answers.
- Answers based on long policies and contracts
- Knowledge-base search (RAG)
- Summaries of large documents
- Sizes
- 3B – 398B-A94B
- Hardware
- from: Laptop
TextGGUF2024–2026
Prime Intellect · USA
Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.
- Reasoning and math tasks
- Programming help
- Agents with tool calling
- Sizes
- 10B – 106B-A12B
- Hardware
- from: 1 GPU
TextRUGGUF2024–2026
UTTER consortium (Unbabel, universities of Lisbon, Edinburgh, Amsterdam and others) · European Union
European language models trained on all EU languages and several others, with a focus on translation. Russian is supported. Permissive license.
- Translation and localization of texts
- Answering questions in different languages
- Draft emails for foreign partners
- Sizes
- 1.7B – 22B
- Hardware
- from: Laptop
Cybersecurity2025–2026
Cisco (Foundation AI) · USA
Cisco models for information security based on Llama 3.1 8B: analysis of vulnerabilities, threats and incidents. Can be deployed inside your own perimeter.
- Analyzing vulnerability and threat reports
- Helping SOC analysts during incidents
- Mapping threats to MITRE ATT&CK
- Sizes
- 8B
- Hardware
- from: Laptop
Text2025
Naver · South Korea
Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.
- Lightweight Korean-English assistant
- Analysis of images and documents
- Text classification
- Sizes
- 0.5B – 32B
- Hardware
- from: Laptop
TextRUGGUF2024–2025
Vikhr Models · Russia
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
- Russian-language assistant on your own PC or server
- Knowledge-base answers (RAG)
- Texts and emails in Russian
- Sizes
- 0.5B – 24B
- Hardware
- from: Laptop
TextGGUF2023–2025
Inception (G42), MBZUAI and Cerebras · UAE
A model family for Arabic and English, including Gulf dialects. Suits companies working with Arabic-speaking customers and government bodies in the region.
- A chatbot in Arabic and English
- Translating and summarising documents in Arabic
- Classifying customer requests
- Sizes
- 256M – 70B
- Hardware
- from: Laptop
Tabular dataGGUF2024–2025
Zhejiang University · China
A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.
- Answering questions about tables and data exports
- Automated data analysis with generated code
- A helper for BI and internal reporting
- Sizes
- 7B – 72B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek's maths models. The first 7B version introduced the GRPO training method; the 685B V2 writes and checks its own olympiad-level proofs.
- Calculations and formula checks
- Checking mathematical workings in reports
- Working through problems step by step
- Sizes
- 7B – 685B
- Hardware
- from: Laptop
TextOllama2023–2025
Nous Research · USA
Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.
- Agents that call functions and APIs
- Data extraction in strict JSON format
- Assistant with flexible role and tone settings
- Sizes
- 3B – 405B
- Hardware
- from: Laptop
TextOllama2025
Deep Cogito · USA
Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.
- A chat assistant with a reasoning mode
- Writing code and calling tools
- Answering complex questions about documents
- Sizes
- 3B – 671B
- Hardware
- from: Laptop
TextRU2025
Avito Tech · Russia
Avito's model based on Qwen3-8B, retrained for Russian: its own tokenizer makes Russian text 15–25% faster. Supports function calling.
- Product and listing descriptions in Russian
- Chatbot that calls internal services
- Request analysis and classification
- Sizes
- 7.9B
- Hardware
- from: Laptop
TextOllama2025
OpenAI · USA
OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.
- AI agent that calls internal systems
- Answers based on internal policies
- Drafts of emails and reports
- Sizes
- 20B, 120B
- Hardware
- from: 1 GPU
TextRU2024–2025
Lomonosov Moscow State University Research Computing Center, LAIR lab (RefalMachine) · Russia
Qwen models adapted for Russian: a new tokenizer plus further training on Russian texts. As a result, Russian text is generated up to twice as fast as with the original model of the same size.
- Russian-language assistant on your own server
- Answers based on company documents (RAG) in Russian
- Analysis and summaries of long Russian texts
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
TextGGUF2025
ByteDance · China
An open ByteDance 36B model with up to 512K tokens of context and an adjustable thinking budget. Fits on a single powerful GPU.
- Analysis of long documents
- Agents with tools
- Corporate assistant
- Sizes
- 36B
- Hardware
- from: 1 GPU
Text2024–2025
xAI · USA
xAI publishes the weights of previous Grok generations. The models are very large and need a GPU cluster, so in practice they are rarely run.
- Research on large models
- Assistant on your own infrastructure
- Text generation and analysis
- Sizes
- 314B (Grok-1), Grok-2 is larger
- Hardware
- from: Cluster
MedicineGGUF2025
Intelligent Internet · UK
Reasoning medical models on Qwen3, designed to run on an ordinary computer. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff with the reasoning shown
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
TextRUOllama2024–2025
Hugging Face · USA
Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.
- Simple on-device assistant
- Classification and routing of requests
- Base for fine-tuning on a narrow task
- Sizes
- 135M – 3B
- Hardware
- from: Laptop
Math and reasoning2025
NVIDIA · USA
NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Working through programming problems
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
TranslationRUGGUF2024–2025
Unbabel · Portugal
Language models tailored for translation and multilingual text work: translating, editing, and assessing translation quality. Russian is supported. Non-commercial license.
- Translation that respects context and terminology
- Post-editing machine translation
- Assessing the quality of a finished translation
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Clouditera · China
A Chinese open family for cybersecurity: reviewing vulnerabilities, analysing logs and traffic, explaining commands and scripts.
- Reviewing vulnerabilities and drafting fix recommendations
- Analysing logs and reconstructing an attack chain
- Explaining suspicious commands and scripts
- Sizes
- 1.5B – 14B
- Hardware
- from: Laptop
CybersecurityGGUF2025
Trendyol · Turkey
Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.
- Reviewing alerts and first-pass incident assessment
- Explaining suspicious activity in reports
- Helping the on-duty shift of a monitoring centre
- Sizes
- 32B и 70B
- Hardware
- from: 1 GPU
TextOllama2025
DeepSeek · China
A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.
- Complex calculations and logic checks
- Analysis of contracts and internal policies
- Help for developers
- Sizes
- 1,5B – 671B
- Hardware
- from: Laptop
CodeRU2025
MTS AI (MWS AI) · Russia
A small coding assistant from MTS AI that understands requests in Russian. Runs locally, with plugins for VS Code and JetBrains.
- Code suggestions and completion in the editor
- Code explanations in Russian
- Drafts of tests and documentation
- Sizes
- 1.5B
- Hardware
- from: Laptop
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
TextRUGGUF2023–2025
Ilya Gusev (IlyaGusev) · Russia
The best-known Russian community fine-tune: open models (Llama, Mistral, Gemma, YandexGPT) trained to act as a Russian-speaking assistant. A convenient starting point for a Russian chatbot on your own server.
- Russian-language chat assistant
- Answers based on the company knowledge base
- Drafts of emails, descriptions and posts in Russian
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of contracts and internal policies
- Sizes
- 32B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Stanford University · USA
A reasoning model trained on just a thousand problems. It can be told to think longer to answer a hard question more accurately.
- Calculations and formula checks
- Working through complex problems step by step
- Training
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Qihoo 360 · China
Reasoning models from Qihoo 360: a standard Qwen2.5 was fine-tuned for long reasoning using an open recipe; data and code are published.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Computer-use agents2024–2025
Salesforce · USA
Salesforce models for function calling and agents: they pick the right tool and fill in its parameters. Strong on benchmarks, but the license is non-commercial.
- Calling APIs and internal services on user request
- Multi-step agents with several tools
- Comparing approaches before choosing a commercial model
- Sizes
- 1B – 8x22B
- Hardware
- from: Laptop
Finance2025
Shanghai University of Finance and Economics (SUFE) · China
A reasoning model for financial tasks based on Qwen2.5-7B: calculations, report analysis, regulatory questions. Trained on Chinese and English data.
- Financial calculations with step-by-step explanations
- Answering questions about financial statements
- Analyzing tables of financial data
- Sizes
- 7B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
NovaSky (Sky Computing Lab, Berkeley) · USA
A Berkeley reasoning model trained for under 450 dollars. It showed that o1-preview-level reasoning can be reproduced with modest resources.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
TextOllama2023–2025
Ai2 · USA
Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.
- Employee assistant on your own server
- Math and precise instruction following
- Reference recipe for your own fine-tuning
- Sizes
- 7B – 405B
- Hardware
- from: Laptop
Cybersecurity2025
Trend Micro · Japan
Trend Micro cybersecurity models based on Llama 3.1 8B, fine-tuned on a corpus of security texts. A reasoning version is available.
- Answering questions about threats and vulnerabilities
- Analyzing cyberattack reports
- A base for fine-tuning for SOC tasks
- Sizes
- 8B
- Hardware
- from: Laptop
TextGGUF2025
HUMAIN (formerly SDAIA) · Saudi Arabia
A Saudi model for Arabic and English, trained from scratch. One 7B version is openly available.
- An Arabic-language assistant
- Answering questions about documents
- Writing and editing texts in Arabic
- Sizes
- 7B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.
- Calculations and formula checks
- Checking calculations in estimates and reports
- Working through problems step by step
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Finance2023–2024
Du Xiaoman (Duxiaoman-DI) · China
A large Chinese model family for the financial industry: advice, document reading and long texts up to 8k-16k. Not investment advice: decisions are made by a specialist.
- Answering customer questions about banking products
- Working through long financial documents
- An internal assistant for a finance company regulations
- Sizes
- 6B – 176B
- Hardware
- from: 1 GPU
TextOllama2024
Nexusflow · USA
Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.
- Research on agents and function calling
- Comparison with commercial models
- Experiments with a chat assistant
- Sizes
- 70B – 72B
- Hardware
- from: 1 GPU
TextRU2024
MTS AI (MWS AI) · Russia
A lightweight Russian-language model from MTS AI for Russian texts: answers, summaries, drafts. A ready version for CPU without a GPU is available. The larger Cotype Pro is not released openly.
- Drafts of emails and descriptions in Russian
- Short document summaries
- Answers to common customer questions
- Sizes
- 1.5B
- Hardware
- from: Laptop
Finance2023–2024
The Fin AI / ChanceFocus · international project
One of the first open model families for financial text: reading statements, news and questions about numbers. Not investment advice: decisions are made by a specialist.
- Reading financial statements and press releases
- Answering questions about numeric data in documents
- Classifying financial texts
- Sizes
- 0.5B – 30B
- Hardware
- from: Laptop
Finance2023–2024
AI4Finance Foundation · USA
An open set of lightweight add-ons for ordinary language models that work with financial texts and news. Not investment advice: decisions are made by a specialist.
- Assessing the tone of financial news and reports
- Tagging mentions of companies and instruments in text
- Preparing digests from a stream of business news
- Sizes
- adapters for 6B - 20B base models
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
Stability AI · UK
Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.
- A lightweight chatbot on an ordinary PC
- Code autocompletion in the editor
- A base for fine-tuning on your own task
- Sizes
- 1.6B – 12B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2023–2024
M42 Health · UAE
Clinical models from Abu Dhabi-based M42, built on Llama and tuned to answer medical questions. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff on clinical questions
- Draft discharge summaries and letters for a doctor to review
- Searching medical literature
- Sizes
- 8B – 70B
- Hardware
- from: Laptop
FinanceNot maintained2024
Writer · USA
A large model for financial documents with a context window of about 131k tokens: it holds long reports whole. Not investment advice: decisions are made by a specialist.
- Working with long annual reports and prospectuses
- Summaries and digests of financial documents
- Finding answers inside a large document pack
- Sizes
- 70B (the model card states 72 billion parameters)
- Hardware
- from: Cluster
TextOllamaNot maintained2023–2024
OpenChat (Tsinghua University) · China
Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.
- Chat assistant on an inexpensive server
- Drafts of emails and replies
- Help with simple code
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
01.AI · China
Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.
- Chat assistant on a single GPU
- Analysis of long documents
- Classification and data extraction from text
- Sizes
- 6B – 34B
- Hardware
- from: Laptop
TextNot maintained2024
Equall · France
Language models for legal texts, fine-tuned on US and European legal corpora (based on Mistral and Mixtral). English only.
- Reviewing English-language contracts
- Spotting risks and non-standard terms
- Drafting legal memos
- Sizes
- 7B – 141B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2024
Avignon University and Nantes University · France
Mistral 7B fine-tuned on PubMed Central papers, plus several merges with the general model. Compact and easy to run. Does not replace a doctor; decisions are made by a specialist.
- Searching and summarising medical papers
- Draft reference materials for staff
- Explaining medical terminology
- Sizes
- 7B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2024
Saama AI Research · India
Llama 3 fine-tuned on medical and biological data. One of the first strong open medical models of 2024. Does not replace a doctor; decisions are made by a specialist.
- Extracting data from medical documents
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 8B, 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Hugging Face (H4) · USA
Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.
- Lightweight chat assistant
- Reference and starting point for your own fine-tuning
- Drafts of texts and replies
- Sizes
- 7B – 141B-A35B
- Hardware
- from: Laptop
TextRUNot maintained2023–2024
Sber (ai-forever) · Russia
Sber's Russian text-to-text model, successor to ruT5 (2021). Small and fast: fine-tuned for summarizing, paraphrasing and fixing errors in Russian text; ready-made SAGE spell-checking versions exist.
- Fixing spelling mistakes and typos in Russian text
- Short summaries and paraphrasing
- Normalizing requests and inquiries before processing
- Sizes
- 95M – 1.7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
CybersecurityGGUFNot maintained2023–2024
ZySec AI · India
A small open assistant for security professionals: questions about standards, reviewing threats and vulnerabilities, drafting internal documents.
- Answering questions about security policies and standards
- First-pass review of threat reports
- Drafting internal protection guidelines
- Sizes
- 2.8B и 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
WizardLM (Microsoft and Peking University) · USA / China
Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.
- Complex multi-step instructions
- Help for developers
- Solving math problems
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Intel · USA
A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.
- A simple chat assistant
- Experiments with running on Intel hardware
- A base for fine-tuning
- Sizes
- 7B
- Hardware
- from: Laptop
FinanceGGUFNot maintained2023
AdaptLLM · not disclosed
Finance-tuned versions of Llama 2: reading industry texts, reports and questions about terminology. Not investment advice: decisions are made by a specialist.
- Reading financial news and reports
- Answering questions about financial terminology
- A base for fine-tuning to your own financial task
- Sizes
- 7B и 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Microsoft Research · USA
Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.
- Research on reasoning methods
- Comparison with modern small models
- Training specialists
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
Tabular dataNot maintained2023
OSU NLP Group, Ohio State University · USA
A general-purpose model for tables: filling gaps, finding rows, matching columns and answering questions about the data.
- Answering questions about tables inside documents
- Matching columns across different tables
- Finding and completing records in reference books
- Sizes
- 7B
- Hardware
- from: Laptop
Text analysisNot maintained2022–2023
Mike Zhang, Rob van der Goot, Barbara Plank (IT University of Copenhagen and LMU Munich) · Denmark
A research line of models for labour market texts: trained on job postings and the European ESCO occupation taxonomy, they pull skills and requirements out of vacancies. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting skills and requirements from vacancy text
- Mapping skills to the single ESCO reference list
- Classifying vacancies and job titles
- Sizes
- 110M – 560M
- Hardware
- from: Laptop
FinanceNot maintained2023
Fudan-DISC, Fudan University · China
A financial assistant made of several fine-tuned experts: advice, calculations, document reading and knowledge-base search. Not investment advice: decisions are made by a specialist.
- In-house advice on financial questions
- Reading financial documents and news
- Prompts for front-office staff
- Sizes
- 13B
- Hardware
- from: 1 GPU
MedicineNot maintained2023
Bo Wang's lab (University of Toronto, Vector Institute) · Canada
An early medical model on Llama 2 70B, trained on dialogues based on medical texts. Now mainly of research interest. Does not replace a doctor; decisions are made by a specialist.
- Research pilots on medical dialogue
- Training materials for staff
- Comparison with newer medical models
- Sizes
- 70B
- Hardware
- from: 1 GPU
TextRUGGUFNot maintained2022–2023
Sber (ai-forever) · Russia
Sber's multilingual model covering 61 languages, including languages of the peoples of Russia and the CIS. Separate fine-tunes exist for Buryat, Yakut, Tatar, Bashkir, Kazakh and others, rare for open models.
- Texts in languages of the peoples of Russia and the CIS
- Base for fine-tuning on a less common language
- Drafts and templates in several languages
- Sizes
- 1.3B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
LMSYS (Berkeley and partners) · USA
One of the first open chat models (2023): LLaMA fine-tuned on user conversations with ChatGPT. A historical milestone; today it is weaker than any modern model of the same size.
- Experiments and team training
- Simple chat assistant for tests
- Comparison with newer models
- Sizes
- 7B – 33B
- Hardware
- from: Laptop
TextRUGGUFNot maintained2023
Sber (ai-forever) · Russia
Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.
- Base for fine-tuning on a narrow Russian-language task
- Generating template Russian texts
- Experiments with Russian-language models without license restrictions
- Sizes
- 13B
- Hardware
- from: 1 GPU
TextNot maintained2022–2023
Google · USA
Compact input-output models trained to follow instructions. Still used as a cheap base for classification, extraction and short answers.
- Classification of requests and documents
- Extracting fields from text
- Short answers and summaries
- Sizes
- 80M – 20B
- Hardware
- from: Laptop
TextNot maintained2022–2023
EleutherAI · USA
Fully open models from the non-profit lab EleutherAI: GPT-NeoX-20B and the Pythia series with published intermediate training checkpoints.
- Base model for fine-tuning
- Research into model behavior
- Simple text generation and completion
- Sizes
- 70M – 20B
- Hardware
- from: Laptop
TextNot maintained2022
Meta · USA
An early open Meta series matching GPT-3 in size. Outdated; useful for research and comparison.
- Research experiments
- Training specialists
- Comparison with modern models
- Sizes
- 125M – 66B (175B on request)
- Hardware
- from: Laptop
TextGGUFNot maintained2022
BigScience (Hugging Face and the community) · France
One of the first large open models, trained by a community of hundreds of researchers in 46 languages. Today it is interesting mostly as a historical milestone.
- Text generation and translation in many languages
- Experiments and team training
- Base model for fine-tuning on a narrow task
- Sizes
- 560M – 176B
- Hardware
- from: Laptop