Open-source models for math and reasoning

Reasoning models solve problems step by step: they calculate, check intermediate conclusions and handle complex logic better than ordinary chat models. They suit calculations, analytics and tasks where accuracy matters. Keep in mind these answers take longer and use more compute, and check the license.

27 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Text2024–2026

K2 (K2-Think, K2-V2, K2-Horizon)

MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE

Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.

  • Reasoning, maths and technical questions
  • Analysing long documents
  • Agents and writing code
Sizes
0.9B – 375B-A23B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningOllama2025–2026

OpenThinker

Open Thoughts (Stanford, Berkeley and other universities) · USA

Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Checking the logic of internal policies
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

Apriel

ServiceNow · USA

ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.

  • A reasoning assistant for internal services
  • Tool calling and enterprise agents
  • Analysing screenshots and documents with images
Sizes
5B – 15B
Hardware
from: Laptop
Commercial use allowedDetails
CodeOllama2025–2026

Rnj-1

Essential AI · USA

An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.

  • Writing and fixing code
  • A developer agent on a single GPU
  • Solving technical and scientific problems
Sizes
8B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningGGUF2025–2026

Goedel-Prover

Princeton University · USA

Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.

  • Formal verification of mathematical workings
  • Verifying code correctness
  • Training
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningGGUF2026

QED-Nano

LM Provers (CMU, Hugging Face, ETH Zurich, Project Numina) · USA, Switzerland, France

A small 4B model on Qwen3 that writes mathematical proofs in plain language almost at the level of large models. Runs on a laptop.

  • Checking the logic of reasoning and workings
  • Step-by-step explanations of solutions
  • Training and olympiad preparation
Sizes
4B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningGGUF2024–2025

DeepSeek-Math

DeepSeek · China

DeepSeek's maths models. The first 7B version introduced the GRPO training method; the 685B V2 writes and checks its own olympiad-level proofs.

  • Calculations and formula checks
  • Checking mathematical workings in reports
  • Working through problems step by step
Sizes
7B – 685B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2025

Hermes

Nous Research · USA

Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.

  • Agents that call functions and APIs
  • Data extraction in strict JSON format
  • Assistant with flexible role and tone settings
Sizes
3B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2025

Cogito

Deep Cogito · USA

Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.

  • A chat assistant with a reasoning mode
  • Writing code and calling tools
  • Answering complex questions about documents
Sizes
3B – 671B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safetyOllama2025

gpt-oss-safeguard

OpenAI · USA

Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.

  • Moderation by internal company rules
  • Labeling disputed messages with an explanation
  • Checking reviews and listings before publishing
Sizes
20B – 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
Math and reasoningGGUF2025

Kimina-Prover

Moonshot AI and Project Numina · China, France

Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.

  • Formal verification of mathematical workings
  • Translating a problem from plain language into Lean
  • Training and olympiad preparation
Sizes
0.6B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
TextRU2024–2025

Ruadapt (RuadaptQwen)

Lomonosov Moscow State University Research Computing Center, LAIR lab (RefalMachine) · Russia

Qwen models adapted for Russian: a new tokenizer plus further training on Russian texts. As a result, Russian text is generated up to twice as fast as with the original model of the same size.

  • Russian-language assistant on your own server
  • Answers based on company documents (RAG) in Russian
  • Analysis and summaries of long Russian texts
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
MedicineGGUF2025

II-Medical

Intelligent Internet · UK

Reasoning medical models on Qwen3, designed to run on an ordinary computer. Does not replace a doctor; decisions are made by a specialist.

  • Reference answers to staff with the reasoning shown
  • Draft discharge summaries for a doctor to review
  • Searching medical literature
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningOllama2025

DeepScaleR, DeepCoder, DeepSWE

Agentica (Berkeley, Sky Computing Lab) and Together AI · USA

Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.

  • Solving maths problems with step-by-step working
  • Generating and checking code
  • An agent for fixing bugs in a repository
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoning2025

AceMath / AceReason

NVIDIA · USA

NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Working through programming problems
Sizes
1.5B – 72B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to SQLGGUF2025

SQL-R1

IDEA Research · China

A text-to-SQL model trained with reinforcement learning: it works through the schema and the conditions step by step before producing a query.

  • Database queries for questions with several conditions
  • Reviewing and fixing other people SQL queries
  • An analyst helper inside a BI system
Sizes
3B – 14B
Hardware
from: Laptop
Commercial use allowedDetails
Finance2025

Fino1 / Fin-o1

The Fin AI · international project

Models that spell out their reasoning before answering financial questions with numbers and tables. Not investment advice: decisions are made by a specialist.

  • Calculation questions on statements with the working shown
  • Working through tasks with tables and numbers from documents
  • Checking calculations made by hand
Sizes
8B и 14B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Math and reasoningGGUF2024–2025

DeepSeek-Prover

DeepSeek · China

DeepSeek models for formal proofs in Lean 4: the proof is checked by a program, not a person. A narrow tool for mathematicians and engineers.

  • Formal verification of mathematical workings
  • Verifying algorithm correctness
  • Training and olympiad preparation
Sizes
7B – 671B
Hardware
from: Laptop
Commercial use with conditionsDetails
Math and reasoningOllama2024–2025

QwQ

Qwen (Alibaba) · China

Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.

  • Calculations and formula checks
  • Complex analytics with step-by-step breakdowns
  • Checking the logic of contracts and internal policies
Sizes
32B
Hardware
from: 1 GPU
Commercial use allowedDetails
Math and reasoningGGUF2025

s1

Stanford University · USA

A reasoning model trained on just a thousand problems. It can be told to think longer to answer a hard question more accurately.

  • Calculations and formula checks
  • Working through complex problems step by step
  • Training
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningGGUF2025

Light-R1

Qihoo 360 · China

Reasoning models from Qihoo 360: a standard Qwen2.5 was fine-tuned for long reasoning using an open recipe; data and code are published.

  • Calculations and formula checks
  • Working through problems step by step
  • A base for your own reasoning fine-tuning
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Finance2025

Fin-R1

Shanghai University of Finance and Economics (SUFE) · China

A reasoning model for financial tasks based on Qwen2.5-7B: calculations, report analysis, regulatory questions. Trained on Chinese and English data.

  • Financial calculations with step-by-step explanations
  • Answering questions about financial statements
  • Analyzing tables of financial data
Sizes
7B
Hardware
from: Laptop
Commercial use allowedDetails
Math and reasoningGGUF2025

Sky-T1

NovaSky (Sky Computing Lab, Berkeley) · USA

A Berkeley reasoning model trained for under 450 dollars. It showed that o1-preview-level reasoning can be reproduced with modest resources.

  • Calculations and formula checks
  • Working through problems step by step
  • A base for your own reasoning fine-tuning
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2025

Tulu

Ai2 · USA

Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.

  • Employee assistant on your own server
  • Math and precise instruction following
  • Reference recipe for your own fine-tuning
Sizes
7B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Math and reasoningOllama2024–2025

Qwen2.5-Math

Qwen (Alibaba) · China

Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.

  • Calculations and formula checks
  • Checking calculations in estimates and reports
  • Working through problems step by step
Sizes
1.5B – 72B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023–2024

WizardLM, WizardCoder, WizardMath

WizardLM (Microsoft and Peking University) · USA / China

Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.

  • Complex multi-step instructions
  • Help for developers
  • Solving math problems
Sizes
7B – 70B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllamaNot maintained2023

Orca 2

Microsoft Research · USA

Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.

  • Research on reasoning methods
  • Comparison with modern small models
  • Training specialists
Sizes
7B – 13B
Hardware
from: Laptop
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment