Open-source models for fact-checking and evaluating AI answers

Judge models check the output of other AI models: they spot fabricated facts, verify answers against sources and score quality. Teams use them to monitor chatbots and assistants before launch and in production. Check which criteria the model evaluates, how it performs in your languages and what the license permits.

12 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Fact-checking and judgesRU2025–2026

POLLUX Judge (ai-forever)

SberDevices (ai-forever) · Russia

Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.

  • Automatic quality checks of Russian chatbot answers
  • Comparing several models before choosing one
  • Checking answers after fine-tuning
Sizes
4B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judges2024–2025

CompassJudger / CompassVerifier

OpenCompass (Shanghai AI Laboratory) · China

A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.

  • Scoring model answers against set criteria
  • Checking an answer against a reference solution
  • Comparing several models on your own data
Sizes
1.5B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judges2024–2025

Skywork-Reward

Skywork (Kunlun Tech) · China

Reward models: they score how good a language model's answer is for the user. Used for fine-tuning your own models and picking the best of several answers.

  • Choosing the best of several bot answers
  • Scoring answer quality during model fine-tuning
  • Comparing models before rollout
Sizes
0.6B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judgesRU2024–2025

Nemotron Reward / GenRM

NVIDIA · USA

Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.

  • Choosing the best of several candidate answers
  • Preparing data to fine-tune your own model
  • Scoring assistant answers in Russian and other languages
Sizes
49B, 70B and 340B
Hardware
from: Cluster
Commercial use with conditionsDetails
Fact-checking and judges2025

Atla Selene

Atla · UK

An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.

  • Scoring chatbot answers against your own criteria
  • Comparing two versions of a prompt or model
  • Filtering out weak answers before they reach a person
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Fact-checking and judges2024

GLIDER (Patronus)

Patronus AI · USA

A small judge: it scores against your criteria and highlights which part of the answer led to that score. The license is non-commercial. The judge itself makes mistakes and does not replace manual review.

  • Scoring answers against your criteria with an explanation
  • Understanding why a score was lowered
  • Bulk review of assistant conversations
Sizes
3.8B (based on Phi-3.5-mini)
Hardware
from: Laptop
Non-commercial onlyDetails
Fact-checking and judges2024

Flow Judge

Flow AI · not disclosed

A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.

  • Checking AI assistant answers against the instruction
  • Bulk scoring of exported conversations
  • Quality control before rolling out changes
Sizes
3.8B (based on Phi-3.5-mini)
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judgesOllamaNot maintained2024

MiniCheck

UT Austin and Bespoke Labs · USA

Checks whether each claim in an AI answer is supported by the source documents. The small versions are free; the larger 7B is in Ollama but non-commercial.

  • Checking RAG bot answers against documents
  • Finding unsupported claims in reports and summaries
  • Automated quality control of AI answers
Sizes
0.4B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Fact-checking and judgesNot maintained2023–2024

HHEM (Vectara)

Vectara · USA

A small model that checks whether an AI answer is grounded in the source text or made up. Runs on a CPU and works well as a filter in RAG systems.

  • Checking knowledge base chatbot answers for fabrications
  • Quality control of document summaries
  • Comparing language models by their tendency to make errors
Sizes
110M
Hardware
from: Laptop
Commercial use allowedDetails
Fact-checking and judgesNot maintained2024

Patronus Lynx

Patronus AI · USA

Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.

  • Finding invented facts in AI assistant answers
  • Checking that answers rest on the attached documents
  • Filtering out answers before they go to a customer
Sizes
8B and 70B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Fact-checking and judgesNot maintained2024

ArmoRM (RLHFlow)

RLHFlow · USA

An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.

  • Choosing the best of several candidate answers
  • Preparing data for model fine-tuning
  • Scoring assistant answers across several attributes
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Fact-checking and judgesNot maintained2023–2024

Prometheus 2

KAIST and LG AI Research (prometheus-eval) · South Korea

An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.

  • Scoring chatbot answers on your own scale
  • Comparing two answer options
  • Quality checks before launching an AI service
Sizes
7B – 8x7B
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment