Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
OpenCompass (Shanghai AI Laboratory) · China
A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring model answers against set criteria
- Checking an answer against a reference solution
- Comparing several models on your own data
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
Skywork (Kunlun Tech) · China
Reward models: they score how good a language model's answer is for the user. Used for fine-tuning your own models and picking the best of several answers.
- Choosing the best of several bot answers
- Scoring answer quality during model fine-tuning
- Comparing models before rollout
- Sizes
- 0.6B – 27B
- Hardware
- from: Laptop
Fact-checking and judgesRU2024–2025
NVIDIA · USA
Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data to fine-tune your own model
- Scoring assistant answers in Russian and other languages
- Sizes
- 49B, 70B and 340B
- Hardware
- from: Cluster
Fact-checking and judges2025
Atla · UK
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring chatbot answers against your own criteria
- Comparing two versions of a prompt or model
- Filtering out weak answers before they reach a person
- Sizes
- 8B
- Hardware
- from: 1 GPU
Fact-checking and judges2024
Patronus AI · USA
A small judge: it scores against your criteria and highlights which part of the answer led to that score. The license is non-commercial. The judge itself makes mistakes and does not replace manual review.
- Scoring answers against your criteria with an explanation
- Understanding why a score was lowered
- Bulk review of assistant conversations
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Fact-checking and judges2024
Flow AI · not disclosed
A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.
- Checking AI assistant answers against the instruction
- Bulk scoring of exported conversations
- Quality control before rolling out changes
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Fact-checking and judgesOllamaNot maintained2024
UT Austin and Bespoke Labs · USA
Checks whether each claim in an AI answer is supported by the source documents. The small versions are free; the larger 7B is in Ollama but non-commercial.
- Checking RAG bot answers against documents
- Finding unsupported claims in reports and summaries
- Automated quality control of AI answers
- Sizes
- 0.4B – 7B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
Vectara · USA
A small model that checks whether an AI answer is grounded in the source text or made up. Runs on a CPU and works well as a filter in RAG systems.
- Checking knowledge base chatbot answers for fabrications
- Quality control of document summaries
- Comparing language models by their tendency to make errors
- Sizes
- 110M
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
Patronus AI · USA
Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.
- Finding invented facts in AI assistant answers
- Checking that answers rest on the attached documents
- Filtering out answers before they go to a customer
- Sizes
- 8B and 70B
- Hardware
- from: 1 GPU
Fact-checking and judgesNot maintained2024
RLHFlow · USA
An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data for model fine-tuning
- Scoring assistant answers across several attributes
- Sizes
- 8B
- Hardware
- from: 1 GPU
Fact-checking and judgesNot maintained2023–2024
KAIST and LG AI Research (prometheus-eval) · South Korea
An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.
- Scoring chatbot answers on your own scale
- Comparing two answer options
- Quality checks before launching an AI service
- Sizes
- 7B – 8x7B
- Hardware
- from: Laptop