Open-source rerankers for search and RAG

A reranker takes documents found by search and reorders them by how well they answer the question. This noticeably improves the accuracy of an assistant that answers from a knowledge base. Check quality in your languages, speed at your document volume and the license terms.

11 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
RerankersGGUF2024–2026

Jina Reranker

Jina AI · Germany

Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.

  • Refining search results before a chatbot answers
  • Sorting retrieved PDF pages and slides
  • Catalog and knowledge base search
Sizes
33M – 2.4B
Hardware
from: Laptop
Commercial use with conditionsDetails
Rerankers2026

R3 (R3-Embedding и R3-Rerank)

Tencent · China

A pair of small Tencent models based on Qwen3 that pick the right skill for an AI agent for a given request: the embedding model finds candidates, the reranker chooses the best one.

  • Choosing a tool or skill for an AI agent
  • Routing requests between bot scenarios
  • Search across a catalog of internal tools
Sizes
0.6B
Hardware
from: Laptop
Commercial use allowedDetails
Rerankers2025–2026

Llama Nemotron Rerank

NVIDIA · USA

A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.

  • Reordering passages before an AI assistant answers
  • Sorting retrieved scan and PDF pages
  • Search across internal policies and instructions
Sizes
1B
Hardware
from: Laptop
Commercial use with conditionsDetails
Rerankers2025

zerank (ZeroEntropy)

ZeroEntropy · USA

Rerankers built on Qwen3. The model card lists the target domains — finance, law, code, medicine, science; the stated language is English.

  • Refining results before an AI assistant answers
  • Sorting search results across contracts and reports
  • Search across technical and scientific documentation
Sizes
zerank-2 — 4B (based on Qwen3-4B), plus a smaller "small" version
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2024–2025

mxbai (Mixedbread Embed и Rerank)

Mixedbread · Germany

Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.

  • Search across a knowledge base
  • Reranking results before a bot answers
  • Product catalog search
Sizes
17M – 1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Rerankers2023–2025

ColBERT (поиск с поздним взаимодействием)

Stanford NLP, later Answer.AI and LightOn · USA and France

A different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.

  • Search across a knowledge base of long documents
  • Reordering retrieved passages
  • Search across policies and technical documentation
Sizes
about 33M – 150M
Hardware
from: Laptop
Commercial use with conditionsDetails
Visual document search2024

MonoQwen2-VL (LightOn)

LightOn · France

A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.

  • Refining search results over scans and PDFs
  • Selecting pages before an AI assistant answers
  • Sorting retrieved slides and reports
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails
RerankersGGUFNot maintained2023–2024

BGE Reranker

BAAI (Beijing Academy of Artificial Intelligence) · China

Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.

  • Refining search results before a chatbot answers
  • Sorting knowledge base search results
  • Selecting the most relevant clauses of contracts and policies
Sizes
278M – 9B
Hardware
from: Laptop
Commercial use allowedDetails
RerankersGGUFNot maintained2023–2024

RankVicuna / RankLLaMA / RankZephyr (Castorini)

University of Waterloo, Castorini group · Canada

Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.

  • Reordering a long list of search results
  • Selecting sources for an AI assistant answer
  • Research comparisons of retrieval approaches
Sizes
7B – 13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
RerankersNot maintained2023

BCEmbedding (Youdao)

NetEase Youdao · China

An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.

  • Search across a knowledge base and reference materials
  • Reordering retrieved passages
  • Picking answers for a support chatbot
Sizes
about 280M
Hardware
from: Laptop
Commercial use allowedDetails
RerankersRUNot maintained2022

Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO)

UKP Lab and the Sentence Transformers community · Germany

The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.

  • Reordering knowledge base search results
  • Selecting passages before a chatbot answers
  • Finding duplicates among tickets and product cards
Sizes
about 4M – 120M
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment