Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO)
The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.
The last open version came out in Jun 2022. The family has not been updated for a long time: the model still works, but do not expect fixes or new sizes.
- Developer
- UKP Lab and the Sentence Transformers community, Germany
- First release
- Mar 2022
- Latest release
- Jun 2022
- Sizes
- about 4M – 120M
- License
- Commercial use allowedApache 2.0
- Russian
- Supported
- Running
- On your own serverAlso runs without a GPU
- Industries
- Customer support, Documents and accounting, Retail and marketplaces, Software development
What it does
- Reordering knowledge base search results
- Selecting passages before a chatbot answers
- Finding duplicates among tickets and product cards
Where it is used
Hardware requirements
Versions
- mmarco-mMiniLMv2-L12 (14 языков, в том числе русский)
- ms-marco TinyBERT-L2, MiniLM-L2/L4/L6/L12, electra-base
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO) be used in a commercial project?
Yes. License: Apache 2.0. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO) need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Some versions also run on an ordinary CPU, without a GPU. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO) support Russian?
Yes, Russian is listed on the model card.
Where can I download Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO) and what does it cost?
The Cross-Encoder MS MARCO (MiniLM, TinyBERT, mMARCO) weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.
DetailsRerankersJina RerankerJina AI · GermanyCommercial use with conditionsStrong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.
DetailsRerankersColBERT (поиск с поздним взаимодействием)Stanford NLP, later Answer.AI and LightOn · USA and FranceCommercial use with conditionsA different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.
DetailsSource: huggingface.co/cross-encoder/mmarco-mMiniLMv2-L12-H384-v1. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


