Open-source embedding models for search and RAG

Embedding models turn text into vectors so you can search by meaning rather than exact words. They power knowledge base search and RAG, where an assistant answers from your own documents. Check quality in your languages, maximum input length, vector size and the license.

43 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Music and soundGGUF2022–2026

MERT

m-a-p (Multimodal Art Projection) · UK / China

A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.

  • Automatic tagging of a music catalog
  • Finding similar tracks
  • Detecting genre, mood and tempo
Sizes
95M – 632M
Hardware
from: Laptop
Non-commercial onlyDetails
Visual document searchGGUF2026

EVIE

Tencent · China

Tencent models based on Qwen3.5 for searching scans and PDFs as images. According to the model card, among the top of the ViDoRe leaderboard at release.

  • Search across scans and PDFs without OCR
  • RAG over reports with tables and charts
  • Search across document archives
Sizes
4.5B – 8B
Hardware
from: 1 GPU
Commercial use allowedDetails
Search and RAGRU2024–2026

FRIDA / Giga-Embeddings

Sber (SberDevices) · Russia

Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.

  • Search across Russian-language documents
  • RAG for chatbots in Russian
  • Classifying requests and reviews
Sizes
480M – 10B-A1.8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRU2025–2026

NVIDIA Nemotron Embed

NVIDIA · USA

NVIDIA embeddings for search and RAG. Nemotron-3-Embed, released in 2026, is under the permissive OpenMDW license and works in many languages.

  • Search across corporate documents
  • RAG for chatbots and assistants
  • Search across images and pages (VL versions)
Sizes
1B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRUGGUF2023–2026

Jina Embeddings

Jina AI · Germany

Strong multilingual embeddings with long context; v5-omni understands text, images and audio. Recent versions are open for non-commercial use only.

  • Search across documents in many languages
  • Search across images and scans
  • Classification and clustering
Sizes
33M – 3.8B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRUOllama2024–2026

Granite Embedding

IBM · USA

Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.

  • Search across corporate documents
  • RAG on a regular server without a GPU
  • Reranking results
Sizes
30M – 311M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGGGUF2025–2026

Octen Embedding

Octen · USA / Singapore

Qwen3-Embedding models fine-tuned by the startup Octen for search in legal, financial and medical texts. As of January 2026 the 8B version topped the RTEB leaderboard.

  • Search across contracts and case law
  • Search across financial reports
  • Search across long documents up to 32K tokens
Sizes
0.6B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGGGUF2026

pplx-embed

Perplexity · USA

Embeddings from the Perplexity search service. Some versions take into account the context of the whole document, not just a single fragment.

  • Search across large document collections
  • RAG that accounts for document context
  • Website and catalog search
Sizes
0.6B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRU2026

Harrier (harrier-oss)

Microsoft · USA

Microsoft's 2026 multilingual embeddings with context up to 32K tokens; Russian is on the language list. The 270M and 0.6B versions run on a regular server, 27B is the most accurate.

  • Multilingual knowledge base search
  • Picking passages for RAG
  • Search across long documents
Sizes
270M – 27B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2025–2026

Qwen3 Embedding / Reranker

Alibaba (Qwen) · China

Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.

  • Knowledge base search for RAG
  • Reranking results before answering
  • Search across scans, slides and screenshots
Sizes
0.6B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAG2026

Voyage 4 nano

Voyage AI (MongoDB) · USA

The only open model in the Voyage 4 line: its vectors are compatible with the paid larger versions, so you can start locally and move to the API later.

  • Document search on your own server
  • RAG for small knowledge bases
  • Finding similar texts
Sizes
about 340M
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2025

Perception Encoder (PE)

Meta · USA

Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.

  • Search photos and videos by description
  • Catalog labeling and tagging
  • Search across audio and video (PE-AV)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2025

MetaCLIP / MetaCLIP 2

Meta · USA

Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.

  • Image search by text
  • Image classification without training
  • Search research and prototypes
Sizes
0.15B – 3.6B
Hardware
from: Laptop
Non-commercial onlyDetails
Search and RAGRUOllama2024–2025

mxbai (Mixedbread Embed и Rerank)

Mixedbread · Germany

Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.

  • Search across a knowledge base
  • Reranking results before a bot answers
  • Product catalog search
Sizes
17M – 1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2025

ModernVBERT / ColModernVBERT

Illuin Technology, EPFL, CentraleSupélec · France

A compact (250M) model for searching document pages as images. According to the authors, it matches models 10 times larger and runs without a GPU.

  • Search across scans and PDFs on a modest server
  • Indexing document archives
  • Search across slides and manuals
Sizes
250M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGOllama2025

EmbeddingGemma

Google · USA

A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.

  • On-device document search
  • RAG without sending data outside
  • Text classification
Sizes
300M
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGOllama2023–2025

BGE (BAAI General Embedding)

BAAI (Beijing Academy of Artificial Intelligence) · China

Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.

  • Search across English-language documents
  • Picking passages for chatbot answers (RAG)
  • Code search (bge-code)
Sizes
24M – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2024–2025

ColPali / ColQwen

Illuin Technology (ViDoRe team) · France

Searches PDFs and scans as images: pages do not need to be OCR'd first, the model finds the right one for a question directly, including tables and charts. Trained on English.

  • Search across scans, presentations and PDFs
  • RAG over documents with tables and charts
  • Search across technical documentation
Sizes
256M – 3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGGGUF2024–2025

JobBERT (TechWolf)

TechWolf · Belgium

A model from a Belgian HR company: it turns job titles into vectors so you can find similar vacancies and resumes. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Matching job titles coming from different sources
  • Finding similar vacancies and resumes by meaning
  • Cleaning up the company job title reference list
Sizes
109M – 278M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRU2023–2025

RuModernBERT и USER (deepvk)

deepvk (VK) · Russia

Russian encoders from the VK team: RuModernBERT reads long texts, USER produces vectors for search, GeRaCl classifies texts by topic without training.

  • Classifying requests without labeled data
  • Knowledge base search in Russian
  • Analyzing long contracts
Sizes
35M – 360M
Hardware
from: Laptop
Commercial use allowedDetails
Rerankers2023–2025

ColBERT (поиск с поздним взаимодействием)

Stanford NLP, later Answer.AI and LightOn · USA and France

A different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.

  • Search across a knowledge base of long documents
  • Reordering retrieved passages
  • Search across policies and technical documentation
Sizes
about 33M – 150M
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRUOllama2024–2025

Nomic Embed

Nomic AI · USA

Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.

  • Search across documents and a knowledge base
  • Code search
  • Search across scans and PDFs without text recognition
Sizes
137M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2025

Nomic Embed Multimodal / ColNomic

Nomic AI · USA

Search across PDF pages and scans as images. The cards list English, Italian, French, German and Spanish — Russian is not among them.

  • Search across an archive of scans and PDFs
  • Search across tables and diagrams inside documents
  • Picking pages for an AI assistant answer
Sizes
3B and 7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2023–2025

SigLIP (наследник CLIP)

Google · USA

Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.

  • Image search by text query
  • Automatic catalog labeling and tagging
  • Filtering prohibited content
Sizes
about 0.2B to 2B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAG2025

CareerBERT

The authors of the CareerBERT paper, German universities · Germany

A German-language model that matches a resume to occupations from the European ESCO reference list and suggests suitable directions. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Suggesting occupations that fit the candidate experience
  • Matching resumes against job descriptions
  • Hints on internal career moves
Sizes
110M
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRU2023–2025

GTE

Alibaba · China

Alibaba embeddings and rerankers for search: from tiny to 7B based on Qwen2. There is a multilingual mGTE version with long context.

  • Semantic search across documents
  • Reranking search results
  • Clustering and classifying texts
Sizes
33M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllama2019–2025

Sentence Transformers (all-MiniLM, paraphrase-multilingual)

UKP Lab (TU Darmstadt), later Hugging Face · Germany

The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.

  • Search across a knowledge base and FAQ
  • Finding similar tickets and duplicates
  • Grouping reviews and requests by topic
Sizes
about 20M – 470M
Hardware
from: Laptop
Commercial use allowedDetails
Visual document search2025

LlamaIndex vdr

LlamaIndex · USA

A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.

  • Search across scans and PDFs without OCR
  • Search across invoices, acts and contracts
  • Picking pages for an AI assistant answer
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails
Music and sound2024

MuQ / MuQ-MuLan

Tencent AI Lab · China

The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.

  • Searching music by text description
  • Tagging tracks by genre and mood
  • Finding similar music
Sizes
300M – 700M
Hardware
from: Laptop
Non-commercial onlyDetails
Search and RAGRUOllama2024

Snowflake Arctic Embed

Snowflake · USA

Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.

  • Search across documents and knowledge bases
  • Picking passages for RAG
  • Search across reports and internal data
Sizes
22M – 568M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAG2024

ConTeXT-Skill-Extraction

TechWolf · Belgium

Finds mentions of skills in a vacancy or resume text and maps them to the company skill reference list. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting skills from a job description
  • Matching candidate skills against requirements
  • Building a competence map across departments
Sizes
109M
Hardware
from: Laptop
Commercial use with conditionsDetails
Visual document search2024

GME (General Multimodal Embedding)

Alibaba (Tongyi Lab) · China

One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.

  • Finding a product by photo
  • Search across a catalogue of images and cards
  • Search across document pages as images
Sizes
2B and 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2024

VLM2Vec

TIGER-Lab · Canada

Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.

  • Search across a mixed archive of texts and images
  • Search across document pages as images
  • Finding similar cards and illustrations
Sizes
about 4B (based on Phi-3.5-V)
Hardware
from: 1 GPU
Commercial use allowedDetails
Visual document search2024

DSE (Document Screenshot Embedding)

University of Waterloo, Tevatron project · Canada

Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.

  • Search across scans and PDFs without OCR
  • Search across presentations and reports with complex layouts
  • Picking pages for an AI assistant answer
Sizes
2B (based on Qwen2-VL)
Hardware
from: 1 GPU
Commercial use allowedDetails
Text analysisRUNot maintained2020–2024

ruBERT, ruRoBERTa, ruELECTRA (ai-forever)

SberDevices (ai-forever) · Russia

Sber's Russian-language encoders trained on large Russian corpora. A base for classifiers, NER and semantic search in Russian.

  • Classifying requests in Russian
  • Extracting names, amounts and dates after fine-tuning
  • Detecting review sentiment
Sizes
about 30M to 430M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUNot maintained2022–2024

E5 / multilingual-e5

Microsoft · USA

Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.

  • Search across a knowledge base and documents
  • Finding answers for a chatbot (RAG)
  • Finding similar requests and duplicates
Sizes
33M – 7B
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRUOllamaNot maintained2024

BGE-M3

BAAI · China

A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.

  • Search across a document base
  • RAG for a chatbot
  • Finding similar requests and duplicates
Sizes
568M
Hardware
from: Laptop
Commercial use allowedDetails
RerankersNot maintained2023

BCEmbedding (Youdao)

NetEase Youdao · China

An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.

  • Search across a knowledge base and reference materials
  • Reordering retrieved passages
  • Picking answers for a support chatbot
Sizes
about 280M
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundNot maintained2023

CLAP (LAION)

LAION · Germany

CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.

  • Search sounds and music by description
  • Automatic tags for an audio library
  • Recognizing sound types (siren, breaking glass, voice)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundNot maintained2022

AST (Audio Spectrogram Transformer)

MIT · USA

A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.

  • Sound event recognition
  • Tagging an audio archive
  • Detecting alarm sounds
Sizes
about 87M
Hardware
from: Laptop
Commercial use allowedDetails
Computer visionNot maintained2021–2022

CLIP (OpenAI)

OpenAI · USA

The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.

  • Image search by text query
  • Automatic tags for a catalog
  • Finding similar images
Sizes
about 0.15B – 0.6B
Hardware
from: Laptop
Commercial use allowedDetails
Computer visionRUNot maintained2022

ruCLIP (ai-forever)

Sber AI and SberDevices (ai-forever) · Russia

A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.

  • Product search by photo and by Russian description
  • Sorting images into categories without labeling
  • Checking that a photo matches its caption
Sizes
150M – 430M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisRUNot maintained2021

rubert-tiny

David Dale (cointegrated) · Russia

A very small Russian-English BERT that runs fast on a regular CPU. Ready-made fine-tuned versions exist for sentiment, toxicity and emotions.

  • Detecting review sentiment
  • Filtering rude chat messages
  • Fast classification of requests
Sizes
12M – 29M
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment