Visual document searchGGUF2026
Tencent · China
Tencent models based on Qwen3.5 for searching scans and PDFs as images. According to the model card, among the top of the ViDoRe leaderboard at release.
- Search across scans and PDFs without OCR
- RAG over reports with tables and charts
- Search across document archives
- Sizes
- 4.5B – 8B
- Hardware
- from: 1 GPU
RerankersGGUF2024–2026
Jina AI · Germany
Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.
- Refining search results before a chatbot answers
- Sorting retrieved PDF pages and slides
- Catalog and knowledge base search
- Sizes
- 33M – 2.4B
- Hardware
- from: Laptop
Rerankers2025–2026
NVIDIA · USA
A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.
- Reordering passages before an AI assistant answers
- Sorting retrieved scan and PDF pages
- Search across internal policies and instructions
- Sizes
- 1B
- Hardware
- from: Laptop
Visual document search2025
Illuin Technology, EPFL, CentraleSupélec · France
A compact (250M) model for searching document pages as images. According to the authors, it matches models 10 times larger and runs without a GPU.
- Search across scans and PDFs on a modest server
- Indexing document archives
- Search across slides and manuals
- Sizes
- 250M
- Hardware
- from: Laptop
Visual document search2024–2025
Illuin Technology (ViDoRe team) · France
Searches PDFs and scans as images: pages do not need to be OCR'd first, the model finds the right one for a question directly, including tables and charts. Trained on English.
- Search across scans, presentations and PDFs
- RAG over documents with tables and charts
- Search across technical documentation
- Sizes
- 256M – 3B
- Hardware
- from: Laptop
Visual document search2025
Nomic AI · USA
Search across PDF pages and scans as images. The cards list English, Italian, French, German and Spanish — Russian is not among them.
- Search across an archive of scans and PDFs
- Search across tables and diagrams inside documents
- Picking pages for an AI assistant answer
- Sizes
- 3B and 7B
- Hardware
- from: 1 GPU
Visual document search2025
LlamaIndex · USA
A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.
- Search across scans and PDFs without OCR
- Search across invoices, acts and contracts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Visual document search2024
Alibaba (Tongyi Lab) · China
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
- Finding a product by photo
- Search across a catalogue of images and cards
- Search across document pages as images
- Sizes
- 2B and 7B
- Hardware
- from: 1 GPU
Visual document search2024
TIGER-Lab · Canada
Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.
- Search across a mixed archive of texts and images
- Search across document pages as images
- Finding similar cards and illustrations
- Sizes
- about 4B (based on Phi-3.5-V)
- Hardware
- from: 1 GPU
Visual document search2024
LightOn · France
A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.
- Refining search results over scans and PDFs
- Selecting pages before an AI assistant answers
- Sorting retrieved slides and reports
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Visual document search2024
University of Waterloo, Tevatron project · Canada
Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.
- Search across scans and PDFs without OCR
- Search across presentations and reports with complex layouts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU