Documents and OCR2026
Jina AI · Germany
Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.
- Converting scans and PDFs to Markdown
- Recognizing tables and formulas
- Parsing invoices, acts and reports
- Sizes
- 3.4B-A0.6B
- Hardware
- from: 1 GPU
Documents and OCR2026
TeleAI (China Telecom) · China
A new lightweight document parsing model that led the OmniDocBench v1.6 benchmark at release. Handles pages photographed on a phone and crumpled pages well. Languages on the card: Chinese, English, Japanese.
- Recognising invoices and delivery notes photographed on a phone
- Recognising tables and formulas
- Converting documents to Markdown for RAG
- Sizes
- about 1.2B
- Hardware
- from: Laptop
Image + textGGUF2024–2026
Alibaba (AIDC-AI) · China
Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.
- Extracting data from invoices, contracts and delivery notes
- Table recognition
- Answering questions about photos and charts
- Sizes
- 0.9B – 80B-A3B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Tencent · China
A lightweight OCR model from Tencent: document parsing, finding text in photos, field extraction and translating text from images. Version 1.5 is faster and runs on an ordinary PC.
- Extracting fields from invoices and delivery notes
- Recognising tables and formulas
- Translating text in photos and scans
- Sizes
- 1B
- Hardware
- from: Laptop
Documents and OCRRU2024–2026
Datalab · USA
A compact OCR toolkit from the makers of Marker and Chandra: text recognition, page layout, reading order and tables. Surya OCR 2 (650M) also runs on a CPU; Russian scored 88.8% in benchmarks.
- Recognizing scans and PDFs, including in Russian
- Page layout: headings, tables, images, reading order
- Recognizing tables by rows and columns
- Sizes
- up to 650M
- Hardware
- from: Laptop
Documents and OCR2026
Baidu · China
Baidu's OCR model building on DeepSeek-OCR ideas: processes multi-page documents and PDFs in a single pass and outputs structured text. Claimed to be multilingual, but the language list is not published.
- Converting multi-page PDFs and scans to text and Markdown
- Recognizing contracts, invoices and reports
- Preparing document archives for search and RAG
- Sizes
- 3.3B
- Hardware
- from: Laptop
Documents and OCRRU2022–2026
Baidu (PaddlePaddle) · China
Classic lightweight PaddleOCR models: detecting and recognizing lines of text plus page layout. They run on CPUs and phones; there is a separate model for East Slavic languages, including Russian.
- Recognizing text on scans, photos and screens
- Reading labels, displays and markings in production and warehouses
- Page layout: tables, formulas, stamps, headings
- Sizes
- from 1.5M to tens of millions of parameters
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Baidu (PaddlePaddle) · China
A compact document parsing model from the popular PaddleOCR toolkit. Per the model card it supports 109 languages, including Russian; version 1.6 leads the OmniDocBench benchmark.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables, formulas and stamps
- Converting scans to Markdown and JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
Shanghai AI Laboratory (OpenDataLab) · China
A popular open tool for converting PDFs to Markdown with its own small model. MinerU2.5-Pro was improved through data alone, without growing in size. Languages on the card: Chinese and English.
- Converting PDF reports and contracts to Markdown
- Recognising tables and formulas
- Preparing documents for RAG and search
- Sizes
- 0.9B – 1.2B
- Hardware
- from: Laptop
Image + textOllama2025–2026
IBM · USA
Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.
- Extracting fields from forms and invoices
- Turning charts and tables into data
- Answering questions about documents
- Sizes
- 2B – 4B
- Hardware
- from: Laptop
Text analysisOllama2024–2026
NuMind · France
Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.
- Extracting company details, amounts and dates from invoices and contracts into JSON
- Parsing receipts, waybills and forms against a set template
- Converting scans to Markdown for search
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
rednote hilab (Xiaohongshu) · China
A multilingual document parsing model: text, tables, formulas and reading order in one pass. dots.mocr also turns charts and diagrams into vector SVG.
- Recognising invoices, contracts and delivery notes
- Converting tables into an editable format
- Converting charts and diagrams into vector format
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2025–2026
Datalab · USA
A strong OCR model from the authors of Marker and Surya: handwriting, forms, tables. Per the model card it supports 90+ languages, with Russian among the examples.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising handwritten forms and questionnaires
- Recognising complex tables
- Sizes
- 5B – 9B
- Hardware
- from: Laptop
Documents and OCRRUGGUF2026
Baidu (Qianfan) · China
A Baidu model that not only recognises a document but also answers questions about it. Per the model card it supports 192 languages, including Cyrillic.
- Recognising invoices, contracts and delivery notes, including in Russian
- Page layout analysis and table recognition
- Answering questions about a document
- Sizes
- 4B
- Hardware
- from: Laptop
Image + text2026
Sukhrob Nurali · not disclosed
A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.
- Moving a resume from PDF into a candidate record
- Filling a candidate database without manual typing
- Parsing resumes with different layouts and styling
- Sizes
- 8B, a fine-tune of Qwen3-VL-8B-Instruct
- Hardware
- from: 1 GPU
Documents and OCROllama2025–2026
DeepSeek · China
An OCR model that compresses a page into a small number of visual tokens, so it processes large volumes quickly. Version 2 better understands reading order.
- Bulk recognition of scanned invoices and contracts
- Table recognition
- Converting PDFs to Markdown for search and RAG
- Sizes
- about 3B
- Hardware
- from: Laptop
Documents and OCRRUOllama2026
Zhipu AI (Z.ai) · China
A lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising tables and formulas
- Extracting fields to JSON
- Sizes
- 0.9B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
LightOn · France
A French 1B OCR model that converts a page into text in one pass and is fast on high volumes. Languages on the card: European languages, Chinese and Japanese; no Russian.
- Recognising invoices and contracts in European languages
- Table recognition
- Converting PDFs to text for search and RAG
- Sizes
- 0.9B – 1B
- Hardware
- from: Laptop
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Documents and OCRGGUF2025
Ai2 (Allen Institute for AI) · USA
A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.
- Bulk digitisation of a PDF archive
- Converting contracts and reports into text
- Preparing documents for search and RAG
- Sizes
- 7B
- Hardware
- from: 1 GPU
Documents and OCRRUGGUF2025
Nanonets · USA / India
A model that converts documents to Markdown with tables, stamps, signatures, checkboxes and watermarks. The OCR2 model card lists Russian among its languages.
- Recognising invoices, contracts and delivery notes, including in Russian
- Recognising stamps, signatures and marks
- Handwriting recognition
- Sizes
- 1.5B – 3B
- Hardware
- from: Laptop
Image + textRU2025
Avito Tech · Russia
Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.
- Product descriptions from photos in Russian
- Checking that a photo matches its description
- Reading brands and text in images
- Sizes
- 7.4B
- Hardware
- from: 1 GPU
Documents and OCR2025
IBM and Hugging Face · USA
Tiny models for the open Docling document converter: they turn a page into markup with tables, formulas and code. Run on an ordinary laptop.
- Converting PDFs and scans to Markdown for search and RAG
- Recognising tables in reports
- Processing invoices and contracts on an ordinary PC
- Sizes
- 256M – 258M
- Hardware
- from: Laptop
Deepfake detection2022–2025
GRIP, University Federico II of Naples · Italy
Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.
- Checking scans of certificates and contracts for edits
- Highlighting altered photo regions for an expert
- Filtering out obviously redrawn documents before manual review
- Sizes
- a transformer model producing a map of suspicious regions
- Hardware
- from: 1 GPU
Documents and OCR2024
StepFun · China
One of the first general-purpose new-generation OCR models: text, formulas, tables, sheet music and diagrams. Small and runs on low-end hardware, but already behind newer models.
- Recognising scanned invoices and contracts
- Converting tables into an editable format
- Recognising formulas and diagrams
- Sizes
- 580M
- Hardware
- from: Laptop
Documents and OCRNot maintained2024
Microsoft · USA
Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.
- Converting resume scans into text that keeps its structure
- Preparing documents for field extraction
- Digitising paper forms
- Sizes
- about 1.4B
- Hardware
- from: 1 GPU
Documents and OCRNot maintained2024
Microsoft · USA
One model for every document task: reading, answering questions about a page, extracting fields, classification. In HR it is used to parse resumes and attached scans. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from a resume and its attachments
- Answering questions about document content
- Classifying incoming documents
- Sizes
- 742M
- Hardware
- from: Laptop
Documents and OCRNot maintained2022–2023
Microsoft · USA
Small models that find tables on PDF and scanned pages and restore their structure: rows, columns, headers. The text inside is read by a separate OCR.
- Finding tables in reports, statements and invoices
- Restoring rows and columns for export to Excel
- Preparing tabular data for analysis and RAG
- Sizes
- 29M
- Hardware
- from: Laptop
Documents and OCRNot maintained2023
Meta · USA
An early model that converts scientific PDFs into text with formulas. Now outdated and outperformed by almost all modern OCR models.
- Converting scientific papers from PDF into text with formulas
- Digitising technical documentation
- Sizes
- 250M – 350M
- Hardware
- from: Laptop
Image + textNot maintained2023
Google · USA
Reads a document or a screenshot as an image and answers with structure: text, fields, answers to questions. In HR it is fine-tuned for resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting data from resumes and forms supplied as images
- Questions about the content of a scan
- Parsing tables and diagrams in documents
- Sizes
- 282M – 1.3B
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
SCUT DLVC Lab, South China University of Technology · China
A light model that takes both the text and the position of blocks on the page into account: trained in one language and transferable to others. Good for tagging fields in resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Tagging fields in resumes and forms
- Extracting data from forms and templates
- Parsing documents in several languages
- Sizes
- about 130M for the English version and about 280M for the multilingual one
- Hardware
- from: Laptop
Documents and OCRNot maintained2021–2022
Microsoft · USA
Recognizes a single line of text, including handwriting. The official weights are English only, but the model is often fine-tuned for other languages; there are community Russian versions.
- Recognizing handwritten lines in questionnaires and forms
- Recognizing printed lines after text detection on the page
- A base for fine-tuning to your own handwriting or font
- Sizes
- 62M – 608M
- Hardware
- from: Laptop
Text analysisNot maintained2020–2022
Microsoft · USA
Classic document understanding models: they take into account the text, its position on the page and the image. They are fine-tuned to extract fields from forms and receipts. Only the first version is free for commercial use.
- Extracting fields from questionnaires, forms and receipts after fine-tuning
- Classifying document types
- Answering questions about a scanned page
- Sizes
- about 110M – 370M
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
NAVER CLOVA · South Korea
Reads a scanned document and returns a filled-in field structure straight away, with no separate OCR step. In HR it is fine-tuned for parsing resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from forms and resumes
- Parsing scans of certificates and diplomas
- Detecting the type of an incoming document
- Sizes
- about 200M
- Hardware
- from: Laptop
Computer visionNot maintained2022
Microsoft · USA
From an image it works out what kind of document it is: resume, diploma, certificate, contract. Helps sort candidate file bundles by type. A human makes the decision about a candidate; automatic screening without review must not be used.
- Sorting incoming candidate documents by type
- Checking that a document package is complete
- Finding the right scan in an archive
- Sizes
- base and large versions
- Hardware
- from: Laptop