Open-source models for document recognition

OCR models turn scans and document photos into text and structured data: invoices, delivery notes, contracts, tables. This removes manual data entry in accounting and ERP systems. Check support for your scripts and handwriting, how well tables and layout are preserved, and the license terms.

36 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Documents and OCR2026

jina-ocr-v1

Jina AI · Germany

Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.

  • Converting scans and PDFs to Markdown
  • Recognizing tables and formulas
  • Parsing invoices, acts and reports
Sizes
3.4B-A0.6B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Documents and OCR2026

TeleOCR

TeleAI (China Telecom) · China

A new lightweight document parsing model that led the OmniDocBench v1.6 benchmark at release. Handles pages photographed on a phone and crumpled pages well. Languages on the card: Chinese, English, Japanese.

  • Recognising invoices and delivery notes photographed on a phone
  • Recognising tables and formulas
  • Converting documents to Markdown for RAG
Sizes
about 1.2B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textGGUF2024–2026

Ovis

Alibaba (AIDC-AI) · China

Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.

  • Extracting data from invoices, contracts and delivery notes
  • Table recognition
  • Answering questions about photos and charts
Sizes
0.9B – 80B-A3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRUGGUF2025–2026

HunyuanOCR

Tencent · China

A lightweight OCR model from Tencent: document parsing, finding text in photos, field extraction and translating text from images. Version 1.5 is faster and runs on an ordinary PC.

  • Extracting fields from invoices and delivery notes
  • Recognising tables and formulas
  • Translating text in photos and scans
Sizes
1B
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCRRU2024–2026

Surya OCR

Datalab · USA

A compact OCR toolkit from the makers of Marker and Chandra: text recognition, page layout, reading order and tables. Surya OCR 2 (650M) also runs on a CPU; Russian scored 88.8% in benchmarks.

  • Recognizing scans and PDFs, including in Russian
  • Page layout: headings, tables, images, reading order
  • Recognizing tables by rows and columns
Sizes
up to 650M
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCR2026

Unlimited-OCR

Baidu · China

Baidu's OCR model building on DeepSeek-OCR ideas: processes multi-page documents and PDFs in a single pass and outputs structured text. Claimed to be multilingual, but the language list is not published.

  • Converting multi-page PDFs and scans to text and Markdown
  • Recognizing contracts, invoices and reports
  • Preparing document archives for search and RAG
Sizes
3.3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRU2022–2026

PP-OCR и PP-DocLayout (классический PaddleOCR)

Baidu (PaddlePaddle) · China

Classic lightweight PaddleOCR models: detecting and recognizing lines of text plus page layout. They run on CPUs and phones; there is a separate model for East Slavic languages, including Russian.

  • Recognizing text on scans, photos and screens
  • Reading labels, displays and markings in production and warehouses
  • Page layout: tables, formulas, stamps, headings
Sizes
from 1.5M to tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2024–2026

MiniCPM-V

OpenBMB (ModelBest and Tsinghua University) · China

Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.

  • On-device text recognition in photos
  • Processing receipts and documents without sending them to the cloud
  • Describing photos and video
Sizes
1.3B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRUGGUF2025–2026

PaddleOCR-VL

Baidu (PaddlePaddle) · China

A compact document parsing model from the popular PaddleOCR toolkit. Per the model card it supports 109 languages, including Russian; version 1.6 leads the OmniDocBench benchmark.

  • Recognising invoices, contracts and delivery notes, including in Russian
  • Recognising tables, formulas and stamps
  • Converting scans to Markdown and JSON
Sizes
0.9B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRGGUF2025–2026

MinerU

Shanghai AI Laboratory (OpenDataLab) · China

A popular open tool for converting PDFs to Markdown with its own small model. MinerU2.5-Pro was improved through data alone, without growing in size. Languages on the card: Chinese and English.

  • Converting PDF reports and contracts to Markdown
  • Recognising tables and formulas
  • Preparing documents for RAG and search
Sizes
0.9B – 1.2B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2025–2026

Granite Vision

IBM · USA

Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.

  • Extracting fields from forms and invoices
  • Turning charts and tables into data
  • Answering questions about documents
Sizes
2B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisOllama2024–2026

NuExtract

NuMind · France

Models for template-based data extraction: give it a document or scan and a JSON field template, get a filled-in JSON back. NuExtract3 (4B) also converts scans to Markdown.

  • Extracting company details, amounts and dates from invoices and contracts into JSON
  • Parsing receipts, waybills and forms against a set template
  • Converting scans to Markdown for search
Sizes
0.5B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRUGGUF2025–2026

dots.ocr

rednote hilab (Xiaohongshu) · China

A multilingual document parsing model: text, tables, formulas and reading order in one pass. dots.mocr also turns charts and diagrams into vector SVG.

  • Recognising invoices, contracts and delivery notes
  • Converting tables into an editable format
  • Converting charts and diagrams into vector format
Sizes
about 3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRUGGUF2025–2026

Chandra OCR

Datalab · USA

A strong OCR model from the authors of Marker and Surya: handwriting, forms, tables. Per the model card it supports 90+ languages, with Russian among the examples.

  • Recognising invoices, contracts and delivery notes, including in Russian
  • Recognising handwritten forms and questionnaires
  • Recognising complex tables
Sizes
5B – 9B
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCRRUGGUF2026

Qianfan-OCR

Baidu (Qianfan) · China

A Baidu model that not only recognises a document but also answers questions about it. Per the model card it supports 192 languages, including Cyrillic.

  • Recognising invoices, contracts and delivery notes, including in Russian
  • Page layout analysis and table recognition
  • Answering questions about a document
Sizes
4B
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2026

Qwen3-VL Resume Parser

Sukhrob Nurali · not disclosed

A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.

  • Moving a resume from PDF into a candidate record
  • Filling a candidate database without manual typing
  • Parsing resumes with different layouts and styling
Sizes
8B, a fine-tune of Qwen3-VL-8B-Instruct
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCROllama2025–2026

DeepSeek-OCR

DeepSeek · China

An OCR model that compresses a page into a small number of visual tokens, so it processes large volumes quickly. Version 2 better understands reading order.

  • Bulk recognition of scanned invoices and contracts
  • Table recognition
  • Converting PDFs to Markdown for search and RAG
Sizes
about 3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRRUOllama2026

GLM-OCR

Zhipu AI (Z.ai) · China

A lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.

  • Recognising invoices, contracts and delivery notes, including in Russian
  • Recognising tables and formulas
  • Extracting fields to JSON
Sizes
0.9B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRGGUF2025–2026

LightOnOCR

LightOn · France

A French 1B OCR model that converts a page into text in one pass and is fast on high volumes. Languages on the card: European languages, Chinese and Japanese; no Russian.

  • Recognising invoices and contracts in European languages
  • Table recognition
  • Converting PDFs to text for search and RAG
Sizes
0.9B – 1B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2023–2025

Qwen-VL

Alibaba (Qwen team) · China

One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.

  • Extracting data from scanned invoices and delivery notes
  • Analysing photos of products and shelves
  • Analysing video and camera footage
Sizes
2B – 235B-A22B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRGGUF2025

olmOCR

Ai2 (Allen Institute for AI) · USA

A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.

  • Bulk digitisation of a PDF archive
  • Converting contracts and reports into text
  • Preparing documents for search and RAG
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCRRUGGUF2025

Nanonets-OCR

Nanonets · USA / India

A model that converts documents to Markdown with tables, stamps, signatures, checkboxes and watermarks. The OCR2 model card lists Russian among its languages.

  • Recognising invoices, contracts and delivery notes, including in Russian
  • Recognising stamps, signatures and marks
  • Handwriting recognition
Sizes
1.5B – 3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textRU2025

A-Vision (Авито)

Avito Tech · Russia

Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.

  • Product descriptions from photos in Russian
  • Checking that a photo matches its description
  • Reading brands and text in images
Sizes
7.4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCR2025

SmolDocling и Granite-Docling

IBM and Hugging Face · USA

Tiny models for the open Docling document converter: they turn a page into markup with tables, formulas and code. Run on an ordinary laptop.

  • Converting PDFs and scans to Markdown for search and RAG
  • Recognising tables in reports
  • Processing invoices and contracts on an ordinary PC
Sizes
256M – 258M
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2022–2025

TruFor

GRIP, University Federico II of Naples · Italy

Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.

  • Checking scans of certificates and contracts for edits
  • Highlighting altered photo regions for an expert
  • Filtering out obviously redrawn documents before manual review
Sizes
a transformer model producing a map of suspicious regions
Hardware
from: 1 GPU
Non-commercial onlyDetails
Documents and OCR2024

GOT-OCR 2.0

StepFun · China

One of the first general-purpose new-generation OCR models: text, formulas, tables, sheet music and diagrams. Small and runs on low-end hardware, but already behind newer models.

  • Recognising scanned invoices and contracts
  • Converting tables into an editable format
  • Recognising formulas and diagrams
Sizes
580M
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2024

Kosmos-2.5

Microsoft · USA

Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Converting resume scans into text that keeps its structure
  • Preparing documents for field extraction
  • Digitising paper forms
Sizes
about 1.4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCRNot maintained2024

UDOP

Microsoft · USA

One model for every document task: reading, answering questions about a page, extracting fields, classification. In HR it is used to parse resumes and attached scans. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting fields from a resume and its attachments
  • Answering questions about document content
  • Classifying incoming documents
Sizes
742M
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2022–2023

Table Transformer

Microsoft · USA

Small models that find tables on PDF and scanned pages and restore their structure: rows, columns, headers. The text inside is read by a separate OCR.

  • Finding tables in reports, statements and invoices
  • Restoring rows and columns for export to Excel
  • Preparing tabular data for analysis and RAG
Sizes
29M
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2023

Nougat

Meta · USA

An early model that converts scientific PDFs into text with formulas. Now outdated and outperformed by almost all modern OCR models.

  • Converting scientific papers from PDF into text with formulas
  • Digitising technical documentation
Sizes
250M – 350M
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textNot maintained2023

Pix2Struct

Google · USA

Reads a document or a screenshot as an image and answers with structure: text, fields, answers to questions. In HR it is fine-tuned for resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting data from resumes and forms supplied as images
  • Questions about the content of a scan
  • Parsing tables and diagrams in documents
Sizes
282M – 1.3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2022

LiLT

SCUT DLVC Lab, South China University of Technology · China

A light model that takes both the text and the position of blocks on the page into account: trained in one language and transferable to others. Good for tagging fields in resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Tagging fields in resumes and forms
  • Extracting data from forms and templates
  • Parsing documents in several languages
Sizes
about 130M for the English version and about 280M for the multilingual one
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2021–2022

TrOCR

Microsoft · USA

Recognizes a single line of text, including handwriting. The official weights are English only, but the model is often fine-tuned for other languages; there are community Russian versions.

  • Recognizing handwritten lines in questionnaires and forms
  • Recognizing printed lines after text detection on the page
  • A base for fine-tuning to your own handwriting or font
Sizes
62M – 608M
Hardware
from: Laptop
Commercial use allowedDetails
Text analysisNot maintained2020–2022

LayoutLM (v1–v3, LayoutXLM)

Microsoft · USA

Classic document understanding models: they take into account the text, its position on the page and the image. They are fine-tuned to extract fields from forms and receipts. Only the first version is free for commercial use.

  • Extracting fields from questionnaires, forms and receipts after fine-tuning
  • Classifying document types
  • Answering questions about a scanned page
Sizes
about 110M – 370M
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCRNot maintained2022

Donut

NAVER CLOVA · South Korea

Reads a scanned document and returns a filled-in field structure straight away, with no separate OCR step. In HR it is fine-tuned for parsing resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting fields from forms and resumes
  • Parsing scans of certificates and diplomas
  • Detecting the type of an incoming document
Sizes
about 200M
Hardware
from: Laptop
Commercial use allowedDetails
Computer visionNot maintained2022

DiT (Document Image Transformer)

Microsoft · USA

From an image it works out what kind of document it is: resume, diploma, certificate, contract. Helps sort candidate file bundles by type. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Sorting incoming candidate documents by type
  • Checking that a document package is complete
  • Finding the right scan in an archive
Sizes
base and large versions
Hardware
from: Laptop
Commercial use with conditionsDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment