Legal stack: contracts, precedent, redaction

Contracts are the case where the cost of a mistake is known in advance, so the chain is built so that every step can be checked. First the document has to be read reliably along with its layout, then similar clauses and your approved wording have to be found, then differences compared and explained, and personal data has to be dealt with before anything leaves the building. The model here assists a lawyer rather than replacing one.

Updated 22 Sep 2026Find a model in 4 questions

Step 1. Read the document with its layout

Contracts arrive as scans and PDFs where half the meaning sits in clause numbering, footnotes, annexes and tables. Document recognition models extract structure as well as text: clause, sub-clause, reference to an annex. Skip this step and the analysis drifts at the numbering level, so a reference to the penalty clause lands in the wrong section, which is unacceptable in legal work.

What does the work

Documents and OCRNot maintained2023

Nougat

Meta · USA

An early model that converts scientific PDFs into text with formulas. Now outdated and outperformed by almost all modern OCR models.

  • Converting scientific papers from PDF into text with formulas
  • Digitising technical documentation
Sizes
250M – 350M
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCRGGUF2025

olmOCR

Ai2 (Allen Institute for AI) · USA

A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.

  • Bulk digitisation of a PDF archive
  • Converting contracts and reports into text
  • Preparing documents for search and RAG
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCRGGUF2025–2026

MinerU

Shanghai AI Laboratory (OpenDataLab) · China

A popular open tool for converting PDFs to Markdown with its own small model. MinerU2.5-Pro was improved through data alone, without growing in size. Languages on the card: Chinese and English.

  • Converting PDF reports and contracts to Markdown
  • Recognising tables and formulas
  • Preparing documents for RAG and search
Sizes
0.9B – 1.2B
Hardware
from: Laptop
Commercial use allowedDetails

Step 2. Find precedent in your own archive

Meaning-based search surfaces approved wording, earlier versions and previous contracts with the same counterparty. Some models search the document page directly, layout included, which suits scans. Without this step a lawyer reads every contract from scratch, and the company renegotiates points it already settled two years ago.

What does the work

Visual document search2024–2025

ColPali / ColQwen

Illuin Technology (ViDoRe team) · France

Searches PDFs and scans as images: pages do not need to be OCR'd first, the model finds the right one for a question directly, including tables and charts. Trained on English.

  • Search across scans, presentations and PDFs
  • RAG over documents with tables and charts
  • Search across technical documentation
Sizes
256M – 3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Search and RAGRUOllamaNot maintained2024

BGE-M3

BAAI · China

A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.

  • Search across a document base
  • RAG for a chatbot
  • Finding similar requests and duplicates
Sizes
568M
Hardware
from: Laptop
Commercial use allowedDetails
RerankersGGUF2024–2026

Jina Reranker

Jina AI · Germany

Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.

  • Refining search results before a chatbot answers
  • Sorting retrieved PDF pages and slides
  • Catalog and knowledge base search
Sizes
33M – 2.4B
Hardware
from: Laptop
Commercial use with conditionsDetails

Step 3. Compare and explain the differences

A language model matches the incoming text against your approved version and lists the differences in plain words: payment terms, liability, termination, jurisdiction. It flags what looks off while the lawyer decides, so every remark must point at a specific clause. Without this part a first pass takes hours, and small edits by a counterparty get noticed after signing.

What does the work

TextNot maintained2024

SaulLM

Equall · France

Language models for legal texts, fine-tuned on US and European legal corpora (based on Mistral and Mixtral). English only.

  • Reviewing English-language contracts
  • Spotting risks and non-standard terms
  • Drafting legal memos
Sizes
7B – 141B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2023–2026

Qwen

Alibaba · China

A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.

  • Chatbot and knowledge-base assistant
  • Replies to emails and customer requests
  • Document parsing and classification
Sizes
0,6B – 2,4T-A95B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2026

DeepSeek

DeepSeek · China

DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.

  • Employee assistant on your own server
  • Analysis of long contracts and reports
  • Agents that work with tools and APIs
Sizes
7B – 1.6T-A49B
Hardware
from: Laptop
Commercial use allowedDetails

Step 4. Redact before sending

A dedicated model finds personal data and identifiers in the text and replaces them with placeholders before the document goes to an outside service or a shared archive. It works narrowly and fast and rewrites nothing else. Without this part, any external processing of contracts becomes a transfer of people data that you will have to account for, and that is no longer a technical question.

What does the work

Moderation and safety2024–2026

GLiNER-PII

GLiNER community (Fastino, Knowledgator, NVIDIA and others) · USA

Small GLiNER-based models for finding personal data: passports, phone numbers, accounts, addresses. Data types are set in words. Russian is not officially supported.

  • Masking personal data before cloud AI
  • Finding passport data and bank details in documents
  • Checking data exports for leaks
Sizes
about 200M to 500M
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safety2024

Piiranha

iiiorg · not disclosed

A popular detector of 17 types of personal data in six European languages. No Russian and a non-commercial license: suitable for trials and research.

  • Finding personal data in texts
  • Comparing the quality of PII detectors
Sizes
278M
Hardware
from: Laptop
Non-commercial onlyDetails
Moderation and safety2026

OpenAI Privacy Filter

OpenAI · USA

Finds and hides personal data: names, addresses, phone numbers, emails, account numbers, passwords. Runs even in the browser. Trained mostly on English.

  • Removing personal data from text before sending it to cloud AI
  • Finding passwords and keys in texts
  • Anonymizing correspondence for analytics
Sizes
1.5B (50M active)
Hardware
from: Laptop
Commercial use allowedDetails

What to check before you start

Here more than anywhere, remember that a model is wrong confidently: a missed clause looks exactly as calm as a caught one. So the chain is deployed as a first pass and a risk highlighter rather than as approval, and every remark is verified by a lawyer against the text. Keep a log of who accepted what, or nobody will be able to trace how a wording ended up in the signed version. Legal questions, personal data and permitted processing are described here in general terms only, with a lawyer reviewing the specific case.

Other stacks

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment