MedicineGGUF2023–2026
FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China
A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Hints for doctors when reviewing images (Vision)
- Sizes
- 7B – 72B
- Hardware
- from: Laptop
TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextGGUF2025–2026
Xiaomi · China
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
- Logic and calculation tasks
- Agents with tools
- Help for developers
- Sizes
- 7B – 1,02T-A42B
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
Documents and OCR2026
Jina AI · Germany
Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.
- Converting scans and PDFs to Markdown
- Recognizing tables and formulas
- Parsing invoices, acts and reports
- Sizes
- 3.4B-A0.6B
- Hardware
- from: 1 GPU
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere · Canada
Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.
- Knowledge-base answers with source citations
- Agents that work with internal systems
- Translation and correspondence in different languages
- Sizes
- 2.5B – 218B-A25B
- Hardware
- from: Laptop
TextOllama2024–2026
NVIDIA · USA
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
- Agents with tool calling
- Reasoning and calculation tasks
- Answers based on long documents
- Sizes
- 4B – 550B-A55B
- Hardware
- from: Laptop
TextRUOllama2025–2026
Liquid AI · USA
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
- Offline assistant on a laptop or phone
- Data extraction from documents
- Tool calling in apps
- Sizes
- 230M – 24B-A2B
- Hardware
- from: Laptop
TextOllama2026
Meta Superintelligence Labs · USA
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
- Agents with tool calling
- Analysis of screenshots, charts and documents
- Multilingual assistant
- Sizes
- 30B
- Hardware
- from: 1 GPU
Computer-use agents2025–2026
XLANG Lab (University of Hong Kong) · China
Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.
- Working in desktop software without an API
- Moving data between systems
- Running user scenarios for tests
- Sizes
- 7B – about 400B (MoE)
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2025–2026
Ant Group (inclusionAI) · China
An Ant Group family for finding elements on screen and completing tasks in phone and computer interfaces. UI-Venus-2 was specifically trained to refuse dangerous actions.
- Automating actions in mobile apps
- Filling in forms in web interfaces
- UI autotests
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Autonomous driving2025–2026
NVIDIA · USA
Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.
- Auto-labeling camera recordings to train your own driver assistance systems
- Analyzing complex road scenes with text explanations
- Testing autopilot systems in simulation on rare scenarios
- Sizes
- 10B – 34B
- Hardware
- from: 1 GPU
Autonomous driving2026
Alibaba (Qwen team) · China
An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.
- A perception and planning prototype for autonomous vehicles on closed sites
- Answering questions about camera recordings when reviewing incidents
- Labeling road scenes to train your own models
- Sizes
- 4B
- Hardware
- from: 1 GPU
TextRUOllama2023–2026
Mistral AI · France
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
- Fast chat responses
- Data extraction from text
- Translation and multilingual work
- Sizes
- 3B – 675B
- Hardware
- from: Laptop
TextGGUF2025–2026
Moonshot AI · China
Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.
- Multi-step agents: search, data collection, reports
- In-depth document analysis
- Help for developers
- Sizes
- 16B-A3B – 2.8T-A104B
- Hardware
- from: 1 GPU
TextOllama2024–2026
LG AI Research · South Korea
Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.
- Corporate assistant
- Working with Korean and English texts
- Analysis of documents and images (4.5)
- Sizes
- 1.2B – 750B-A37B
- Hardware
- from: Laptop
TextGGUF2025–2026
Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland
Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.
- Multilingual assistant
- Answers based on documents
- Analysis of images and scans (v1.5)
- Sizes
- 0.5B – 70B
- Hardware
- from: Laptop
TextGGUF2026
Thinking Machines Lab · USA
Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.
- Flagship-level corporate assistant
- Analysis of documents, images and audio
- Programming help
- Sizes
- 276B-A12B, 975B-A41B
- Hardware
- from: Cluster
Image + textGGUF2024–2026
Alibaba (AIDC-AI) · China
Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.
- Extracting data from invoices, contracts and delivery notes
- Table recognition
- Answering questions about photos and charts
- Sizes
- 0.9B – 80B-A3B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
Microsoft · USA
Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.
- Filling in web forms and applications
- Collecting data from web portals without an API
- Checking websites against scenarios
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Image + text2026
OpenMOSS (Fudan University) · China
An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.
- Analyzing long videos and finding events by time
- Real-time streaming video analysis
- Understanding photos and documents
- Sizes
- about 11B
- Hardware
- from: 1 GPU
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
TextRUGGUF2025–2026
MiniMax · China
Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.
- Analysis of large document archives in a single request
- Agents with tools
- Help for developers
- Sizes
- 230B-A10B – 456B-A46B
- Hardware
- from: Cluster
Image + textOllama2024–2026
Moondream (M87 Labs) · USA
A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.
- Finding and counting objects in photos
- Checking photos from field reports
- Captions and tags for a catalogue
- Sizes
- 2B – 9B-A2B
- Hardware
- from: Laptop
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
TextGGUF2025–2026
StepFun · China
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
- High-load agents
- Analysis of documents with diagrams and screenshots
- Help for developers
- Sizes
- 8B – 321B
- Hardware
- from: 1 GPU
Image + textOllama2023–2026
LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China
The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.
- Answering questions about photos and screenshots
- Describing products from a photo
- Frame-by-frame video analysis
- Sizes
- 0.5B – 72B
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Image + textGGUF2025–2026
Kuaishou · China
Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.
- Analysing and describing short videos
- Reviewing clips and content
- Finding the right moment in a video
- Sizes
- 8B – 671B-A37B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025–2026
H Company · France
A French model family for controlling a browser and computer: precisely finds the right element on screen and handles multi-step tasks. The latest Holo3 and 3.1 are open under Apache 2.0.
- Working in web portals and legacy software without an API
- Filling in forms and applications
- Testing interfaces against scenarios
- Sizes
- 0.8B – 235B-A22B
- Hardware
- from: Laptop
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
Image + textOllama2025–2026
IBM · USA
Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.
- Extracting fields from forms and invoices
- Turning charts and tables into data
- Answering questions about documents
- Sizes
- 2B – 4B
- Hardware
- from: Laptop
Computer-use agents2024–2026
Show Lab (National University of Singapore) · Singapore
A lightweight model for working with interfaces: finds buttons and fields by description and performs actions on the web and on a phone. ShowUI-π can drag with the mouse.
- Clicking and filling in forms from a task description
- Web UI autotests
- An assistant on a low-end computer without the cloud
- Sizes
- 2B (ShowUI), about 500M (ShowUI-π)
- Hardware
- from: Laptop
TextGGUF2025–2026
ServiceNow · USA
ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.
- A reasoning assistant for internal services
- Tool calling and enterprise agents
- Analysing screenshots and documents with images
- Sizes
- 5B – 15B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Zhejiang University · China
A medical model for text, images, 3D scans and video: from a light 4B to a large MoE. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images and CT scans
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 4B – 235B-A22B
- Hardware
- from: Laptop
Text2025–2026
Reka AI · USA
Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.
- Photo and video analysis (Edge)
- Object detection in images
- Reasoning tasks (Flash)
- Sizes
- 7B – 21B
- Hardware
- from: Laptop
Image + textGGUF2023–2026
Shanghai AI Laboratory (OpenGVLab) · China
A large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.
- Understanding documents, diagrams and charts
- Answering questions about photos
- Video analysis
- Sizes
- 1B – 241B-A28B
- Hardware
- from: Laptop
Image + text2024–2026
Ai2 (Allen Institute for AI) · USA
Fully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.
- Counting products and objects in photos
- Pointing to where an item is in an image
- Video analysis
- Sizes
- 1B-A7B – 72B
- Hardware
- from: Laptop
Voice assistantsGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.
- Voice assistant on your own server
- Analyzing videos and documents
- Voice answers about a camera image
- Sizes
- 8B – 9B
- Hardware
- from: Laptop
Computer-use agents2025–2026
Alibaba (Tongyi Lab, X-PLUG) · China
Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.
- Automating actions in mobile apps
- Working in desktop software without an API
- Testing apps against scenarios
- Sizes
- 2B – 32B
- Hardware
- from: Laptop
Image + text2026
Sukhrob Nurali · not disclosed
A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.
- Moving a resume from PDF into a candidate record
- Filling a candidate database without manual typing
- Parsing resumes with different layouts and styling
- Sizes
- 8B, a fine-tune of Qwen3-VL-8B-Instruct
- Hardware
- from: 1 GPU
MedicineOllama2025–2026
Google · USA
Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.
- Draft discharge summaries and reports for a doctor to review
- Hints for doctors when reviewing images
- Searching and summarising medical literature
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Medicine2023–2026
Stanford AIMI · USA
Stanford models for chest X-rays: they describe the image and prepare a draft report. Does not replace a doctor; decisions are made by a specialist.
- A draft X-ray description for the radiologist
- Hints for doctors when reviewing images
- Checking reports for completeness
- Sizes
- 3B – 8B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Alibaba DAMO Academy · China
Alibaba's medical model based on Qwen2.5-VL: understands many types of medical images and medical text, and can reason step by step. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Microsoft · USA
Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.
- Assistant on a laptop or your own server
- Reasoning and calculation tasks
- Analysis of images and diagrams (vision versions)
- Sizes
- 1.3B – 42B-A6.6B
- Hardware
- from: Laptop
Computer-use agentsGGUF2026
Meituan · China
Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.
- Working in office and legacy software without an API
- Moving data between systems
- Running test scenarios
- Sizes
- 8B – 32B
- Hardware
- from: 1 GPU
Text2025
Naver · South Korea
Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.
- Lightweight Korean-English assistant
- Analysis of images and documents
- Text classification
- Sizes
- 0.5B – 32B
- Hardware
- from: Laptop
Image + text2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
Vision models from Zhipu: first CogVLM, then the GLM-V line. GLM-4.6V can call tools based on images and act as an agent operating an interface.
- Answering questions about photos and documents
- An agent that operates an interface from screenshots
- Analysing charts and reports
- Sizes
- 9B – 106B-A12B
- Hardware
- from: Laptop
Computer-use agents2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.
- Automating actions in mobile apps
- Working in web interfaces without an API
- Testing apps against scenarios
- Sizes
- 9B – 18B
- Hardware
- from: 1 GPU
Computer-use agentsGGUF2025
Alibaba (Tongyi-MAI) · China
Compact Alibaba models for working in smartphone and computer interfaces: they find elements and complete multi-step tasks. The small size allows running on an ordinary GPU.
- Automating actions in mobile apps
- Working in software without an API
- UI autotests
- Sizes
- 2B – 8B
- Hardware
- from: Laptop
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Image + textRU2025
Avito Tech · Russia
Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.
- Product descriptions from photos in Russian
- Checking that a photo matches its description
- Reading brands and text in images
- Sizes
- 7.4B
- Hardware
- from: 1 GPU
Voice assistantsRUGGUF2025
Alibaba (Qwen) · China
Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.
- Voice assistant for customers
- Analyzing calls and videos
- Voice answers about documents and images
- Sizes
- 3B – 30B-A3B
- Hardware
- from: Laptop
Image + textGGUF2025
Moonshot AI · China
An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.
- Analysing long PDFs and presentations
- Answering questions about video
- Operating interfaces from screenshots
- Sizes
- 16B-A3B
- Hardware
- from: 1 GPU
Image generationGGUF2025
ByteDance Seed · China
A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.
- Photo editing in a conversation
- Answering questions about an image
- Image generation with explanations
- Sizes
- 14B-A7B
- Hardware
- from: 1 GPU
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
Image + textGGUF2024–2025
Hugging Face · France / USA
The smallest vision models from Hugging Face, starting at 256M; they run in a browser and on a phone. SmolVLM2 also understands video.
- Describing photos and video on low-end hardware
- Reading simple documents
- Embedding in mobile and offline apps
- Sizes
- 256M – 2.2B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
ByteDance Seed · China
A model that looks at a screenshot and controls the mouse and keyboard itself: clicks, fills in fields, navigates menus. The first generation and 1.5-7B are open; UI-TARS-2 weights were not released.
- Working in legacy software without an API
- Filling in forms and moving data between systems
- UI autotests from plain-language scenarios
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
Microsoft Research · USA
An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.
- Pilots in interface control
- Research projects spanning screens and robotics
- Analyzing screenshots with an action plan
- Sizes
- 8B
- Hardware
- from: 1 GPU
Image + textGGUF2024–2025
DeepSeek · China
A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.
- Answering questions about images
- Draft illustrations from a description
- Experiments with a unified vision and generation model
- Sizes
- 1B – 7B
- Hardware
- from: Laptop
Image + text2023–2025
Alibaba DAMO Academy · China
Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
- Video description and short summary
- Finding a moment in a recording by question
- Tagging a video archive
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Image + textGGUF2024
Google · USA
Google's vision model built on Gemma, designed as a base for fine-tuning on a narrow task: captions, object detection, reading text.
- Fine-tuning for your own recognition task
- Finding objects in photos
- Reading text in images
- Sizes
- 3B – 28B
- Hardware
- from: Laptop
Image + textNot maintained2023–2024
Hugging Face · France / USA
Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.
- Answering questions about images
- Analysing documents and screenshots
- A base for fine-tuning
- Sizes
- 8B – 80B
- Hardware
- from: 1 GPU
Documents and OCRNot maintained2024
Microsoft · USA
Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.
- Converting resume scans into text that keeps its structure
- Preparing documents for field extraction
- Digitising paper forms
- Sizes
- about 1.4B
- Hardware
- from: 1 GPU
Image + textNot maintained2024
Microsoft · USA
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Sizes
- 0.23B – 0.77B
- Hardware
- from: Laptop
Documents and OCRNot maintained2024
Microsoft · USA
One model for every document task: reading, answering questions about a page, extracting fields, classification. In HR it is used to parse resumes and attached scans. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from a resume and its attachments
- Answering questions about document content
- Classifying incoming documents
- Sizes
- 742M
- Hardware
- from: Laptop
MedicineNot maintained2023
Shanghai Jiao Tong University and Shanghai AI Lab · China
An early general-purpose radiology model: understands 2D and 3D images (CT, MRI) together with text. More of a research base. Does not replace a doctor; decisions are made by a specialist.
- Research pilots on CT and MRI analysis
- Hints for doctors when reviewing images
- A base for fine-tuning on the clinic's own images
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Image + textNot maintained2023
Google · USA
Reads a document or a screenshot as an image and answers with structure: text, fields, answers to questions. In HR it is fine-tuned for resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting data from resumes and forms supplied as images
- Questions about the content of a scan
- Parsing tables and diagrams in documents
- Sizes
- 282M – 1.3B
- Hardware
- from: Laptop
Documents and OCRNot maintained2022
NAVER CLOVA · South Korea
Reads a scanned document and returns a filled-in field structure straight away, with no separate OCR step. In HR it is fine-tuned for parsing resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.
- Extracting fields from forms and resumes
- Parsing scans of certificates and diplomas
- Detecting the type of an incoming document
- Sizes
- about 200M
- Hardware
- from: Laptop