Vision-language models: image and text

These models look at an image and answer questions about it: describe a product photo, read a screenshot, chart or diagram, spot details in a picture. They power photo report checks, product listings and visual support. When choosing, check answer quality in your languages, the license and the GPU memory required.

72 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
MedicineGGUF2023–2026

HuatuoGPT

FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China

A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.

  • Draft discharge summaries for a doctor to review
  • Searching medical literature
  • Hints for doctors when reviewing images (Vision)
Sizes
7B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2026

DeepSeek

DeepSeek · China

DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.

  • Employee assistant on your own server
  • Analysis of long contracts and reports
  • Agents that work with tools and APIs
Sizes
7B – 1.6T-A49B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2023–2026

InternLM / Intern-S

Shanghai AI Laboratory · China

Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.

  • Research assistant: papers, formulas, data
  • Analysis of scientific and technical documents
  • Corporate chat on small models
Sizes
1.8B – about 1T
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

MiMo

Xiaomi · China

Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.

  • Logic and calculation tasks
  • Agents with tools
  • Help for developers
Sizes
7B – 1,02T-A42B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2024–2026

Aya

Cohere Labs · Canada

Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.

  • Translation and correspondence in less common languages
  • Multilingual chat assistant
  • Analysis of images with text (Vision)
Sizes
3.3B – 35B
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCR2026

jina-ocr-v1

Jina AI · Germany

Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.

  • Converting scans and PDFs to Markdown
  • Recognizing tables and formulas
  • Parsing invoices, acts and reports
Sizes
3.4B-A0.6B
Hardware
from: 1 GPU
Non-commercial onlyDetails
TextRUOllama2023–2026

Qwen

Alibaba · China

A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.

  • Chatbot and knowledge-base assistant
  • Replies to emails and customer requests
  • Document parsing and classification
Sizes
0,6B – 2,4T-A95B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2023–2026

GLM (ChatGLM)

Zhipu AI (Z.ai) · China

One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.

  • Corporate chat assistant
  • Agents for routine office tasks
  • Help for developers
Sizes
1.5B – 744B-A40B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2024–2026

Cohere Command

Cohere · Canada

Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.

  • Knowledge-base answers with source citations
  • Agents that work with internal systems
  • Translation and correspondence in different languages
Sizes
2.5B – 218B-A25B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2024–2026

NVIDIA Nemotron

NVIDIA · USA

NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).

  • Agents with tool calling
  • Reasoning and calculation tasks
  • Answers based on long documents
Sizes
4B – 550B-A55B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2025–2026

Liquid LFM

Liquid AI · USA

Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.

  • Offline assistant on a laptop or phone
  • Data extraction from documents
  • Tool calling in apps
Sizes
230M – 24B-A2B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextOllama2026

Muse Glimmer

Meta Superintelligence Labs · USA

An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.

  • Agents with tool calling
  • Analysis of screenshots, charts and documents
  • Multilingual assistant
Sizes
30B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agents2025–2026

OpenCUA / Qwen-CUA

XLANG Lab (University of Hong Kong) · China

Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.

  • Working in desktop software without an API
  • Moving data between systems
  • Running user scenarios for tests
Sizes
7B – about 400B (MoE)
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2025–2026

UI-Venus

Ant Group (inclusionAI) · China

An Ant Group family for finding elements on screen and completing tasks in phone and computer interfaces. UI-Venus-2 was specifically trained to refuse dangerous actions.

  • Automating actions in mobile apps
  • Filling in forms in web interfaces
  • UI autotests
Sizes
2B – 72B
Hardware
from: Laptop
Commercial use with conditionsDetails
Autonomous driving2025–2026

NVIDIA Alpamayo

NVIDIA · USA

Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.

  • Auto-labeling camera recordings to train your own driver assistance systems
  • Analyzing complex road scenes with text explanations
  • Testing autopilot systems in simulation on rare scenarios
Sizes
10B – 34B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Autonomous driving2026

Qwen-Drive

Alibaba (Qwen team) · China

An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.

  • A perception and planning prototype for autonomous vehicles on closed sites
  • Answering questions about camera recordings when reviewing incidents
  • Labeling road scenes to train your own models
Sizes
4B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextRUOllama2023–2026

Mistral

Mistral AI · France

European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).

  • Fast chat responses
  • Data extraction from text
  • Translation and multilingual work
Sizes
3B – 675B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

Kimi

Moonshot AI · China

Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.

  • Multi-step agents: search, data collection, reports
  • In-depth document analysis
  • Help for developers
Sizes
16B-A3B – 2.8T-A104B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextOllama2024–2026

EXAONE

LG AI Research · South Korea

Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.

  • Corporate assistant
  • Working with Korean and English texts
  • Analysis of documents and images (4.5)
Sizes
1.2B – 750B-A37B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

Apertus

Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland

Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.

  • Multilingual assistant
  • Answers based on documents
  • Analysis of images and scans (v1.5)
Sizes
0.5B – 70B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2026

Inkling

Thinking Machines Lab · USA

Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.

  • Flagship-level corporate assistant
  • Analysis of documents, images and audio
  • Programming help
Sizes
276B-A12B, 975B-A41B
Hardware
from: Cluster
Commercial use allowedDetails
Image + textGGUF2024–2026

Ovis

Alibaba (AIDC-AI) · China

Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.

  • Extracting data from invoices, contracts and delivery notes
  • Table recognition
  • Answering questions about photos and charts
Sizes
0.9B – 80B-A3B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agentsGGUF2025–2026

Fara

Microsoft · USA

Small Microsoft models for working in the browser: they look at the page and click, type and scroll. Designed to run directly on a work computer without the cloud.

  • Filling in web forms and applications
  • Collecting data from web portals without an API
  • Checking websites against scenarios
Sizes
4B – 27B
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2026

MOSS-VL

OpenMOSS (Fudan University) · China

An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.

  • Analyzing long videos and finding events by time
  • Real-time streaming video analysis
  • Understanding photos and documents
Sizes
about 11B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2024–2026

Gemma

Google · USA

Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.

  • Offline assistant on a laptop
  • Reading photos of documents and receipts
  • Customer request classification
Sizes
270M – 31B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUGGUF2025–2026

MiniMax

MiniMax · China

Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.

  • Analysis of large document archives in a single request
  • Agents with tools
  • Help for developers
Sizes
230B-A10B – 456B-A46B
Hardware
from: Cluster
Commercial use with conditionsDetails
Image + textOllama2024–2026

Moondream

Moondream (M87 Labs) · USA

A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.

  • Finding and counting objects in photos
  • Checking photos from field reports
  • Captions and tags for a catalogue
Sizes
2B – 9B-A2B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + text2023–2026

InternVideo

Shanghai AI Lab (OpenGVLab) · China

A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.

  • Searching a video archive with a text query
  • Action recognition in video
  • Answering questions about a long recording
Sizes
small encoders – 9B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

Step

StepFun · China

StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.

  • High-load agents
  • Analysis of documents with diagrams and screenshots
  • Help for developers
Sizes
8B – 321B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textOllama2023–2026

LLaVA

LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China

The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.

  • Answering questions about photos and screenshots
  • Describing products from a photo
  • Frame-by-frame video analysis
Sizes
0.5B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2024–2026

MiniCPM-V

OpenBMB (ModelBest and Tsinghua University) · China

Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.

  • On-device text recognition in photos
  • Processing receipts and documents without sending them to the cloud
  • Describing photos and video
Sizes
1.3B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textGGUF2025–2026

Keye-VL

Kuaishou · China

Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.

  • Analysing and describing short videos
  • Reviewing clips and content
  • Finding the right moment in a video
Sizes
8B – 671B-A37B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agentsGGUF2025–2026

Holo

H Company · France

A French model family for controlling a browser and computer: precisely finds the right element on screen and handles multi-step tasks. The latest Holo3 and 3.1 are open under Apache 2.0.

  • Working in web portals and legacy software without an API
  • Filling in forms and applications
  • Testing interfaces against scenarios
Sizes
0.8B – 235B-A22B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

ERNIE 4.5

Baidu · China

Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.

  • Corporate assistant
  • Analysis of documents and images
  • Customer request classification
Sizes
0.3B – 424B-A47B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2025–2026

Granite Vision

IBM · USA

Compact IBM models for business documents: tables, charts, forms, field-value pairs. The model card openly warns that it works best with English.

  • Extracting fields from forms and invoices
  • Turning charts and tables into data
  • Answering questions about documents
Sizes
2B – 4B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2024–2026

ShowUI

Show Lab (National University of Singapore) · Singapore

A lightweight model for working with interfaces: finds buttons and fields by description and performs actions on the web and on a phone. ShowUI-π can drag with the mouse.

  • Clicking and filling in forms from a task description
  • Web UI autotests
  • An assistant on a low-end computer without the cloud
Sizes
2B (ShowUI), about 500M (ShowUI-π)
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

Apriel

ServiceNow · USA

ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.

  • A reasoning assistant for internal services
  • Tool calling and enterprise agents
  • Analysing screenshots and documents with images
Sizes
5B – 15B
Hardware
from: Laptop
Commercial use allowedDetails
MedicineGGUF2025–2026

Hulu-Med

Zhejiang University · China

A medical model for text, images, 3D scans and video: from a light 4B to a large MoE. Does not replace a doctor; decisions are made by a specialist.

  • Hints for doctors when reviewing images and CT scans
  • Draft reports and discharge summaries
  • Searching medical literature
Sizes
4B – 235B-A22B
Hardware
from: Laptop
Commercial use allowedDetails
Text2025–2026

Reka Flash и Reka Edge

Reka AI · USA

Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.

  • Photo and video analysis (Edge)
  • Object detection in images
  • Reasoning tasks (Flash)
Sizes
7B – 21B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textGGUF2023–2026

InternVL

Shanghai AI Laboratory (OpenGVLab) · China

A large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.

  • Understanding documents, diagrams and charts
  • Answering questions about photos
  • Video analysis
Sizes
1B – 241B-A28B
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2024–2026

Molmo

Ai2 (Allen Institute for AI) · USA

Fully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.

  • Counting products and objects in photos
  • Pointing to where an item is in an image
  • Video analysis
Sizes
1B-A7B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistantsGGUF2025–2026

MiniCPM-o

OpenBMB (ModelBest, Tsinghua University) · China

A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.

  • Voice assistant on your own server
  • Analyzing videos and documents
  • Voice answers about a camera image
Sizes
8B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2025–2026

GUI-Owl / Mobile-Agent

Alibaba (Tongyi Lab, X-PLUG) · China

Models for controlling phones and computers from the Mobile-Agent project: they work with Android, Windows, macOS and the browser; version 1.5 has a reasoning mode.

  • Automating actions in mobile apps
  • Working in desktop software without an API
  • Testing apps against scenarios
Sizes
2B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2026

Qwen3-VL Resume Parser

Sukhrob Nurali · not disclosed

A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.

  • Moving a resume from PDF into a candidate record
  • Filling a candidate database without manual typing
  • Parsing resumes with different layouts and styling
Sizes
8B, a fine-tune of Qwen3-VL-8B-Instruct
Hardware
from: 1 GPU
Commercial use allowedDetails
MedicineOllama2025–2026

MedGemma

Google · USA

Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.

  • Draft discharge summaries and reports for a doctor to review
  • Hints for doctors when reviewing images
  • Searching and summarising medical literature
Sizes
4B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
Medicine2023–2026

CheXagent / CheXOne

Stanford AIMI · USA

Stanford models for chest X-rays: they describe the image and prepare a draft report. Does not replace a doctor; decisions are made by a specialist.

  • A draft X-ray description for the radiologist
  • Hints for doctors when reviewing images
  • Checking reports for completeness
Sizes
3B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
MedicineGGUF2025–2026

Lingshu

Alibaba DAMO Academy · China

Alibaba's medical model based on Qwen2.5-VL: understands many types of medical images and medical text, and can reason step by step. Does not replace a doctor; decisions are made by a specialist.

  • Hints for doctors when reviewing images
  • Draft reports and discharge summaries
  • Searching medical literature
Sizes
7B – 32B
Hardware
from: Laptop
Commercial use allowedDetails
TextRUOllama2023–2026

Phi

Microsoft · USA

Small Microsoft models trained on carefully selected data: strong at logic and math for their modest size. Versions with images and speech are available.

  • Assistant on a laptop or your own server
  • Reasoning and calculation tasks
  • Analysis of images and diagrams (vision versions)
Sizes
1.3B – 42B-A6.6B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agentsGGUF2026

EvoCUA

Meituan · China

Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.

  • Working in office and legacy software without an API
  • Moving data between systems
  • Running test scenarios
Sizes
8B – 32B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text2025

HyperCLOVA X SEED

Naver · South Korea

Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.

  • Lightweight Korean-English assistant
  • Analysis of images and documents
  • Text classification
Sizes
0.5B – 32B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + text2023–2025

CogVLM и GLM-V

Zhipu AI (Z.ai) and Tsinghua University · China

Vision models from Zhipu: first CogVLM, then the GLM-V line. GLM-4.6V can call tools based on images and act as an agent operating an interface.

  • Answering questions about photos and documents
  • An agent that operates an interface from screenshots
  • Analysing charts and reports
Sizes
9B – 106B-A12B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2023–2025

CogAgent / AutoGLM

Zhipu AI (Z.ai) and Tsinghua University · China

One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.

  • Automating actions in mobile apps
  • Working in web interfaces without an API
  • Testing apps against scenarios
Sizes
9B – 18B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer-use agentsGGUF2025

MAI-UI

Alibaba (Tongyi-MAI) · China

Compact Alibaba models for working in smartphone and computer interfaces: they find elements and complete multi-step tasks. The small size allows running on an ordinary GPU.

  • Automating actions in mobile apps
  • Working in software without an API
  • UI autotests
Sizes
2B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2023–2025

Qwen-VL

Alibaba (Qwen team) · China

One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.

  • Extracting data from scanned invoices and delivery notes
  • Analysing photos of products and shelves
  • Analysing video and camera footage
Sizes
2B – 235B-A22B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textRU2025

A-Vision (Авито)

Avito Tech · Russia

Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.

  • Product descriptions from photos in Russian
  • Checking that a photo matches its description
  • Reading brands and text in images
Sizes
7.4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice assistantsRUGGUF2025

Qwen Omni

Alibaba (Qwen) · China

Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.

  • Voice assistant for customers
  • Analyzing calls and videos
  • Voice answers about documents and images
Sizes
3B – 30B-A3B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textGGUF2025

Kimi-VL

Moonshot AI · China

An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.

  • Analysing long PDFs and presentations
  • Answering questions about video
  • Operating interfaces from screenshots
Sizes
16B-A3B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025

BAGEL

ByteDance Seed · China

A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.

  • Photo editing in a conversation
  • Answering questions about an image
  • Image generation with explanations
Sizes
14B-A7B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2023–2025

Llama

Meta · USA

The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.

  • Assistant for employees
  • Summaries of meetings and documents
  • Base for industry-specific fine-tuning
Sizes
1B – 405B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textGGUF2024–2025

SmolVLM

Hugging Face · France / USA

The smallest vision models from Hugging Face, starting at 256M; they run in a browser and on a phone. SmolVLM2 also understands video.

  • Describing photos and video on low-end hardware
  • Reading simple documents
  • Embedding in mobile and offline apps
Sizes
256M – 2.2B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agentsGGUF2025

UI-TARS

ByteDance Seed · China

A model that looks at a screenshot and controls the mouse and keyboard itself: clicks, fills in fields, navigates menus. The first generation and 1.5-7B are open; UI-TARS-2 weights were not released.

  • Working in legacy software without an API
  • Filling in forms and moving data between systems
  • UI autotests from plain-language scenarios
Sizes
2B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agentsGGUF2025

Magma

Microsoft Research · USA

An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.

  • Pilots in interface control
  • Research projects spanning screens and robotics
  • Analyzing screenshots with an action plan
Sizes
8B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textGGUF2024–2025

Janus

DeepSeek · China

A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.

  • Answering questions about images
  • Draft illustrations from a description
  • Experiments with a unified vision and generation model
Sizes
1B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + text2023–2025

VideoLLaMA

Alibaba DAMO Academy · China

Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.

  • Video description and short summary
  • Finding a moment in a recording by question
  • Tagging a video archive
Sizes
2B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textGGUF2024

PaliGemma

Google · USA

Google's vision model built on Gemma, designed as a base for fine-tuning on a narrow task: captions, object detection, reading text.

  • Fine-tuning for your own recognition task
  • Finding objects in photos
  • Reading text in images
Sizes
3B – 28B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textNot maintained2023–2024

IDEFICS

Hugging Face · France / USA

Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.

  • Answering questions about images
  • Analysing documents and screenshots
  • A base for fine-tuning
Sizes
8B – 80B
Hardware
from: 1 GPU
Commercial use allowedDetails
Documents and OCRNot maintained2024

Kosmos-2.5

Microsoft · USA

Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Converting resume scans into text that keeps its structure
  • Preparing documents for field extraction
  • Digitising paper forms
Sizes
about 1.4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image + textNot maintained2024

Florence-2

Microsoft · USA

A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.

  • Reading text in photos
  • Finding and highlighting objects
  • Automatic photo captions
Sizes
0.23B – 0.77B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2024

UDOP

Microsoft · USA

One model for every document task: reading, answering questions about a page, extracting fields, classification. In HR it is used to parse resumes and attached scans. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting fields from a resume and its attachments
  • Answering questions about document content
  • Classifying incoming documents
Sizes
742M
Hardware
from: Laptop
Commercial use allowedDetails
MedicineNot maintained2023

RadFM

Shanghai Jiao Tong University and Shanghai AI Lab · China

An early general-purpose radiology model: understands 2D and 3D images (CT, MRI) together with text. More of a research base. Does not replace a doctor; decisions are made by a specialist.

  • Research pilots on CT and MRI analysis
  • Hints for doctors when reviewing images
  • A base for fine-tuning on the clinic's own images
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image + textNot maintained2023

Pix2Struct

Google · USA

Reads a document or a screenshot as an image and answers with structure: text, fields, answers to questions. In HR it is fine-tuned for resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting data from resumes and forms supplied as images
  • Questions about the content of a scan
  • Parsing tables and diagrams in documents
Sizes
282M – 1.3B
Hardware
from: Laptop
Commercial use allowedDetails
Documents and OCRNot maintained2022

Donut

NAVER CLOVA · South Korea

Reads a scanned document and returns a filled-in field structure straight away, with no separate OCR step. In HR it is fine-tuned for parsing resumes and forms. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Extracting fields from forms and resumes
  • Parsing scans of certificates and diplomas
  • Detecting the type of an incoming document
Sizes
about 200M
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment