MiniCPM-V or Qwen-VL: which to choose for image recognition
MiniCPM-V from OpenBMB is a compact line from 1.3B to 8B that runs on a phone or laptop, so receipts and documents can be processed on the device without sending anything to the cloud. Qwen-VL from Alibaba reaches 235B-A22B and covers heavier scenarios: camera footage analysis and an agent that operates an interface from screenshots. Both are in Ollama and both run without a GPU in their smaller sizes. MiniCPM-V is now Apache 2.0, though early versions had their own license requiring registration for commercial use; with Qwen-VL, Apache 2.0 applies only to Qwen3-VL, while older versions carry their own Qwen licenses. MiniCPM-V was updated later: version 4.6 shipped in 2026-05.
Comparison based on catalog data
| Parameter | MiniCPM-V | Qwen-VL |
|---|---|---|
| Category | Image + text, Documents and OCR | Image + text, Documents and OCR |
| Developer | OpenBMB (ModelBest and Tsinghua University), China | Alibaba (Qwen team), China |
| Releases | Jan 2024 – May 2026 | Aug 2023 – Oct 2025 |
| Sizes | 1.3B – 8B | 2B – 235B-A22B |
| Hardware | Laptop, 1 GPU | Laptop, 1 GPU, Cluster |
| Commercial use | Commercial use allowed | Commercial use allowed |
| License | Apache 2.0 (early versions had their own MiniCPM license requiring registration for commercial use) | Qwen3-VL: Apache 2.0; older Qwen-VL and some Qwen2/2.5-VL models (3B, 72B) have their own Qwen licenses |
| Russian | Not stated | Not stated |
| Ollama | Yes | Yes |
| Without GPU | Yes | Yes |
| Tasks |
|
|
Choose MiniCPM-V if
- The model has to run inside a mobile app or on an employee laptop
- Data cannot leave the device: receipts and documents are processed locally
- You need the smallest possible model: version 4.6 is just 1.3B
Choose Qwen-VL if
- You need headroom: Qwen-VL goes up to 235B-A22B and runs on a cluster
- Your task is analysing video and camera footage
- You need an agent that acts on interface screenshots
Other comparisons
- Qwen-VL or InternVL: which vision model to choose
- Qwen or GigaChat: which to choose for business
- Qwen or Llama: which to choose for business
- DeepSeek or Qwen: which to choose for business
- Gemma or Llama: which to choose for business
- Mistral or Llama: which to choose for business
- PaddleOCR-VL or DeepSeek-OCR: which to choose for business
- gpt-oss or Qwen: which one to run on your own server
- Llama or DeepSeek: which to run inside your own perimeter
- Mistral or Qwen: which to choose for business
- Gemma or Phi: small models for modest hardware
- YandexGPT or Qwen: which to choose for Russian
Need a model for your task?
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.
- SelectThe model and size for your task and hardware budget
- DeployOn your server or in a closed network, with an API
- Fine-tuneOn your data, or connect a knowledge base
- IntegrateInto your CRM, ERP, bot, website or team chat


