Qwen-VL or InternVL: which vision model to choose
Both families read documents, answer questions about photos and analyse video, and both scale from a CPU to a cluster. Qwen-VL from Alibaba is available in Ollama, which makes deployment the simplest; the current Qwen3-VL is Apache 2.0, while older Qwen-VL and some Qwen2/2.5-VL models (3B, 72B) carry their own Qwen licenses, so check that before rollout. InternVL from Shanghai AI Laboratory was updated more recently (2026-03 vs 2025-10) and starts at 1B, and the 4B InternVL-U under MIT also generates and edits images. The catalog lists Qwen-VL for accounting, retail, e-commerce and support, and InternVL for document workflow, manufacturing quality control and e-commerce.
Comparison based on catalog data
| Parameter | Qwen-VL | InternVL |
|---|---|---|
| Category | Image + text, Documents and OCR | Image + text |
| Developer | Alibaba (Qwen team), China | Shanghai AI Laboratory (OpenGVLab), China |
| Releases | Aug 2023 – Oct 2025 | Dec 2023 – Mar 2026 |
| Sizes | 2B – 235B-A22B | 1B – 241B-A28B |
| Hardware | Laptop, 1 GPU, Cluster | Laptop, 1 GPU, Cluster |
| Commercial use | Commercial use allowed | Commercial use allowed |
| License | Qwen3-VL: Apache 2.0; older Qwen-VL and some Qwen2/2.5-VL models (3B, 72B) have their own Qwen licenses | InternVL3.5: Apache 2.0; InternVL-U: MIT; some large versions inherit the base model's license (for example, Qwen) |
| Russian | Not stated | Not stated |
| Ollama | Yes | No |
| Without GPU | Yes | Yes |
| Tasks |
|
|
Choose Qwen-VL if
- You want a simple install: Qwen-VL is in Ollama, InternVL is not
- You need an agent that operates an interface from screenshots
- You process scanned invoices and delivery notes, shelf photos and camera footage
Choose InternVL if
- You want one model for both understanding and image generation: InternVL-U (4B) under MIT
- You want the most recent release: InternVL was last updated 2026-03
- Your work is manufacturing quality control, diagrams and charts in documents
Other comparisons
- MiniCPM-V or Qwen-VL: which to choose for image recognition
- Qwen or GigaChat: which to choose for business
- Qwen or Llama: which to choose for business
- DeepSeek or Qwen: which to choose for business
- Gemma or Llama: which to choose for business
- Mistral or Llama: which to choose for business
- PaddleOCR-VL or DeepSeek-OCR: which to choose for business
- gpt-oss or Qwen: which one to run on your own server
- Llama or DeepSeek: which to run inside your own perimeter
- Mistral or Qwen: which to choose for business
- Gemma or Phi: small models for modest hardware
- YandexGPT or Qwen: which to choose for Russian
Need a model for your task?
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.
- SelectThe model and size for your task and hardware budget
- DeployOn your server or in a closed network, with an API
- Fine-tuneOn your data, or connect a knowledge base
- IntegrateInto your CRM, ERP, bot, website or team chat


