MiniCPM-V
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- Developer
- OpenBMB (ModelBest and Tsinghua University), China
- First release
- Jan 2024
- Latest release
- May 2026
- Sizes
- 1.3B – 8B
- License
- Commercial use allowedApache 2.0 (early versions had their own MiniCPM license requiring registration for commercial use)
- Russian
- Not stated
- Ready-made builds
- GGUF, AWQ, GPTQ, MLX (Apple)
- Running
- Available in OllamaAlso runs without a GPU
- Industries
- Documents and accounting, Retail and marketplaces, Customer support
What it does
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- A built-in assistant for a mobile app
Where it is used
Hardware requirements
Versions
- MiniCPM-V 4.6 и 4.6-Thinking
- MiniCPM-V 4.5
- MiniCPM-V 4.0
- MiniCPM-V 2.6
- MiniCPM-Llama3-V 2.5
- MiniCPM-V 2.0
- MiniCPM-V
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems. Quantization compresses a model so it takes less video memory and runs on more modest hardware. Answers change slightly, so quality is checked on your own examples.
Frequently asked questions
Can MiniCPM-V be used in a commercial project?
Yes. License: Apache 2.0 (early versions had their own MiniCPM license requiring registration for commercial use). It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does MiniCPM-V need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Some versions also run on an ordinary CPU, without a GPU. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does MiniCPM-V support Russian?
The model card does not list languages, so Russian support cannot be promised — it has to be tested on your own examples.
Where can I download MiniCPM-V and what does it cost?
The MiniCPM-V weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Comparisons
Similar models
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
DetailsImage + textSmolVLMHugging Face · France / USACommercial use allowedThe smallest vision models from Hugging Face, starting at 256M; they run in a browser and on a phone. SmolVLM2 also understands video.
DetailsVoice assistantsMiniCPM-oOpenBMB (ModelBest, Tsinghua University) · ChinaCommercial use allowedA small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.
DetailsSource: huggingface.co/openbmb/MiniCPM-V-4.6. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


