CogVLM и GLM-V
Vision models from Zhipu: first CogVLM, then the GLM-V line. GLM-4.6V can call tools based on images and act as an agent operating an interface.
- Developer
- Zhipu AI (Z.ai) and Tsinghua University, China
- First release
- Nov 2023
- Latest release
- Dec 2025
- Sizes
- 9B – 106B-A12B
- License
- Commercial use allowedGLM-4.1V / 4.5V / 4.6V: MIT; CogVLM and CogVLM2: their own licenses (CogVLM2 is based on Llama 3)
- Russian
- Not supported
- Running
- On your own serverNeeds a GPU
- Industries
- Documents and accounting, Science and research
What it does
- Answering questions about photos and documents
- An agent that operates an interface from screenshots
- Analysing charts and reports
- Captions for photos and video
Where it is used
Hardware requirements
Versions
- GLM-4.6V и GLM-4.6V-Flash 9B
- GLM-4.5V 106B-A12B
- GLM-4.1V-9B-Thinking
- CogVLM2 19B
- CogAgent
- CogVLM-17B
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can CogVLM и GLM-V be used in a commercial project?
Yes. License: GLM-4.1V / 4.5V / 4.6V: MIT; CogVLM and CogVLM2: their own licenses (CogVLM2 is based on Llama 3). It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does CogVLM и GLM-V need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does CogVLM и GLM-V support Russian?
No. The model card lists its languages and Russian is not among them.
Where can I download CogVLM и GLM-V and what does it cost?
The CogVLM и GLM-V weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
DetailsImage + textQwen-VLAlibaba (Qwen team) · ChinaCommercial use allowedOne of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
DetailsDocuments and OCRGLM-OCRZhipu AI (Z.ai) · ChinaCommercial use allowedA lightweight OCR model from Zhipu for document parsing. The model card lists Russian among supported languages; built for high load and low-end hardware.
DetailsSource: huggingface.co/zai-org/GLM-4.6V. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


