Keye-VL
Vision models from Kuaishou focused on short videos. Keye-VL-2.0 (30B, 3B active) understands well what happens in a clip and when.
- Developer
- Kuaishou, China
- First release
- Jun 2025
- Latest release
- May 2026
- Sizes
- 8B – 671B-A37B
- License
- Commercial use allowedApache 2.0
- Russian
- Not supported
- Ready-made builds
- GGUF
- Running
- On your own serverNeeds a GPU
- Industries
- Media and production, Marketing and content
What it does
- Analysing and describing short videos
- Reviewing clips and content
- Finding the right moment in a video
- Answering questions about photos
Where it is used
Hardware requirements
Versions
- Keye-VL-2.0-30B-A3B
- Keye-VL-671B-A37B
- Keye-VL-1.5-8B
- Keye-VL-8B-Preview
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems. Quantization compresses a model so it takes less video memory and runs on more modest hardware. Answers change slightly, so quality is checked on your own examples.
Frequently asked questions
Can Keye-VL be used in a commercial project?
Yes. License: Apache 2.0. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does Keye-VL need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does Keye-VL support Russian?
No. The model card lists its languages and Russian is not among them.
Where can I download Keye-VL and what does it cost?
The Keye-VL weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
DetailsImage + textKimi-VLMoonshot AI · ChinaCommercial use allowedAn efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.
DetailsImage + textInternVLShanghai AI Laboratory (OpenGVLab) · ChinaCommercial use allowedA large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.
DetailsSource: huggingface.co/Kwai-Keye/Keye-VL-2.0-30B-A3B. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


