Image + text

Kimi-VL

An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.

Developer
Moonshot AI, China
First release
Apr 2025
Latest release
Jun 2025
Sizes
16B-A3B
License
Commercial use allowedMIT
Russian
Not stated
Ready-made builds
GGUF, MLX (Apple)
Running
On your own serverNeeds a GPU
Industries
Documents and accounting, Science and research

What it does

  • Analysing long PDFs and presentations
  • Answering questions about video
  • Operating interfaces from screenshots

Where it is used

Document workflowAnalyticsCustomer support

Hardware requirements

LaptopLaptop or regular PC, up to 8 GB of VRAM — smaller versions
no versions
1 GPUOne GPU with 16–80 GB — mid-size versions
fits
ClusterServer with several GPUs — flagship versions
no versions

Versions

  1. Kimi-VL-A3B-Thinking-2506
  2. Kimi-VL-A3B Instruct и Thinking

How to run it

On your own serverWeights are downloaded from Hugging Face and served with vLLM (load and API) or llama.cpp (modest hardware). The model then runs in a closed network with no per-request fees.Hugging Face
How much hardware you needCalculate the VRAM for the model size, context length and number of concurrent requests.Open the hardware calculator

I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems. Quantization compresses a model so it takes less video memory and runs on more modest hardware. Answers change slightly, so quality is checked on your own examples.

Frequently asked questions

Can Kimi-VL be used in a commercial project?

Yes. License: MIT. It allows commercial use, but it is still worth having a lawyer review the license before launch.

What hardware does Kimi-VL need?

At minimum: One GPU with 16–80 GB — mid-size versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.

Does Kimi-VL support Russian?

The model card does not list languages, so Russian support cannot be promised — it has to be tested on your own examples.

Where can I download Kimi-VL and what does it cost?

The Kimi-VL weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.

How I deploy it for clients

  1. SelectionI pick the model size for your task and hardware and test it on your examples.
  2. DeploymentI deploy it on your server or in a closed network and provide an API.
  3. Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
  4. IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.

Similar models

Source: huggingface.co/moonshotai/Kimi-VL-A3B-Thinking-2506. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.