MOSS-VL
An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.
- Developer
- OpenMOSS (Fudan University), China
- First release
- Apr 2026
- Latest release
- Jul 2026
- Sizes
- about 11B
- License
- Commercial use allowedApache 2.0
- Russian
- Not supported
- Running
- On your own serverNeeds a GPU
- Industries
- Media and production, Security, Documents and accounting
What it does
- Analyzing long videos and finding events by time
- Real-time streaming video analysis
- Understanding photos and documents
Where it is used
Hardware requirements
Versions
- MOSS-VL-Realtime
- MOSS-VL 0708 (Base, Instruct)
- MOSS-VL 0408 (Base, Instruct)
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can MOSS-VL be used in a commercial project?
Yes. License: Apache 2.0. It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does MOSS-VL need?
At minimum: One GPU with 16–80 GB — mid-size versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does MOSS-VL support Russian?
No. The model card lists its languages and Russian is not among them.
Where can I download MOSS-VL and what does it cost?
The MOSS-VL weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
DetailsImage + textVideoLLaMAAlibaba DAMO Academy · ChinaCommercial use allowedModels that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
DetailsImage + textQwen-VLAlibaba (Qwen team) · ChinaCommercial use allowedOne of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
DetailsSource: huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


