MolmoAct
A fully open robot control model that first "reasons" about space and trajectory, then acts. Its reasoning can be checked.
- Developer
- Ai2 (Allen Institute for AI), USA
- First release
- Aug 2025
- Latest release
- May 2026
- Sizes
- 5B – 8B
- License
- Commercial use allowedApache 2.0 (code and first version; the license for the MolmoAct2 base checkpoint is not stated on the model card)
- Running
- On your own serverNeeds a GPU
- Industries
- Manufacturing and logistics, Science and research
What it does
- Controlling a robot arm with explainable steps
- Fine-tuning for your own robot
- Research pilots
Where it is used
Hardware requirements
Versions
- MolmoAct 2
- MolmoAct
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can MolmoAct be used in a commercial project?
Yes. License: Apache 2.0 (code and first version; the license for the MolmoAct2 base checkpoint is not stated on the model card). It allows commercial use, but it is still worth having a lawyer review the license before launch.
What hardware does MolmoAct need?
At minimum: One GPU with 16–80 GB — mid-size versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does MolmoAct support Russian?
Language does not matter for this model: it does not work with text.
Where can I download MolmoAct and what does it cost?
The MolmoAct weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
The first large open vision-language-action model: a robot arm carries out commands like "put the apple in the bowl". OFT makes it several times faster.
DetailsRoboticsπ0 / π0.5 (openpi)Physical Intelligence · USACommercial use with conditionsRobot control models from Physical Intelligence: folding laundry, tidying up, handling objects. π0.5 copes better in unfamiliar settings.
DetailsImage + textMolmoAi2 (Allen Institute for AI) · USACommercial use allowedFully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.
DetailsSource: huggingface.co/allenai/MolmoAct2. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


