VGGT
Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.
- Developer
- Meta and the University of Oxford (VGG), USA / UK
- First release
- Mar 2025
- Latest release
- May 2026
- Sizes
- about 1.2B
- License
- Commercial use with conditionsVGGT-1B — CC-BY-NC 4.0; VGGT-1B-Commercial — own license, commercial use allowed (except military use); VGGT-Omega — non-commercial research
- Running
- On your own serverNeeds a GPU
- Industries
- Manufacturing and logistics, Media and production, Science and research
What it does
- 3D model of a room or object from a photo series
- Camera pose estimation for photogrammetry
- Point cloud for measurements and comparison with the plan
- Groundwork for 3D scenes and Gaussian splats
Where it is used
Hardware requirements
Versions
- VGGT-Omega
- VGGT-1B-Commercial
- VGGT-1B
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can VGGT be used in a commercial project?
With conditions. License: VGGT-1B — CC-BY-NC 4.0; VGGT-1B-Commercial — own license, commercial use allowed (except military use); VGGT-Omega — non-commercial research. Restrictions vary — region, company revenue, attribution requirements. Have a lawyer check the terms before a commercial launch.
What hardware does VGGT need?
At minimum: One GPU with 16–80 GB — mid-size versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does VGGT support Russian?
Language does not matter for this model: it does not work with text.
Where can I download VGGT and what does it cost?
The VGGT weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.
Details3DPi3 (π³)Shanghai AI Lab · ChinaNon-commercial onlyReconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.
Details3DDUSt3R / MASt3RNAVER LABS Europe · France (NAVER, South Korea)Non-commercial onlyThe family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.
DetailsSource: github.com/facebookresearch/vggt. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


