Sapiens
Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.
- Developer
- Meta, USA
- First release
- Aug 2024
- Latest release
- May 2026
- Sizes
- 0.1B – 5B
- License
- Commercial use with conditionsSapiens — CC-BY-NC 4.0; Sapiens2 — own Meta license, commercial use allowed with conditions
- Running
- On your own serverNeeds a GPU
- Industries
- Media and production, Retail and marketplaces, Healthcare
What it does
- Pose and body keypoint detection
- Segmentation of body parts and clothing
- Separating a person from the background
- Preparing data for try-on and avatars
Where it is used
Hardware requirements
Versions
- Sapiens2 Matting 1B
- Sapiens2 (0.1B – 5B)
- Sapiens (0.3B – 2B)
How to run it
I can set this up end to end: pick the model size, deploy it on your server and connect it to your systems.
Frequently asked questions
Can Sapiens be used in a commercial project?
With conditions. License: Sapiens — CC-BY-NC 4.0; Sapiens2 — own Meta license, commercial use allowed with conditions. Restrictions vary — region, company revenue, attribution requirements. Have a lawyer check the terms before a commercial launch.
What hardware does Sapiens need?
At minimum: Laptop or regular PC, up to 8 GB of VRAM — smaller versions. Without a GPU the model is not practical. You can calculate the exact VRAM for your model size and context in the hardware calculator.
Does Sapiens support Russian?
Language does not matter for this model: it does not work with text.
Where can I download Sapiens and what does it cost?
The Sapiens weights are open and free to download. You only pay for the hardware it runs on and for the setup. Source links are at the bottom of this page.
How I deploy it for clients
- SelectionI pick the model size for your task and hardware and test it on your examples.
- DeploymentI deploy it on your server or in a closed network and provide an API.
- Fine-tuningI fine-tune it on your data (LoRA) or connect a knowledge base — whichever is cheaper for the task.
- IntegrationI connect it to your CRM, ERP, bot, website or team chat and set up monitoring.
Similar models
A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.
DetailsComputer visionSegment Anything (SAM)Meta · USACommercial use with conditionsSelects any object in photos and videos with a click or a box. The basis for background removal and object counting.
DetailsComputer visionDepth AnythingByteDance and the University of Hong Kong · ChinaCommercial use with conditionsEstimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.
DetailsSource: github.com/facebookresearch/sapiens2. Data checked against the model card on 22 Sep 2026. Have a lawyer review the license before commercial launch.


