Open-source computer vision models

Computer vision models detect, segment and count objects in photos and video: people in a store, products on a shelf, defects on a line. Many are lightweight and run on a camera or a small server. Look at speed, whether you can fine-tune on your own data, and the license terms.

43 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Faces2021–2026

InsightFace

InsightFace (deepinsight) · China

The most widely used open toolkit for face detection and recognition. Many identity-preserving image generators are built on it. The pretrained weights are non-commercial.

  • Detecting and comparing faces in photos
  • Face-based access in prototypes
  • Face processing as part of other AI systems
Sizes
packages from 16 MB to 407 MB
Hardware
from: Laptop
Non-commercial onlyDetails
3D2024–2026

DUSt3R / MASt3R

NAVER LABS Europe · France (NAVER, South Korea)

The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.

  • 3D scene from several photos without calibration
  • Point matching between images
  • Mapping from video (SLAM)
Sizes
0.57B – 0.69B
Hardware
from: Laptop
Non-commercial onlyDetails
Autonomous driving2020–2026

comma.ai openpilot (supercombo)

comma.ai · USA

An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.

  • A research testbed for driver assistance systems
  • Studying driver attention monitoring with an in-cabin camera
  • Comparison with your own lane-keeping algorithms
Sizes
compact, designed for an in-vehicle device
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2023–2026

TrustMark

Adobe Research and University of Surrey · USA

An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.

  • Marking images on the way out of your own pipeline
  • Checking the provenance of a submitted image
  • Linking with content provenance metadata
Sizes
model types Q and P with different mark capacity
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2024–2026

MoGe

Microsoft Research · USA

Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.

  • Measuring rooms and objects from photos
  • 3D point cloud from a single shot
  • Preparing data for robots and AR
Sizes
ViT-S – ViT-G
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2025–2026

Community Forensics

University of Michigan · USA

A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.

  • Checking submitted photos and illustrations
  • Filtering AI images in a content flow
  • Flagging suspicious images for manual review
Sizes
22M
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2023–2026

UniversalFakeDetect

University of Wisconsin-Madison · USA

An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.

  • Checking images from new, unfamiliar generators
  • A baseline when comparing detectors
  • Fast rollout of a check without training a large model
Sizes
a linear classifier on top of CLIP ViT-L/14
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2025–2026

RF-DETR

Roboflow · USA

Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.

  • Object detection in video and photos
  • Precise outlines of parts and defects
  • Fine-tuning for your own object classes
Sizes
Nano – 2XL
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2023–2026

InternVideo

Shanghai AI Lab (OpenGVLab) · China

A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.

  • Searching a video archive with a text query
  • Action recognition in video
  • Answering questions about a long recording
Sizes
small encoders – 9B
Hardware
from: Laptop
Commercial use allowedDetails
3D2025–2026

VGGT

Meta and the University of Oxford (VGG) · USA / UK

Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.

  • 3D model of a room or object from a photo series
  • Camera pose estimation for photogrammetry
  • Point cloud for measurements and comparison with the plan
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2024–2026

Sapiens

Meta · USA

Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.

  • Pose and body keypoint detection
  • Segmentation of body parts and clothing
  • Separating a person from the background
Sizes
0.1B – 5B
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer vision2023–2026

Segment Anything (SAM)

Meta · USA

Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.

  • Background removal from product photos
  • Counting objects in photos
  • Data labeling for training
Sizes
91M – ~0,85B
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer vision2023–2026

YOLO (Ultralytics)

Ultralytics · USA

The most widely used real-time object detector: finds and marks items in video even on modest hardware. YOLOv5 came out back in 2020; the catalog starts from YOLOv8.

  • Counting people, cars and goods on video
  • Checking hard hats and workwear
  • Spotting defects on the production line
Sizes
2.4M – 68M
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer visionGGUF2024–2025

Depth Anything

ByteDance and the University of Hong Kong · China

Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.

  • Estimating distances and volumes from a camera
  • Depth effects for photo and video
  • Navigation for robots and drones
Sizes
25M – 1.4B
Hardware
from: Laptop
Commercial use with conditionsDetails
3D2025

MapAnything

Meta and Carnegie Mellon University · USA

A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.

  • 3D reconstruction of an object or room from photos
  • Exporting the scene to COLMAP format for further processing
  • Depth and camera pose estimation
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2025

Pi3 (π³)

Shanghai AI Lab · China

Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.

  • 3D scene reconstruction from video
  • Camera pose estimation from frames
  • Point clouds for research and prototypes
Sizes
0.96B – 1.4B
Hardware
from: 1 GPU
Non-commercial onlyDetails
3D2025

HY-Motion

Tencent Hunyuan · China

Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.

  • Character animation from a text description
  • Draft animation for games and videos
  • Motion library for avatars
Sizes
0.46B – 1B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2025

Perception Encoder (PE)

Meta · USA

Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.

  • Search photos and videos by description
  • Catalog labeling and tagging
  • Search across audio and video (PE-AV)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
3DGGUF2025

SAM 3D

Meta · USA

Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.

  • 3D model of an item from a catalog photo
  • Estimating body pose and shape from a photo
  • Try-on and AR scenarios
Sizes
size not stated on the model card
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2023–2025

MetaCLIP / MetaCLIP 2

Meta · USA

Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.

  • Image search by text
  • Image classification without training
  • Search research and prototypes
Sizes
0.15B – 3.6B
Hardware
from: Laptop
Non-commercial onlyDetails
Computer vision2023–2025

Grounding DINO / Rex-Omni

IDEA Research · China

Finds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.

  • Finding objects by description without labeling
  • Automatic data labeling for training
  • Checking photos against requirements
Sizes
172M – 3B
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer vision2023–2025

DINOv2 / DINOv3

Meta · USA

Foundation models that turn an image into a numeric "fingerprint". They are used to build similar-image search, classification and segmentation without large labeled datasets.

  • Finding similar products and photos
  • Image classification on small datasets
  • Base for your own quality-control models
Sizes
21M – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Faces2025

LVFace

ByteDance · China

Transformer-based face recognition from ByteDance, one of the most accurate open models on benchmarks. Weights are published in ONNX format but are non-commercial.

  • Comparing faces and searching a photo database
  • Research on recognition accuracy
  • Access control prototypes
Sizes
ViT-T – ViT-L
Hardware
from: Laptop
Non-commercial onlyDetails
Deepfake detection2023–2025

DeepfakeBench

The Chinese University of Hong Kong, Shenzhen (SCLBD) · China

Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.

  • First-pass check of a submitted video or selfie
  • Comparing several detectors on your own data
  • Fine-tuning a detector for your own flow of applications
Sizes
Xception- and EfficientNet-class detectors, tens of millions of parameters
Hardware
from: Laptop
Non-commercial onlyDetails
Medicine2025

MedSigLIP

Google · USA

Google's lightweight encoder for medical images and text, the same one inside MedGemma. Sorts images and finds similar ones. Does not replace a doctor; decisions are made by a specialist.

  • Finding similar images in a clinic's archive
  • Pre-sorting images for a doctor
  • A base for your own image classifiers
Sizes
about 0.9B
Hardware
from: Laptop
Commercial use with conditionsDetails
Photo editingGGUF2024–2025

BiRefNet

Nankai University · China

An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.

  • Bulk background removal from product photos
  • Precise masks for design and print
  • Cutting out people with hair for advertising
Sizes
about 220M (lightweight lite versions available)
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2024–2025

Watermark Anything (WAM)

Meta · USA

An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.

  • Marking generated and edited images
  • Finding a marked fragment inside a collage
  • Tracking which parts of a picture were made by AI
Sizes
a mark encoder and decoder for images
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detection2024–2025

AIDE

Xiaohongshu, USTC and Shanghai Jiao Tong University · China

An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.

  • Checking realistic AI images without obvious artifacts
  • Comparing detectors on hard examples
  • Fine-tuning for your own type of content
Sizes
several experts based on ConvNeXt and CLIP
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Deepfake detection2022–2025

TruFor

GRIP, University Federico II of Naples · Italy

Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.

  • Checking scans of certificates and contracts for edits
  • Highlighting altered photo regions for an expert
  • Filtering out obviously redrawn documents before manual review
Sizes
a transformer model producing a map of suspicious regions
Hardware
from: 1 GPU
Non-commercial onlyDetails
Moderation and safety2023–2025

NSFW-классификаторы (Falconsai, Freepik)

Falconsai and Freepik · USA and Spain

Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.

  • Filtering user photos and avatars
  • Checking generated images before publishing
  • Labeling a media library
Sizes
86M
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2025

SigLIP (наследник CLIP)

Google · USA

Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.

  • Image search by text query
  • Automatic catalog labeling and tagging
  • Filtering prohibited content
Sizes
about 0.2B to 2B
Hardware
from: Laptop
Commercial use allowedDetails
Computer-use agents2024–2025

OmniParser

Microsoft · USA

Breaks a screenshot down into buttons, fields and icons with labels so a regular language model can understand and control the screen. It does not click itself; it serves as the agent's eyes.

  • Mapping legacy software screens for automation
  • Preparing an agent to work in an interface
  • Checking that the required elements are on screen
Sizes
under 1B (detector + captioning)
Hardware
from: Laptop
Commercial use with conditionsDetails
Medicine2024–2025

Bioptimus H-optimus

Bioptimus · France

Pathology foundation models from France's Bioptimus with 1.1B parameters, plus the compact H0-mini. The first version is open under Apache 2.0. Does not replace a doctor; decisions are made by a specialist.

  • Histology slide patch features for research models
  • A prototype for sorting slides by tissue type
  • Research projects linking morphology and molecular data
Sizes
86M (H0-mini) – 1,1B
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer vision2022–2025

ViTPose

University of Sydney and JD Explore Academy · Australia / China

A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.

  • Body keypoints in photos and video
  • Motion analysis in sports and rehabilitation
  • Monitoring work postures and safety practices
Sizes
33M – about 1B
Hardware
from: Laptop
Commercial use allowedDetails
Medicine2024–2025

MahmoodLab UNI / UNI 2

Mahmood Lab (Mass General Brigham, Harvard) · USA

A foundation model for histology slides: turns patches of digital slides into features for tissue classification. Does not replace a doctor; decisions are made by a specialist.

  • Research classifiers of tissue types from digital slides
  • Finding similar cases in a slide archive
  • Preparing features for research prognosis models
Sizes
about 300M (UNI) – about 680M (UNI2-h)
Hardware
from: Laptop
Non-commercial onlyDetails
Medicine2024

MahmoodLab CONCH / TITAN

Mahmood Lab (Mass General Brigham, Harvard) · USA

Image-plus-text models for pathology: search slides by an English description, classify without fine-tuning; TITAN describes a whole slide. Does not replace a doctor; decisions are made by a specialist.

  • Text-query search across a slide archive for research
  • Draft slide descriptions for research projects
  • Tissue classification without labels at the start of a study
Sizes
about 160M – 300M
Hardware
from: Laptop
Non-commercial onlyDetails
MedicineNot maintained2024

Paige Virchow

Paige · USA

Paige's pathology foundation model, trained on millions of digital slides. The first version is Apache 2.0, the second is for research only. Does not replace a doctor; decisions are made by a specialist.

  • Slide patch features for research classifiers
  • Selecting slides for re-review in research projects
  • Comparison with other pathology models on your own archive
Sizes
about 632M
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textNot maintained2024

Florence-2

Microsoft · USA

A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.

  • Reading text in photos
  • Finding and highlighting objects
  • Automatic photo captions
Sizes
0.23B – 0.77B
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detectionNot maintained2023–2024

IML-ViT

Sichuan University and co-authors · China

An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.

  • Finding pasted and erased fragments in photos
  • Checking document scans for edits
  • A baseline when comparing manipulation-localization models
Sizes
a Vision Transformer based model
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer visionNot maintained2021–2022

CLIP (OpenAI)

OpenAI · USA

The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.

  • Image search by text query
  • Automatic tags for a catalog
  • Finding similar images
Sizes
about 0.15B – 0.6B
Hardware
from: Laptop
Commercial use allowedDetails
Computer visionNot maintained2022

DiT (Document Image Transformer)

Microsoft · USA

From an image it works out what kind of document it is: resume, diploma, certificate, contract. Helps sort candidate file bundles by type. A human makes the decision about a candidate; automatic screening without review must not be used.

  • Sorting incoming candidate documents by type
  • Checking that a document package is complete
  • Finding the right scan in an archive
Sizes
base and large versions
Hardware
from: Laptop
Commercial use with conditionsDetails
Computer visionRUNot maintained2022

ruCLIP (ai-forever)

Sber AI and SberDevices (ai-forever) · Russia

A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.

  • Product search by photo and by Russian description
  • Sorting images into categories without labeling
  • Checking that a photo matches its caption
Sizes
150M – 430M
Hardware
from: Laptop
Commercial use allowedDetails
Deepfake detectionNot maintained2020

Silent-Face-Anti-Spoofing (MiniFASNet)

MiniVision Technology · China

Practically the only fully open weight set for single-frame face liveness: it tells a live person from a photo, a screen or a mask. It errs in both directions - a person must be able to appeal a rejection.

  • Liveness check when signing in by selfie
  • Protecting an access system from a photo on a phone
  • A check during remote customer identification
Sizes
two models of about 1.8 MB each
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment