Faces2021–2026
InsightFace (deepinsight) · China
The most widely used open toolkit for face detection and recognition. Many identity-preserving image generators are built on it. The pretrained weights are non-commercial.
- Detecting and comparing faces in photos
- Face-based access in prototypes
- Face processing as part of other AI systems
- Sizes
- packages from 16 MB to 407 MB
- Hardware
- from: Laptop
3D2024–2026
NAVER LABS Europe · France (NAVER, South Korea)
The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.
- 3D scene from several photos without calibration
- Point matching between images
- Mapping from video (SLAM)
- Sizes
- 0.57B – 0.69B
- Hardware
- from: Laptop
Autonomous driving2020–2026
comma.ai · USA
An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.
- A research testbed for driver assistance systems
- Studying driver attention monitoring with an in-cabin camera
- Comparison with your own lane-keeping algorithms
- Sizes
- compact, designed for an in-vehicle device
- Hardware
- from: Laptop
Deepfake detection2023–2026
Adobe Research and University of Surrey · USA
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
- Marking images on the way out of your own pipeline
- Checking the provenance of a submitted image
- Linking with content provenance metadata
- Sizes
- model types Q and P with different mark capacity
- Hardware
- from: Laptop
Computer vision2024–2026
Microsoft Research · USA
Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.
- Measuring rooms and objects from photos
- 3D point cloud from a single shot
- Preparing data for robots and AR
- Sizes
- ViT-S – ViT-G
- Hardware
- from: Laptop
Deepfake detection2025–2026
University of Michigan · USA
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
- Checking submitted photos and illustrations
- Filtering AI images in a content flow
- Flagging suspicious images for manual review
- Sizes
- 22M
- Hardware
- from: Laptop
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
Computer vision2025–2026
Roboflow · USA
Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.
- Object detection in video and photos
- Precise outlines of parts and defects
- Fine-tuning for your own object classes
- Sizes
- Nano – 2XL
- Hardware
- from: Laptop
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
3D2025–2026
Meta and the University of Oxford (VGG) · USA / UK
Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.
- 3D model of a room or object from a photo series
- Camera pose estimation for photogrammetry
- Point cloud for measurements and comparison with the plan
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
Computer vision2024–2026
Meta · USA
Meta's models for analyzing people in photos: pose keypoints, body part segmentation, normals and depth. Sapiens2 was trained at high resolution and adds human matting.
- Pose and body keypoint detection
- Segmentation of body parts and clothing
- Separating a person from the background
- Sizes
- 0.1B – 5B
- Hardware
- from: Laptop
Computer vision2023–2026
Meta · USA
Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.
- Background removal from product photos
- Counting objects in photos
- Data labeling for training
- Sizes
- 91M – ~0,85B
- Hardware
- from: Laptop
Computer vision2023–2026
Ultralytics · USA
The most widely used real-time object detector: finds and marks items in video even on modest hardware. YOLOv5 came out back in 2020; the catalog starts from YOLOv8.
- Counting people, cars and goods on video
- Checking hard hats and workwear
- Spotting defects on the production line
- Sizes
- 2.4M – 68M
- Hardware
- from: Laptop
Computer visionGGUF2024–2025
ByteDance and the University of Hong Kong · China
Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.
- Estimating distances and volumes from a camera
- Depth effects for photo and video
- Navigation for robots and drones
- Sizes
- 25M – 1.4B
- Hardware
- from: Laptop
3D2025
Meta and Carnegie Mellon University · USA
A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.
- 3D reconstruction of an object or room from photos
- Exporting the scene to COLMAP format for further processing
- Depth and camera pose estimation
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
3D2025
Shanghai AI Lab · China
Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.
- 3D scene reconstruction from video
- Camera pose estimation from frames
- Point clouds for research and prototypes
- Sizes
- 0.96B – 1.4B
- Hardware
- from: 1 GPU
3D2025
Tencent Hunyuan · China
Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.
- Character animation from a text description
- Draft animation for games and videos
- Motion library for avatars
- Sizes
- 0.46B – 1B
- Hardware
- from: 1 GPU
Computer vision2025
Meta · USA
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
- Search photos and videos by description
- Catalog labeling and tagging
- Search across audio and video (PE-AV)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
3DGGUF2025
Meta · USA
Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.
- 3D model of an item from a catalog photo
- Estimating body pose and shape from a photo
- Try-on and AR scenarios
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Computer vision2023–2025
Meta · USA
Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.
- Image search by text
- Image classification without training
- Search research and prototypes
- Sizes
- 0.15B – 3.6B
- Hardware
- from: Laptop
Computer vision2023–2025
IDEA Research · China
Finds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.
- Finding objects by description without labeling
- Automatic data labeling for training
- Checking photos against requirements
- Sizes
- 172M – 3B
- Hardware
- from: Laptop
Computer vision2023–2025
Meta · USA
Foundation models that turn an image into a numeric "fingerprint". They are used to build similar-image search, classification and segmentation without large labeled datasets.
- Finding similar products and photos
- Image classification on small datasets
- Base for your own quality-control models
- Sizes
- 21M – 7B
- Hardware
- from: Laptop
Faces2025
ByteDance · China
Transformer-based face recognition from ByteDance, one of the most accurate open models on benchmarks. Weights are published in ONNX format but are non-commercial.
- Comparing faces and searching a photo database
- Research on recognition accuracy
- Access control prototypes
- Sizes
- ViT-T – ViT-L
- Hardware
- from: Laptop
Deepfake detection2023–2025
The Chinese University of Hong Kong, Shenzhen (SCLBD) · China
Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.
- First-pass check of a submitted video or selfie
- Comparing several detectors on your own data
- Fine-tuning a detector for your own flow of applications
- Sizes
- Xception- and EfficientNet-class detectors, tens of millions of parameters
- Hardware
- from: Laptop
Medicine2025
Google · USA
Google's lightweight encoder for medical images and text, the same one inside MedGemma. Sorts images and finds similar ones. Does not replace a doctor; decisions are made by a specialist.
- Finding similar images in a clinic's archive
- Pre-sorting images for a doctor
- A base for your own image classifiers
- Sizes
- about 0.9B
- Hardware
- from: Laptop
Photo editingGGUF2024–2025
Nankai University · China
An open MIT-licensed model for precise object segmentation and background removal. RMBG-2.0 is built on it. Versions for 2K and for hair and semi-transparent edges.
- Bulk background removal from product photos
- Precise masks for design and print
- Cutting out people with hair for advertising
- Sizes
- about 220M (lightweight lite versions available)
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
- Marking generated and edited images
- Finding a marked fragment inside a collage
- Tracking which parts of a picture were made by AI
- Sizes
- a mark encoder and decoder for images
- Hardware
- from: Laptop
Deepfake detection2024–2025
Xiaohongshu, USTC and Shanghai Jiao Tong University · China
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
- Checking realistic AI images without obvious artifacts
- Comparing detectors on hard examples
- Fine-tuning for your own type of content
- Sizes
- several experts based on ConvNeXt and CLIP
- Hardware
- from: 1 GPU
Deepfake detection2022–2025
GRIP, University Federico II of Naples · Italy
Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.
- Checking scans of certificates and contracts for edits
- Highlighting altered photo regions for an expert
- Filtering out obviously redrawn documents before manual review
- Sizes
- a transformer model producing a map of suspicious regions
- Hardware
- from: 1 GPU
Moderation and safety2023–2025
Falconsai and Freepik · USA and Spain
Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.
- Filtering user photos and avatars
- Checking generated images before publishing
- Labeling a media library
- Sizes
- 86M
- Hardware
- from: Laptop
Computer vision2023–2025
Google · USA
Models that map images and text into a shared space: you can search photos by words and classify images without training. OpenAI's CLIP (2021) is the predecessor.
- Image search by text query
- Automatic catalog labeling and tagging
- Filtering prohibited content
- Sizes
- about 0.2B to 2B
- Hardware
- from: Laptop
Computer-use agents2024–2025
Microsoft · USA
Breaks a screenshot down into buttons, fields and icons with labels so a regular language model can understand and control the screen. It does not click itself; it serves as the agent's eyes.
- Mapping legacy software screens for automation
- Preparing an agent to work in an interface
- Checking that the required elements are on screen
- Sizes
- under 1B (detector + captioning)
- Hardware
- from: Laptop
Medicine2024–2025
Bioptimus · France
Pathology foundation models from France's Bioptimus with 1.1B parameters, plus the compact H0-mini. The first version is open under Apache 2.0. Does not replace a doctor; decisions are made by a specialist.
- Histology slide patch features for research models
- A prototype for sorting slides by tissue type
- Research projects linking morphology and molecular data
- Sizes
- 86M (H0-mini) – 1,1B
- Hardware
- from: Laptop
Computer vision2022–2025
University of Sydney and JD Explore Academy · Australia / China
A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.
- Body keypoints in photos and video
- Motion analysis in sports and rehabilitation
- Monitoring work postures and safety practices
- Sizes
- 33M – about 1B
- Hardware
- from: Laptop
Medicine2024–2025
Mahmood Lab (Mass General Brigham, Harvard) · USA
A foundation model for histology slides: turns patches of digital slides into features for tissue classification. Does not replace a doctor; decisions are made by a specialist.
- Research classifiers of tissue types from digital slides
- Finding similar cases in a slide archive
- Preparing features for research prognosis models
- Sizes
- about 300M (UNI) – about 680M (UNI2-h)
- Hardware
- from: Laptop
Medicine2024
Mahmood Lab (Mass General Brigham, Harvard) · USA
Image-plus-text models for pathology: search slides by an English description, classify without fine-tuning; TITAN describes a whole slide. Does not replace a doctor; decisions are made by a specialist.
- Text-query search across a slide archive for research
- Draft slide descriptions for research projects
- Tissue classification without labels at the start of a study
- Sizes
- about 160M – 300M
- Hardware
- from: Laptop
MedicineNot maintained2024
Paige · USA
Paige's pathology foundation model, trained on millions of digital slides. The first version is Apache 2.0, the second is for research only. Does not replace a doctor; decisions are made by a specialist.
- Slide patch features for research classifiers
- Selecting slides for re-review in research projects
- Comparison with other pathology models on your own archive
- Sizes
- about 632M
- Hardware
- from: Laptop
Image + textNot maintained2024
Microsoft · USA
A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.
- Reading text in photos
- Finding and highlighting objects
- Automatic photo captions
- Sizes
- 0.23B – 0.77B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023–2024
Sichuan University and co-authors · China
An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.
- Finding pasted and erased fragments in photos
- Checking document scans for edits
- A baseline when comparing manipulation-localization models
- Sizes
- a Vision Transformer based model
- Hardware
- from: 1 GPU
Computer visionNot maintained2021–2022
OpenAI · USA
The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.
- Image search by text query
- Automatic tags for a catalog
- Finding similar images
- Sizes
- about 0.15B – 0.6B
- Hardware
- from: Laptop
Computer visionNot maintained2022
Microsoft · USA
From an image it works out what kind of document it is: resume, diploma, certificate, contract. Helps sort candidate file bundles by type. A human makes the decision about a candidate; automatic screening without review must not be used.
- Sorting incoming candidate documents by type
- Checking that a document package is complete
- Finding the right scan in an archive
- Sizes
- base and large versions
- Hardware
- from: Laptop
Computer visionRUNot maintained2022
Sber AI and SberDevices (ai-forever) · Russia
A Russian version of CLIP: matches images with Russian captions. Lets you search photos by description and sort images into categories without training.
- Product search by photo and by Russian description
- Sorting images into categories without labeling
- Checking that a photo matches its caption
- Sizes
- 150M – 430M
- Hardware
- from: Laptop
Deepfake detectionNot maintained2020
MiniVision Technology · China
Practically the only fully open weight set for single-frame face liveness: it tells a live person from a photo, a screen or a mask. It errs in both directions - a person must be able to appeal a rejection.
- Liveness check when signing in by selfie
- Protecting an access system from a photo on a phone
- A check during remote customer identification
- Sizes
- two models of about 1.8 MB each
- Hardware
- from: Laptop