Faces2021–2026
InsightFace (deepinsight) · China
The most widely used open toolkit for face detection and recognition. Many identity-preserving image generators are built on it. The pretrained weights are non-commercial.
- Detecting and comparing faces in photos
- Face-based access in prototypes
- Face processing as part of other AI systems
- Sizes
- packages from 16 MB to 407 MB
- Hardware
- from: Laptop
Deepfake detection2023–2026
Adobe Research and University of Surrey · USA
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
- Marking images on the way out of your own pipeline
- Checking the provenance of a submitted image
- Linking with content provenance metadata
- Sizes
- model types Q and P with different mark capacity
- Hardware
- from: Laptop
Deepfake detection2025–2026
University of Michigan · USA
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
- Checking submitted photos and illustrations
- Filtering AI images in a content flow
- Flagging suspicious images for manual review
- Sizes
- 22M
- Hardware
- from: Laptop
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
Voice: speakers and soundGGUF2022–2026
WeNet community · China
A set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.
- Voice verification of a customer during a call
- Finding repeat calls from the same person
- Splitting a recording by speaker
- Sizes
- from a few to tens of millions of parameters
- Hardware
- from: Laptop
Image + text2026
OpenMOSS (Fudan University) · China
An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.
- Analyzing long videos and finding events by time
- Real-time streaming video analysis
- Understanding photos and documents
- Sizes
- about 11B
- Hardware
- from: 1 GPU
Computer vision2025–2026
Roboflow · USA
Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.
- Object detection in video and photos
- Precise outlines of parts and defects
- Fine-tuning for your own object classes
- Sizes
- Nano – 2XL
- Hardware
- from: Laptop
Moderation and safety2024–2026
GLiNER community (Fastino, Knowledgator, NVIDIA and others) · USA
Small GLiNER-based models for finding personal data: passports, phone numbers, accounts, addresses. Data types are set in words. Russian is not officially supported.
- Masking personal data before cloud AI
- Finding passport data and bank details in documents
- Checking data exports for leaks
- Sizes
- about 200M to 500M
- Hardware
- from: Laptop
Image + textOllama2024–2026
Moondream (M87 Labs) · USA
A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.
- Finding and counting objects in photos
- Checking photos from field reports
- Captions and tags for a catalogue
- Sizes
- 2B – 9B-A2B
- Hardware
- from: Laptop
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
Moderation and safety2024–2026
NVIDIA · USA
NVIDIA content filters for bots, with separate models for keeping the conversation on topic and detecting jailbreaks. Safety Guard v3 was trained on 9 languages; Russian was tested only without fine-tuning.
- Checking bot requests and replies
- Keeping the bot within its topic
- Detecting attempts to bypass rules
- Sizes
- 4B – 8B
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2026
IBM · USA
IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.
- Checking bot requests and replies
- Finding made-up facts in knowledge-base answers
- Checking your own rules written as text
- Sizes
- 38M – 8B
- Hardware
- from: Laptop
Moderation and safety2026
OpenAI · USA
Finds and hides personal data: names, addresses, phone numbers, emails, account numbers, passwords. Runs even in the browser. Trained mostly on English.
- Removing personal data from text before sending it to cloud AI
- Finding passwords and keys in texts
- Anonymizing correspondence for analytics
- Sizes
- 1.5B (50M active)
- Hardware
- from: Laptop
Computer vision2023–2026
Ultralytics · USA
The most widely used real-time object detector: finds and marks items in video even on modest hardware. YOLOv5 came out back in 2020; the catalog starts from YOLOv8.
- Counting people, cars and goods on video
- Checking hard hats and workwear
- Spotting defects on the production line
- Sizes
- 2.4M – 68M
- Hardware
- from: Laptop
Cybersecurity2025–2026
Cisco (Foundation AI) · USA
Cisco models for information security based on Llama 3.1 8B: analysis of vulnerabilities, threats and incidents. Can be deployed inside your own perimeter.
- Analyzing vulnerability and threat reports
- Helping SOC analysts during incidents
- Mapping threats to MITRE ATT&CK
- Sizes
- 8B
- Hardware
- from: Laptop
Moderation and safety2025
ServiceNow · USA
A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.
- Screening chatbot requests for attacks and jailbreaks
- Filtering harmful model answers
- Monitoring the actions of AI agents that use tools
- Sizes
- 8B
- Hardware
- from: Laptop
Computer vision2025
Meta · USA
Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.
- Search photos and videos by description
- Catalog labeling and tagging
- Search across audio and video (PE-AV)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
- Marking video created or processed by AI
- Finding your own mark in re-uploaded clips
- Protecting ad materials from being reused as someone else's
- Sizes
- a mark of 96 to 1024 bits
- Hardware
- from: Laptop
Computer vision2023–2025
IDEA Research · China
Finds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.
- Finding objects by description without labeling
- Automatic data labeling for training
- Checking photos against requirements
- Sizes
- 172M – 3B
- Hardware
- from: Laptop
Moderation and safetyOllama2025
OpenAI · USA
Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.
- Moderation by internal company rules
- Labeling disputed messages with an explanation
- Checking reviews and listings before publishing
- Sizes
- 20B – 120B
- Hardware
- from: 1 GPU
Voice: speakers and sound2022–2025
NVIDIA · USA
NVIDIA models for "who is speaking": TitaNet recognizes a specific person's voice, Sortformer splits a recording into up to 4 speakers, including live during a call.
- Real-time speaker tagging in conversations
- Checking that the same person is calling (voiceprint)
- Preparing meeting transcripts
- Sizes
- 23M (TitaNet) – 117M (Sortformer)
- Hardware
- from: Laptop
Deepfake detection2025
National Institute of Informatics, Yamagishi Lab · Japan
Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.
- Checking audio recordings for synthesis
- Fine-tuning for your own language and recording channel
- Comparing several encoders on your own data
- Sizes
- 0,3B – 2B
- Hardware
- from: Laptop
Moderation and safetyRUGGUF2025
Alibaba (Qwen) · China
Safety filters for 119 languages, Russian among them. The Stream version checks a bot's reply while it is being generated and can cut it off on the fly.
- Filtering bot requests in Russian
- Stopping a dangerous reply during generation
- Labeling messages by risk category
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Deepfake detection2023–2025
IBM Research and The Chinese University of Hong Kong · USA
An AI-text detector trained together with a paraphraser: it was deliberately taught not to give up when the text has been rewritten. It errs in both directions; a human reviews the output.
- Checking texts that may have been rewritten after generation
- First-pass filtering in a newsroom or admissions office
- Comparison against simpler detectors
- Sizes
- about 355M (RoBERTa-large)
- Hardware
- from: Laptop
Faces2025
ByteDance · China
Transformer-based face recognition from ByteDance, one of the most accurate open models on benchmarks. Weights are published in ONNX format but are non-commercial.
- Comparing faces and searching a photo database
- Research on recognition accuracy
- Access control prototypes
- Sizes
- ViT-T – ViT-L
- Hardware
- from: Laptop
Deepfake detection2023–2025
The Chinese University of Hong Kong, Shenzhen (SCLBD) · China
Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.
- First-pass check of a submitted video or selfie
- Comparing several detectors on your own data
- Fine-tuning a detector for your own flow of applications
- Sizes
- Xception- and EfficientNet-class detectors, tens of millions of parameters
- Hardware
- from: Laptop
Deepfake detection2024–2025
Meta · USA
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
- Marking generated and edited images
- Finding a marked fragment inside a collage
- Tracking which parts of a picture were made by AI
- Sizes
- a mark encoder and decoder for images
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Clouditera · China
A Chinese open family for cybersecurity: reviewing vulnerabilities, analysing logs and traffic, explaining commands and scripts.
- Reviewing vulnerabilities and drafting fix recommendations
- Analysing logs and reconstructing an attack chain
- Explaining suspicious commands and scripts
- Sizes
- 1.5B – 14B
- Hardware
- from: Laptop
CybersecurityGGUF2025
Trendyol · Turkey
Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.
- Reviewing alerts and first-pass incident assessment
- Explaining suspicious activity in reports
- Helping the on-duty shift of a monitoring centre
- Sizes
- 32B и 70B
- Hardware
- from: 1 GPU
Deepfake detection2024–2025
Xiaohongshu, USTC and Shanghai Jiao Tong University · China
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
- Checking realistic AI images without obvious artifacts
- Comparing detectors on hard examples
- Fine-tuning for your own type of content
- Sizes
- several experts based on ConvNeXt and CLIP
- Hardware
- from: 1 GPU
Deepfake detection2022–2025
GRIP, University Federico II of Naples · Italy
Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.
- Checking scans of certificates and contracts for edits
- Highlighting altered photo regions for an expert
- Filtering out obviously redrawn documents before manual review
- Sizes
- a transformer model producing a map of suspicious regions
- Hardware
- from: 1 GPU
Moderation and safetyOllama2023–2025
Meta · USA
Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.
- Checking user questions to the bot
- Checking bot replies before sending
- Reporting which rule category was violated
- Sizes
- 1B – 12B
- Hardware
- from: Laptop
Moderation and safety2024–2025
Meta · USA
Tiny classifiers that catch attempts to hack a bot: prompt injections and rule bypassing. The 86M version is multilingual, 22M is English only.
- Protecting a bot from prompt injections
- Checking emails and documents that reach an AI agent
- Fast filter in front of a large model
- Sizes
- 22M – 86M
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
Songting Liu (Plachtaa) · China
Voice conversion without training: transfers timbre from a 1–30 second sample, can sing and work in real time; V2 also changes accent. Use only with the voice owner's consent.
- Re-voicing a video with a different voice
- Voice anonymization in recordings
- Real-time voice for streams
- Sizes
- about 70M – 200M
- Hardware
- from: Laptop
Moderation and safety2023–2025
Falconsai and Freepik · USA and Spain
Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.
- Filtering user photos and avatars
- Checking generated images before publishing
- Labeling a media library
- Sizes
- 86M
- Hardware
- from: Laptop
CybersecurityGGUF2023–2025
Kindo · USA
One of the best-known open families for security and DevSecOps work: reviewing code for weaknesses, test scenarios, explaining attacks.
- Finding weak spots in code and configurations
- Reviewing incidents and explaining attack techniques
- Drafting scripts and procedures for the security team
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2025
Google · USA
Gemma-based filters: they check text for dangerous and offensive content, and ShieldGemma 2 checks images. Focused on English.
- Moderating user messages
- Checking bot replies
- Checking generated images before publishing
- Sizes
- 2B – 27B
- Hardware
- from: Laptop
Deepfake detection2025
Desklib · India
A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.
- Checking submitted articles and reports
- Filtering templated reviews and applications
- First-pass check of student work
- Sizes
- 0.4B (DeBERTa-v3-large)
- Hardware
- from: Laptop
Cybersecurity2025
Trend Micro · Japan
Trend Micro cybersecurity models based on Llama 3.1 8B, fine-tuned on a corpus of security texts. A reasoning version is available.
- Answering questions about threats and vulnerabilities
- Analyzing cyberattack reports
- A base for fine-tuning for SOC tasks
- Sizes
- 8B
- Hardware
- from: Laptop
Image + text2023–2025
Alibaba DAMO Academy · China
Models that watch a video and answer questions about it: what happens, when, who does what. VideoLLaMA 3 at 2B and 7B is among the strongest in its size class.
- Video description and short summary
- Finding a moment in a recording by question
- Tagging a video archive
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Deepfake detection2024
Meta · USA
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
- Marking speech synthesized by your service
- Finding your own mark in third-party publications
- Checking whether synthesis was mixed into a call recording
- Sizes
- a watermark generator and detector, 16-bit message
- Hardware
- from: Laptop
Cybersecurity2024
cybersectony · not disclosed
A very light classifier for emails and links showing signs of phishing. It errs in both directions, so borderline emails are still reviewed by a person.
- Flagging suspicious incoming emails
- Checking links from correspondence before opening them
- A first-level filter in a mail gateway
- Sizes
- about 66M
- Hardware
- from: Laptop
Moderation and safety2024
iiiorg · not disclosed
A popular detector of 17 types of personal data in six European languages. No Russian and a non-commercial license: suitable for trials and research.
- Finding personal data in texts
- Comparing the quality of PII detectors
- Sizes
- 278M
- Hardware
- from: Laptop
Moderation and safetyGGUFNot maintained2024
Ai2 (Allen Institute for AI) · USA
An open Ai2 filter: in a single pass it determines whether a request is harmful, whether a reply is harmful, and whether the bot refused needlessly. Works in English.
- Checking requests to the bot
- Checking bot replies
- Finding unnecessary bot refusals on harmless questions
- Sizes
- 7B
- Hardware
- from: 1 GPU
Deepfake detectionNot maintained2024
UC Santa Barbara and co-authors · USA
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
- Checking long articles and reports as a whole
- Filtering machine text in a publication flow
- Comparing detectors on your own data
- Sizes
- about 150M (Longformer-base)
- Hardware
- from: Laptop
CybersecurityGGUFNot maintained2023–2024
ZySec AI · India
A small open assistant for security professionals: questions about standards, reviewing threats and vulnerabilities, drafting internal documents.
- Answering questions about security policies and standards
- First-pass review of threat reports
- Drafting internal protection guidelines
- Sizes
- 2.8B и 7B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023–2024
Sichuan University and co-authors · China
An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.
- Finding pasted and erased fragments in photos
- Checking document scans for edits
- A baseline when comparing manipulation-localization models
- Sizes
- a Vision Transformer based model
- Hardware
- from: 1 GPU
Deepfake detectionNot maintained2023
Hello-SimpleAI · China
One of the first open AI-text classifiers, trained on the HC3 corpus of paired human and ChatGPT answers. It errs in both directions: its output is a reason to talk to the author, not proof.
- First-pass check of student work
- Filtering templated applications and reviews
- Flagging suspicious texts for manual review
- Sizes
- about 125M (RoBERTa-base)
- Hardware
- from: Laptop
Music and soundNot maintained2023
LAION · Germany
CLIP for audio: maps audio and text into a shared space. Lets you search sounds and music by description and classify them without training. Text must be in English.
- Search sounds and music by description
- Automatic tags for an audio library
- Recognizing sound types (siren, breaking glass, voice)
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Deepfake detectionNot maintained2022–2023
EURECOM · France
A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.
- Spotting synthetic speech in calls
- Checking voice messages and recordings
- Fine-tuning for your own data and codecs
- Sizes
- about 0.3B (wav2vec 2.0 XLS-R encoder)
- Hardware
- from: Laptop
CybersecurityNot maintained2022–2023
Ehsan Aghaei and co-authors · USA
A compact language encoder trained on cybersecurity texts: tagging threat reports, finding entities and classification.
- Tagging threat reports and vulnerability bulletins
- Extracting entities from security texts
- Classifying and searching an internal incident base
- Sizes
- около 125M
- Hardware
- from: Laptop
Text analysisRUNot maintained2022–2023
David Dale (cointegrated) and the community · Russia
Ready-made tiny rubert-tiny models for Russian text: detect rudeness and insults, sentiment and emotions. They run on a CPU in milliseconds.
- Filtering insults in Russian chats and comments
- Labeling reviews as positive, neutral or negative
- Spotting irritated customers in requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
Music and soundNot maintained2022
MIT · USA
A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.
- Sound event recognition
- Tagging an audio archive
- Detecting alarm sounds
- Sizes
- about 87M
- Hardware
- from: Laptop
Computer visionNot maintained2021–2022
OpenAI · USA
The 2021 model that first linked images and text: search photos by words and classify them without training. English only; SigLIP 2 or PE are usually chosen today.
- Image search by text query
- Automatic tags for a catalog
- Finding similar images
- Sizes
- about 0.15B – 0.6B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2021
NAVER Clova AI Research and EURECOM · South Korea
The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.
- Voice check during phone authentication
- Filtering replays and synthesis in a voice menu
- A baseline when comparing voice detectors
- Sizes
- weight files of 0.4 and 1.3 MB
- Hardware
- from: Laptop
Deepfake detectionNot maintained2020
MiniVision Technology · China
Practically the only fully open weight set for single-frame face liveness: it tells a live person from a photo, a screen or a mask. It errs in both directions - a person must be able to appeal a rejection.
- Liveness check when signing in by selfie
- Protecting an access system from a photo on a phone
- A check during remote customer identification
- Sizes
- two models of about 1.8 MB each
- Hardware
- from: Laptop