These models answer one question: is the material genuine? Was a face swapped in a video, is a live person in front of the camera, is the voice synthetic, was the image or text generated? Know the limits: any such detector errs in both directions, so its output is a reason to look closer, not proof. They run on your own server so disputed material never leaves it.
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
Marking images on the way out of your own pipeline
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
Checking images from new, unfamiliar generators
A baseline when comparing detectors
Fast rollout of a check without training a large model
A watermark for video and images that survives re-encoding and cropping. The detector errs in both directions: a missing mark does not prove a forgery, and finding one is a reason for a human to check.
Marking video created or processed by AI
Finding your own mark in re-uploaded clips
Protecting ad materials from being reused as someone else's
National Institute of Informatics, Yamagishi Lab · Japan
Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.
Checking audio recordings for synthesis
Fine-tuning for your own language and recording channel
IBM Research and The Chinese University of Hong Kong · USA
An AI-text detector trained together with a paraphraser: it was deliberately taught not to give up when the text has been rewritten. It errs in both directions; a human reviews the output.
Checking texts that may have been rewritten after generation
First-pass filtering in a newsroom or admissions office
The Chinese University of Hong Kong, Shenzhen (SCLBD) · China
Dozens of open face-swap detectors for video and photo under one codebase with ready weights. A detector errs in both directions: its output is a reason for a human to check, not proof of a forgery.
First-pass check of a submitted video or selfie
Comparing several detectors on your own data
Fine-tuning a detector for your own flow of applications
Sizes
Xception- and EfficientNet-class detectors, tens of millions of parameters
An image watermark that can be applied to individual regions: the model shows which part of the image is marked. It errs in both directions - a human reviews the result.
Xiaohongshu, USTC and Shanghai Jiao Tong University · China
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
Checking realistic AI images without obvious artifacts
Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.
Checking scans of certificates and contracts for edits
Highlighting altered photo regions for an expert
Filtering out obviously redrawn documents before manual review
Sizes
a transformer model producing a map of suspicious regions
A recent open AI-text detector on DeBERTa-v3-large, trained on the RAID dataset, with a separate version for academic work. It errs in both directions - a human always reviews the result.
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
Marking speech synthesized by your service
Finding your own mark in third-party publications
Checking whether synthesis was mixed into a call recording
Sizes
a watermark generator and detector, 16-bit message
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.
Finding pasted and erased fragments in photos
Checking document scans for edits
A baseline when comparing manipulation-localization models
One of the first open AI-text classifiers, trained on the HC3 corpus of paired human and ChatGPT answers. It errs in both directions: its output is a reason to talk to the author, not proof.
A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.
The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.
Practically the only fully open weight set for single-frame face liveness: it tells a live person from a photo, a screen or a mask. It errs in both directions - a person must be able to appeal a rejection.
Liveness check when signing in by selfie
Protecting an access system from a photo on a phone