Open-source content moderation models

Moderation models screen messages, reviews and chatbot replies for toxicity, prohibited topics and jailbreak attempts. They sit as a filter before content is published or before an assistant answers. Check which violation categories the model distinguishes, how it performs in your languages, and whether you can customize the policy.

16 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
TextRUOllama2023–2026

Mistral

Mistral AI · France

European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).

  • Fast chat responses
  • Data extraction from text
  • Translation and multilingual work
Sizes
3B – 675B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safety2024–2026

GLiNER-PII

GLiNER community (Fastino, Knowledgator, NVIDIA and others) · USA

Small GLiNER-based models for finding personal data: passports, phone numbers, accounts, addresses. Data types are set in words. Russian is not officially supported.

  • Masking personal data before cloud AI
  • Finding passport data and bank details in documents
  • Checking data exports for leaks
Sizes
about 200M to 500M
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safety2024–2026

Aegis / Nemotron Safety Guard

NVIDIA · USA

NVIDIA content filters for bots, with separate models for keeping the conversation on topic and detecting jailbreaks. Safety Guard v3 was trained on 9 languages; Russian was tested only without fine-tuning.

  • Checking bot requests and replies
  • Keeping the bot within its topic
  • Detecting attempts to bypass rules
Sizes
4B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safetyOllama2024–2026

Granite Guardian

IBM · USA

IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.

  • Checking bot requests and replies
  • Finding made-up facts in knowledge-base answers
  • Checking your own rules written as text
Sizes
38M – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safety2026

OpenAI Privacy Filter

OpenAI · USA

Finds and hides personal data: names, addresses, phone numbers, emails, account numbers, passwords. Runs even in the browser. Trained mostly on English.

  • Removing personal data from text before sending it to cloud AI
  • Finding passwords and keys in texts
  • Anonymizing correspondence for analytics
Sizes
1.5B (50M active)
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safety2025

AprielGuard

ServiceNow · USA

A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.

  • Screening chatbot requests for attacks and jailbreaks
  • Filtering harmful model answers
  • Monitoring the actions of AI agents that use tools
Sizes
8B
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safetyOllama2025

gpt-oss-safeguard

OpenAI · USA

Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.

  • Moderation by internal company rules
  • Labeling disputed messages with an explanation
  • Checking reviews and listings before publishing
Sizes
20B – 120B
Hardware
from: 1 GPU
Commercial use allowedDetails
Moderation and safetyRUGGUF2025

Qwen3Guard

Alibaba (Qwen) · China

Safety filters for 119 languages, Russian among them. The Stream version checks a bot's reply while it is being generated and can cut it off on the fly.

  • Filtering bot requests in Russian
  • Stopping a dangerous reply during generation
  • Labeling messages by risk category
Sizes
0.6B – 8B
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safetyOllama2023–2025

Llama Guard

Meta · USA

Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.

  • Checking user questions to the bot
  • Checking bot replies before sending
  • Reporting which rule category was violated
Sizes
1B – 12B
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safety2024–2025

Prompt Guard

Meta · USA

Tiny classifiers that catch attempts to hack a bot: prompt injections and rule bypassing. The 86M version is multilingual, 22M is English only.

  • Protecting a bot from prompt injections
  • Checking emails and documents that reach an AI agent
  • Fast filter in front of a large model
Sizes
22M – 86M
Hardware
from: Laptop
Commercial use with conditionsDetails
Moderation and safety2023–2025

NSFW-классификаторы (Falconsai, Freepik)

Falconsai and Freepik · USA and Spain

Small models that tell explicit images from regular ones. The Freepik model distinguishes four levels of explicitness. They run on a CPU.

  • Filtering user photos and avatars
  • Checking generated images before publishing
  • Labeling a media library
Sizes
86M
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safetyOllama2024–2025

ShieldGemma

Google · USA

Gemma-based filters: they check text for dangerous and offensive content, and ShieldGemma 2 checks images. Focused on English.

  • Moderating user messages
  • Checking bot replies
  • Checking generated images before publishing
Sizes
2B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
Cybersecurity2024

Phishing Email Detection DistilBERT

cybersectony · not disclosed

A very light classifier for emails and links showing signs of phishing. It errs in both directions, so borderline emails are still reviewed by a person.

  • Flagging suspicious incoming emails
  • Checking links from correspondence before opening them
  • A first-level filter in a mail gateway
Sizes
about 66M
Hardware
from: Laptop
Commercial use allowedDetails
Moderation and safety2024

Piiranha

iiiorg · not disclosed

A popular detector of 17 types of personal data in six European languages. No Russian and a non-commercial license: suitable for trials and research.

  • Finding personal data in texts
  • Comparing the quality of PII detectors
Sizes
278M
Hardware
from: Laptop
Non-commercial onlyDetails
Moderation and safetyGGUFNot maintained2024

WildGuard

Ai2 (Allen Institute for AI) · USA

An open Ai2 filter: in a single pass it determines whether a request is harmful, whether a reply is harmful, and whether the bot refused needlessly. Works in English.

  • Checking requests to the bot
  • Checking bot replies
  • Finding unnecessary bot refusals on harmless questions
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Text analysisRUNot maintained2022–2023

Русские классификаторы токсичности и тональности

David Dale (cointegrated) and the community · Russia

Ready-made tiny rubert-tiny models for Russian text: detect rudeness and insults, sentiment and emotions. They run on a CPU in milliseconds.

  • Filtering insults in Russian chats and comments
  • Labeling reviews as positive, neutral or negative
  • Spotting irritated customers in requests
Sizes
12M – 29M
Hardware
from: Laptop
Commercial use with conditionsDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment