Moderation models screen messages, reviews and chatbot replies for toxicity, prohibited topics and jailbreak attempts. They sit as a filter before content is published or before an assistant answers. Check which violation categories the model distinguishes, how it performs in your languages, and whether you can customize the policy.
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
GLiNER community (Fastino, Knowledgator, NVIDIA and others) · USA
Small GLiNER-based models for finding personal data: passports, phone numbers, accounts, addresses. Data types are set in words. Russian is not officially supported.
Masking personal data before cloud AI
Finding passport data and bank details in documents
NVIDIA content filters for bots, with separate models for keeping the conversation on topic and detecting jailbreaks. Safety Guard v3 was trained on 9 languages; Russian was tested only without fine-tuning.
IBM judge models: they catch harm, profanity and jailbreak attempts, and in RAG and agents check whether an answer is grounded in the documents. You can state your own rule in words.
Finds and hides personal data: names, addresses, phone numbers, emails, account numbers, passwords. Runs even in the browser. Trained mostly on English.
Removing personal data from text before sending it to cloud AI
Safety filters for 119 languages, Russian among them. The Stream version checks a bot's reply while it is being generated and can cut it off on the fly.
Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.
A very light classifier for emails and links showing signs of phishing. It errs in both directions, so borderline emails are still reviewed by a person.
Flagging suspicious incoming emails
Checking links from correspondence before opening them
An open Ai2 filter: in a single pass it determines whether a request is harmful, whether a reply is harmful, and whether the bot refused needlessly. Works in English.
Checking requests to the bot
Checking bot replies
Finding unnecessary bot refusals on harmless questions