Speech to textRUGGUF2025–2026
Microsoft · USA
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
- Transcribing long meetings with speaker labels
- Voicing podcasts and dialogues
- Real-time voice for assistants
- Sizes
- 0.5B – 9B
- Hardware
- from: Laptop
TextOllama2023–2026
DeepSeek · China
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
- Employee assistant on your own server
- Analysis of long contracts and reports
- Agents that work with tools and APIs
- Sizes
- 7B – 1.6T-A49B
- Hardware
- from: Laptop
TextRU2024–2026
Sber · Russia
Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.
- Russian-language employee assistant on your own server
- Customer replies and request handling in Russian
- Working with contracts and internal policies
- Sizes
- 10B-A1.8B – 702B-A36B
- Hardware
- from: Laptop
TextRU2022–2026
Yandex · Russia
Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.
- Russian-language assistant and chatbot
- Answers based on the company knowledge base
- Base for industry-specific fine-tuning
- Sizes
- 8B – 100B
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere Labs · Canada
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
- Translation and correspondence in less common languages
- Multilingual chat assistant
- Analysis of images with text (Vision)
- Sizes
- 3.3B – 35B
- Hardware
- from: Laptop
Text2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact text models that run directly on a device: laptop, phone or mini PC. The 1B and 2B MiniCPM5 models focus on tool calling and long context.
- A local chat assistant without the cloud
- Data extraction and text classification
- Tool calling and simple agents on low-end hardware
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Alibaba · China
A family of language models with strong Russian language support, from small versions for a laptop to a flagship on par with commercial APIs.
- Chatbot and knowledge-base assistant
- Replies to emails and customer requests
- Document parsing and classification
- Sizes
- 0,6B – 2,4T-A95B
- Hardware
- from: Laptop
Search and RAGRU2024–2026
Sber (SberDevices) · Russia
Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.
- Search across Russian-language documents
- RAG for chatbots in Russian
- Classifying requests and reviews
- Sizes
- 480M – 10B-A1.8B
- Hardware
- from: Laptop
Speech to textGGUF2025–2026
IBM · USA
IBM speech models for recognizing and translating speech in English, several European languages and Japanese. Designed for enterprise use.
- Transcribing business meetings
- Translating speech into text in another language
- Voice assistants
- Sizes
- 470M – 8B
- Hardware
- from: Laptop
TextOllama2023–2026
Zhipu AI (Z.ai) · China
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
- Corporate chat assistant
- Agents for routine office tasks
- Help for developers
- Sizes
- 1.5B – 744B-A40B
- Hardware
- from: Laptop
Text2024–2026
Tencent · China
Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.
- Corporate assistant
- Translation and multilingual texts
- Agents with tools
- Sizes
- 0.5B – 770B-A49B
- Hardware
- from: Laptop
Text2025–2026
Ant Group (inclusionAI) · China
An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.
- Corporate assistant
- Agents for office processes
- Financial analytics (Fin version available)
- Sizes
- 7.9B-A1.3B – 1T
- Hardware
- from: Laptop
TextRUOllama2024–2026
Cohere · Canada
Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.
- Knowledge-base answers with source citations
- Agents that work with internal systems
- Translation and correspondence in different languages
- Sizes
- 2.5B – 218B-A25B
- Hardware
- from: Laptop
TextRUOllama2025–2026
Liquid AI · USA
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
- Offline assistant on a laptop or phone
- Data extraction from documents
- Tool calling in apps
- Sizes
- 230M – 24B-A2B
- Hardware
- from: Laptop
TextOllama2026
Meta Superintelligence Labs · USA
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
- Agents with tool calling
- Analysis of screenshots, charts and documents
- Multilingual assistant
- Sizes
- 30B
- Hardware
- from: 1 GPU
TextRUOllama2023–2026
Mistral AI · France
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
- Fast chat responses
- Data extraction from text
- Translation and multilingual work
- Sizes
- 3B – 675B
- Hardware
- from: Laptop
Search and RAGRU2025–2026
NVIDIA · USA
NVIDIA embeddings for search and RAG. Nemotron-3-Embed, released in 2026, is under the permissive OpenMDW license and works in many languages.
- Search across corporate documents
- RAG for chatbots and assistants
- Search across images and pages (VL versions)
- Sizes
- 1B – 8B
- Hardware
- from: Laptop
Speech to textRUGGUF2024–2026
Sber · Russia
Sber's models for Russian speech recognition, among the most accurate for Russian. Includes emotion recognition, v3 with punctuation, and a multilingual version (Russian, Kazakh, Kyrgyz, Uzbek).
- Transcribing calls in Russian
- Meeting minutes
- Voice control of services
- Sizes
- 220M – 600M
- Hardware
- from: Laptop
TextGGUF2025–2026
Meituan · China
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
- Agents for orders and service processes
- Corporate assistant
- Analysis of long documents
- Sizes
- 1B – 1.6T-A48B
- Hardware
- from: Laptop
TextGGUF2025–2026
Kakao · South Korea
Compact Korean-English models from Kakao. Kanana 2 30B-A3B is fast thanks to MoE; small 1–3B versions suit a regular PC.
- Support chatbot
- Customer request classification
- Lightweight assistant on your own PC
- Sizes
- 1.3B – 30B-A3B
- Hardware
- from: Laptop
TextRU2024–2026
T-Bank · Russia
T-Bank models fine-tuned from Qwen for Russian: they write and reason in Russian noticeably better than the original. T-Lite is 8B, T-Pro 32B on one GPU; T-Search is a multi-step search agent in Russian and English.
- Russian-language support chatbot
- Analysis of requests and documents in Russian
- Answers based on the company knowledge base
- Sizes
- 7B – 36B-A3B
- Hardware
- from: Laptop
Voice: speakers and soundGGUF2022–2026
WeNet community · China
A set of ready-made voiceprint models: checks whether the same person speaks in two recordings and helps split a recording by speaker. One of the models is built into pyannote 3.x.
- Voice verification of a customer during a call
- Finding repeat calls from the same person
- Splitting a recording by speaker
- Sizes
- from a few to tens of millions of parameters
- Hardware
- from: Laptop
RerankersGGUF2024–2026
Jina AI · Germany
Strong multilingual rerankers; m0 also ranks pages as images (scans, slides). The latest versions are open for non-commercial use only.
- Refining search results before a chatbot answers
- Sorting retrieved PDF pages and slides
- Catalog and knowledge base search
- Sizes
- 33M – 2.4B
- Hardware
- from: Laptop
TextOllama2024–2026
Google · USA
Compact Google models that run well on a single computer; larger versions understand images. Includes CodeGemma for code, FunctionGemma 270M for function calling and the fast DiffusionGemma.
- Offline assistant on a laptop
- Reading photos of documents and receipts
- Customer request classification
- Sizes
- 270M – 31B
- Hardware
- from: Laptop
Speech to textRUGGUF2026
Alibaba (Qwen) · China
Speech recognition models from the Qwen team for 50+ languages, including Russian. They handle noise, singing and accents well.
- Transcribing calls and meetings
- Video subtitles
- Multilingual recognition
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
Speech to textGGUF2026
Cohere · Canada
Cohere's speech recognition model for 14 languages (Russian is not on the list), with a separate version for Arabic. Built for accurate transcription of business recordings.
- Transcribing meetings and interviews
- Subtitles
- Searching an audio archive
- Sizes
- 2B
- Hardware
- from: Laptop
Search and RAGRUGGUF2023–2026
Jina AI · Germany
Strong multilingual embeddings with long context; v5-omni understands text, images and audio. Recent versions are open for non-commercial use only.
- Search across documents in many languages
- Search across images and scans
- Classification and clustering
- Sizes
- 33M – 3.8B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Cognitive Computations (Eric Hartford) · USA
Uncensored fine-tunes of Llama, Mistral, Qwen and others that fulfill almost any request. Filtering and moderation are fully on the deployer; do not show it to customers without your own filter.
- Assistant that does not refuse legal but sensitive topics
- Internal tools under a strict system prompt
- Role-play and creative scenarios
- Sizes
- 0.5B – 405B
- Hardware
- from: Laptop
Speech to textRU2023–2026
NVIDIA · USA
Fast NVIDIA speech recognition models, including streaming ones for real-time use. Parakeet TDT v3 and Nemotron 3.5 ASR understand Russian.
- Transcribing calls and meetings
- Video subtitles
- Real-time voice input
- Sizes
- 110M – 2.5B
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
Resemble AI · USA
Speech synthesis with voice cloning and adjustable expressiveness. The multilingual version supports 23 languages, including Russian; Turbo and Flash are sped up for live dialogue.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- about 350M – 500M
- Hardware
- from: Laptop
Text to speechRU2025–2026
OpenMOSS (Fudan University) · China
A speech synthesis family: multi-voice dialogue voicing (TTSD), fast synthesis for live conversation and the tiny Nano. Version 1.5 supports 30+ languages, including Russian.
- Voicing podcasts and dialogues
- Voice for an assistant
- Voice cloning
- Sizes
- 100M – 8.5B
- Hardware
- from: Laptop
TextGGUF2025–2026
StepFun · China
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
- High-load agents
- Analysis of documents with diagrams and screenshots
- Help for developers
- Sizes
- 8B – 321B
- Hardware
- from: 1 GPU
TranslationRU2025–2026
Tencent · China
Tencent translators for 33 languages; the first version won the WMT25 competition. Russian is supported. The small 1.8B version runs on a laptop; the new Hy-MT2 is under Apache 2.0.
- Translating documents while keeping formatting
- Translation with a set glossary of terms
- Translating correspondence with Chinese partners
- Sizes
- 1.8B – 30B-A3B
- Hardware
- from: Laptop
Moderation and safety2024–2026
NVIDIA · USA
NVIDIA content filters for bots, with separate models for keeping the conversation on topic and detecting jailbreaks. Safety Guard v3 was trained on 9 languages; Russian was tested only without fine-tuning.
- Checking bot requests and replies
- Keeping the bot within its topic
- Detecting attempts to bypass rules
- Sizes
- 4B – 8B
- Hardware
- from: Laptop
Image + textOllama2024–2026
OpenBMB (ModelBest and Tsinghua University) · China
Compact vision models that run even on a phone or laptop. Good at reading text in photos and understanding video; version 4.6 is only 1.3B.
- On-device text recognition in photos
- Processing receipts and documents without sending them to the cloud
- Describing photos and video
- Sizes
- 1.3B – 8B
- Hardware
- from: Laptop
Text to speechRU2025–2026
Supertone · South Korea
Very fast, lightweight speech synthesis that runs directly on the device, without a GPU or the cloud. Supertonic 3 speaks 31 languages, including Russian.
- Voicing voice bot replies on an ordinary server
- Voiceover in offline and mobile apps
- Reading texts and notifications aloud
- Sizes
- about 99M
- Hardware
- from: Laptop
Rerankers2025–2026
NVIDIA · USA
A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.
- Reordering passages before an AI assistant answers
- Sorting retrieved scan and PDF pages
- Search across internal policies and instructions
- Sizes
- 1B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2026
IBM · USA
Lightweight IBM embeddings for enterprise search, trained on data with clear rights. R2, released in 2026, became multilingual.
- Search across corporate documents
- RAG on a regular server without a GPU
- Reranking results
- Sizes
- 30M – 311M
- Hardware
- from: Laptop
Text to speechRUGGUF2025–2026
OpenBMB (ModelBest, Tsinghua University) · China
Speech synthesis with voice cloning and natural intonation. VoxCPM2 supports 30 languages, including Russian.
- Voice cloning
- Voicing videos and audiobooks
- Voice for an assistant
- Sizes
- 0.5B – 2.3B
- Hardware
- from: Laptop
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
TranslationRU2020–2026
Helsinki-NLP, University of Helsinki · Finland
More than a thousand small translators, each for its own language pair. Russian-English and back are available. Fast even on a regular CPU.
- Bulk translation of short texts
- Translation right on the server without a GPU
- Translating reviews and requests before analysis
- Sizes
- 25M – 240M
- Hardware
- from: Laptop
Text to speech2025–2026
Kyutai · France
Streaming speech recognition and synthesis models from the makers of Moshi: they start speaking and transcribing without waiting for the end of a phrase. Pocket TTS (100M) runs on a CPU. English, French and a few other European languages, no Russian.
- Streaming speech transcription for voice bots
- Voicing replies with minimal delay
- Speech synthesis on a server without a GPU (Pocket TTS)
- Sizes
- 100M (Pocket TTS) – 2.6B
- Hardware
- from: Laptop
TextGGUF2025–2026
ServiceNow · USA
ServiceNow 15B models with step-by-step reasoning that fit on a single GPU. From version 1.5 they also understand images and are good at calling tools.
- A reasoning assistant for internal services
- Tool calling and enterprise agents
- Analysing screenshots and documents with images
- Sizes
- 5B – 15B
- Hardware
- from: Laptop
Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
Speech to textRUGGUF2025–2026
Mistral AI · France
Mistral's speech models: they understand audio, transcribe and answer questions about a recording. The Realtime version recognizes speech live and supports Russian; speech synthesis is also available.
- Transcribing and summarizing recordings
- Asking questions about audio
- Real-time recognition
- Sizes
- 3B – 24B
- Hardware
- from: Laptop
TextGGUF2024–2026
Sarvam AI · India
Indian models focused on 22 languages of India. Sarvam 30B and 105B (2026) are MoE models with strong reasoning and agent skills.
- Multilingual customer support
- Reasoning and calculation tasks
- Agents with tool calling
- Sizes
- 2B – 105B-A10B
- Hardware
- from: Laptop
Voice: speakers and soundRU2026
FireRedTeam (Xiaohongshu) · China
A speech and sound event detector: tells apart speech, singing and music. In a 102-language test (the FLEURS set, which includes Russian) it beat Silero VAD and TEN VAD. Has a streaming mode.
- Cutting recordings before speech recognition
- Separating speech from music and singing in broadcasts and videos
- Speech detection in voice bots
- Sizes
- compact, exact size not stated
- Hardware
- from: Laptop
Search and RAGRU2026
Microsoft · USA
Microsoft's 2026 multilingual embeddings with context up to 32K tokens; Russian is on the language list. The 270M and 0.6B versions run on a regular server, 27B is the most accurate.
- Multilingual knowledge base search
- Picking passages for RAG
- Search across long documents
- Sizes
- 270M – 27B
- Hardware
- from: Laptop
Voice assistants2024–2026
Kyutai · France
A voice assistant that listens and speaks at the same time, with no delay for recognition and synthesis. Hibiki does simultaneous speech-to-speech translation between several European languages.
- Real-time voice conversation partner
- Simultaneous speech translation
- Zero-latency voice interfaces
- Sizes
- 2B – 7B
- Hardware
- from: Laptop
Avatars2024–2026
Ant Group · China
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
- Presenter video from a photo and voice
- Avatar with gestures for presentations
- Voiced characters
- Sizes
- up to 1.3B
- Hardware
- from: 1 GPU
AvatarsGGUF2025–2026
Alibaba (Quark) · China
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
- Live avatar for customer dialogue
- Endless broadcasts with a presenter
- Interactive characters
- Sizes
- 14B
- Hardware
- from: 1 GPU
Search and RAGRUOllama2025–2026
Alibaba (Qwen) · China
Embeddings and rerankers based on Qwen3, among the best open ones for multilingual search, including Russian. VL versions search images, screenshots and video.
- Knowledge base search for RAG
- Reranking results before answering
- Search across scans, slides and screenshots
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Search and RAG2026
Voyage AI (MongoDB) · USA
The only open model in the Voyage 4 line: its vectors are compatible with the paid larger versions, so you can start locally and move to the API later.
- Document search on your own server
- RAG for small knowledge bases
- Finding similar texts
- Sizes
- about 340M
- Hardware
- from: Laptop
Text to speechRUGGUF2026
Alibaba (Qwen) · China
Speech synthesis in 10 languages, including Russian: voice cloning from 3 seconds, ready-made voices and creating a voice from a text description.
- Voice for a bot or assistant
- Cloning a brand voice
- Choosing a voice by description
- Sizes
- 0.6B – 1.7B
- Hardware
- from: Laptop
TextRUOllama2023–2026
Technology Innovation Institute (TII) · UAE
A family from Abu Dhabi: from the early Falcon 40B and 180B to hybrid Falcon-H1 and tiny Falcon-H1-Tiny models of 90–600M parameters for devices.
- Assistant and answers based on documents
- Running on low-end hardware and devices
- Tool calling in simple agents
- Sizes
- 90M – 180B
- Hardware
- from: Laptop
TextGGUF2024–2026
AI21 Labs · Israel
A hybrid of Transformer and Mamba with a window of up to 256K tokens: handles long documents faster than conventional models. Jamba2 focuses on accurate, source-based answers.
- Answers based on long policies and contracts
- Knowledge-base search (RAG)
- Summaries of large documents
- Sizes
- 3B – 398B-A94B
- Hardware
- from: Laptop
TextRUGGUF2024–2026
UTTER consortium (Unbabel, universities of Lisbon, Edinburgh, Amsterdam and others) · European Union
European language models trained on all EU languages and several others, with a focus on translation. Russian is supported. Permissive license.
- Translation and localization of texts
- Answering questions in different languages
- Draft emails for foreign partners
- Sizes
- 1.7B – 22B
- Hardware
- from: Laptop
TranslationRUOllama2026
Google · USA
Translators based on Gemma 3 for 55 languages that can also translate text in images. Russian is supported. The 4B version fits on a laptop.
- Translating documents and correspondence
- Translating text from screenshots and photos
- Localizing websites and apps
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Voice: speakers and soundRU2025–2026
Daily (Pipecat) · USA
Uses intonation to tell whether a person has finished a thought or just paused, so a voice bot does not interrupt. Version 3 is 8 MB, runs on a CPU and understands 23 languages, including Russian.
- Voice bot does not interrupt the customer during pauses
- Fast reply when the customer has really finished
- An add-on to a standard speech detector in voice assistants
- Sizes
- 8M (v3) – 580M (v1)
- Hardware
- from: Laptop
Voice assistantsGGUF2026
NVIDIA · USA
A voice conversation partner based on Moshi that listens and speaks at the same time and can be interrupted. The role is set by text, the voice by a sample recording. English only.
- A voice assistant with a set role
- A conversation simulator for staff training
- Voice interfaces without delay
- Sizes
- 7B
- Hardware
- from: 1 GPU
Speech to text2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Alibaba's set of fast speech recognition models, primarily for Chinese and Asian languages. SenseVoice also detects emotions and sound events.
- Transcribing calls
- Detecting emotions in the voice
- Recognizing laughter, music and other sounds
- Sizes
- about 230M – 800M
- Hardware
- from: Laptop
Text to speechRU2024–2025
Alibaba (Tongyi, FunAudioLLM) · China
Speech synthesis with voice cloning from a short sample and streaming output for live dialogue. Version 3 supports 9 languages, including Russian.
- Voice for a bot or assistant
- Cloning a brand voice
- Voicing videos
- Sizes
- 300M – 0.5B
- Hardware
- from: Laptop
Text2025
Naver · South Korea
Open smaller models from Korea's Naver: from 0.5B to 32B, including reasoning Think versions and multimodal versions that understand images.
- Lightweight Korean-English assistant
- Analysis of images and documents
- Text classification
- Sizes
- 0.5B – 32B
- Hardware
- from: Laptop
TextRUGGUF2024–2025
Vikhr Models · Russia
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
- Russian-language assistant on your own PC or server
- Knowledge-base answers (RAG)
- Texts and emails in Russian
- Sizes
- 0.5B – 24B
- Hardware
- from: Laptop
Moderation and safety2025
ServiceNow · USA
A guard model that catches both harmful content and attacks on AI (prompt injection, jailbreaks), including when agents use tools.
- Screening chatbot requests for attacks and jailbreaks
- Filtering harmful model answers
- Monitoring the actions of AI agents that use tools
- Sizes
- 8B
- Hardware
- from: Laptop
TextGGUF2023–2025
Inception (G42), MBZUAI and Cerebras · UAE
A model family for Arabic and English, including Gulf dialects. Suits companies working with Arabic-speaking customers and government bodies in the region.
- A chatbot in Arabic and English
- Translating and summarising documents in Arabic
- Classifying customer requests
- Sizes
- 256M – 70B
- Hardware
- from: Laptop
TextOllama2023–2025
Nous Research · USA
Nous Research fine-tunes on top of Llama, Mistral, Qwen and Seed-OSS. Valued for precise instruction following, function calling and strict JSON output; they refuse less often than the originals; Hermes 4 has a reasoning mode.
- Agents that call functions and APIs
- Data extraction in strict JSON format
- Assistant with flexible role and tone settings
- Sizes
- 3B – 405B
- Hardware
- from: Laptop
Text to speechRU2022–2025
Silero · Russia
Lightweight Russian speech synthesis that runs on a regular CPU. Version v5 added CIS languages and languages of Russia's peoples: Tatar, Bashkir, Yakut, Kazakh and others.
- Voicing voice bot replies
- Reading texts in Russian
- Voices in the languages of Russia's peoples
- Sizes
- tens of megabytes
- Hardware
- from: Laptop
Voice: speakers and soundRU2020–2025
Silero · Russia
The most popular open speech detector: tells voice apart from silence and noise. Processes an audio chunk in under a millisecond on a single CPU core; trained on recordings in more than 6,000 languages.
- Cutting calls and recordings before speech recognition
- Detecting when the customer is speaking in a voice bot
- Filtering out silence and noise to save on transcription
- Sizes
- about 2 MB
- Hardware
- from: Laptop
TextOllama2025
Deep Cogito · USA
Fine-tuned Llama, Qwen and DeepSeek models with a hybrid mode: answer immediately or reason first. The 671B v2.1 flagship spends noticeably fewer tokens on reasoning than DeepSeek R1.
- A chat assistant with a reasoning mode
- Writing code and calling tools
- Answering complex questions about documents
- Sizes
- 3B – 671B
- Hardware
- from: Laptop
Moderation and safetyOllama2025
OpenAI · USA
Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.
- Moderation by internal company rules
- Labeling disputed messages with an explanation
- Checking reviews and listings before publishing
- Sizes
- 20B – 120B
- Hardware
- from: 1 GPU
Image + textOllama2023–2025
Alibaba (Qwen team) · China
One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.
- Extracting data from scanned invoices and delivery notes
- Analysing photos of products and shelves
- Analysing video and camera footage
- Sizes
- 2B – 235B-A22B
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
NVIDIA · USA
NVIDIA models for "who is speaking": TitaNet recognizes a specific person's voice, Sortformer splits a recording into up to 4 speakers, including live during a call.
- Real-time speaker tagging in conversations
- Checking that the same person is calling (voiceprint)
- Preparing meeting transcripts
- Sizes
- 23M (TitaNet) – 117M (Sortformer)
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Mixedbread · Germany
Embeddings and rerankers from Germany's Mixedbread. mxbai-embed-large is one of the most downloaded English search models; the v2 rerankers cover 100+ languages, including Russian.
- Search across a knowledge base
- Reranking results before a bot answers
- Product catalog search
- Sizes
- 17M – 1.5B
- Hardware
- from: Laptop
TextRU2025
Avito Tech · Russia
Avito's model based on Qwen3-8B, retrained for Russian: its own tokenizer makes Russian text 15–25% faster. Supports function calling.
- Product and listing descriptions in Russian
- Chatbot that calls internal services
- Request analysis and classification
- Sizes
- 7.9B
- Hardware
- from: Laptop
Search and RAGOllama2025
Google · USA
A small multilingual embedding model based on Gemma 3 that runs even on a phone or laptop without internet.
- On-device document search
- RAG without sending data outside
- Text classification
- Sizes
- 300M
- Hardware
- from: Laptop
Speech to textRU2023–2025
Alpha Cephei · Russia
Offline Russian speech recognition that runs even on a Raspberry Pi or a phone, without internet. Streaming models for live audio and simple Russian speech synthesis, Vosk TTS, are available.
- Transcribing Russian calls and recordings without the cloud
- Voice control in apps and kiosks
- Low-latency streaming speech recognition
- Sizes
- about 45 MB – 1.8 GB
- Hardware
- from: Laptop
Voice assistantsRUGGUF2025
Alibaba (Qwen) · China
Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.
- Voice assistant for customers
- Analyzing calls and videos
- Voice answers about documents and images
- Sizes
- 3B – 30B-A3B
- Hardware
- from: Laptop
Moderation and safetyRUGGUF2025
Alibaba (Qwen) · China
Safety filters for 119 languages, Russian among them. The Stream version checks a bot's reply while it is being generated and can cut it off on the fly.
- Filtering bot requests in Russian
- Stopping a dangerous reply during generation
- Labeling messages by risk category
- Sizes
- 0.6B – 8B
- Hardware
- from: Laptop
Voice: speakers and sound2022–2025
pyannoteAI (Hervé Bredin) · France
The most widely used open tool for splitting a recording by speaker: who spoke and when. Usually paired with speech recognition. Weights are issued after a short form on HF.
- Tagging calls: which part is the agent, which is the customer
- Meeting minutes with speaker labels
- Preparing recordings for transcription and analysis
- Sizes
- a few million parameters
- Hardware
- from: Laptop
Search and RAGOllama2023–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
Some of the most popular embeddings for search and RAG. The main v1.5 versions target English and Chinese; for Russian, BAAI has a separate model, bge-m3.
- Search across English-language documents
- Picking passages for chatbot answers (RAG)
- Code search (bge-code)
- Sizes
- 24M – 9B
- Hardware
- from: Laptop
TextOllama2025
OpenAI · USA
OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.
- AI agent that calls internal systems
- Answers based on internal policies
- Drafts of emails and reports
- Sizes
- 20B, 120B
- Hardware
- from: 1 GPU
TextRU2024–2025
Lomonosov Moscow State University Research Computing Center, LAIR lab (RefalMachine) · Russia
Qwen models adapted for Russian: a new tokenizer plus further training on Russian texts. As a result, Russian text is generated up to twice as fast as with the original model of the same size.
- Russian-language assistant on your own server
- Answers based on company documents (RAG) in Russian
- Analysis and summaries of long Russian texts
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Voice: speakers and sound2024–2025
Alibaba (Tongyi Lab) · China
Alibaba's set of speech cleanup models: noise suppression, separating overlapping voices, upscaling audio to 48 kHz, and isolating a voice using video of the speaker's face.
- Noise suppression in conversation recordings
- Separating two voices speaking at once
- Improving old and phone recordings
- Sizes
- under 1B
- Hardware
- from: Laptop
Text to speechRU2025
ESpeech (independent group of Russian-speaking developers) · Russia
Russian speech synthesis with voice cloning based on the F5-TTS architecture, trained on Russian speech datasets collected by the authors. Stress is placed automatically. Several variants, including a "podcaster" one.
- Voicing videos and audiobooks in Russian
- Cloning a narrator's voice from a sample
- Voice for a bot or assistant in Russian
- Sizes
- about 340M
- Hardware
- from: Laptop
Speech to textRU2025
T-Bank · Russia
A compact T-Bank streaming model for recognizing Russian speech in phone calls. Works in real time even without a GPU.
- Transcribing phone calls
- Voice robots on the line
- Call quality control
- Sizes
- 72M
- Hardware
- from: Laptop
TextRUOllama2024–2025
Hugging Face · USA
Tiny open Hugging Face models for phones and laptops. SmolLM3 (3B) can reason and handle long context; the full training recipe is open.
- Simple on-device assistant
- Classification and routing of requests
- Base for fine-tuning on a narrow task
- Sizes
- 135M – 3B
- Hardware
- from: Laptop
Voice: speakers and sound2025
Agora (TEN project) · USA / China
A lightweight speech detector for real-time voice assistants: it notices the start and end of a phrase faster than Silero VAD. Runs on servers, phones and in the browser.
- Zero-lag speech detection in a voice bot
- Fast assistant response at the end of a phrase
- Use in mobile apps and the browser
- Sizes
- very small, the library is smaller than Silero VAD
- Hardware
- from: Laptop
Voice: speakers and soundRU2023–2025
Community (xbgoose and others), Dusha dataset from SberDevices · Russia
Models that detect emotion from voice in Russian speech: neutral, anger, positive, sadness. Trained on the open Dusha dataset from SberDevices.
- Finding calls with irritated customers
- Assessing the tone of operator conversations
- Prioritizing complaints in a call center
- Sizes
- 21M – 316M
- Hardware
- from: Laptop
TranslationRUGGUF2024–2025
Unbabel · Portugal
Language models tailored for translation and multilingual text work: translating, editing, and assessing translation quality. Russian is supported. Non-commercial license.
- Translation that respects context and terminology
- Post-editing machine translation
- Assessing the quality of a finished translation
- Sizes
- 2B – 72B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
Skywork (Kunlun Tech) · China
Reward models: they score how good a language model's answer is for the user. Used for fine-tuning your own models and picking the best of several answers.
- Choosing the best of several bot answers
- Scoring answer quality during model fine-tuning
- Comparing models before rollout
- Sizes
- 0.6B – 27B
- Hardware
- from: Laptop
Text analysisRU2023–2025
deepvk (VK) · Russia
Russian encoders from the VK team: RuModernBERT reads long texts, USER produces vectors for search, GeRaCl classifies texts by topic without training.
- Classifying requests without labeled data
- Knowledge base search in Russian
- Analyzing long contracts
- Sizes
- 35M – 360M
- Hardware
- from: Laptop
Rerankers2023–2025
Stanford NLP, later Answer.AI and LightOn · USA and France
A different search principle: every word of the question is compared with every word of the document, not the two texts as a whole. The index is heavier than with ordinary embeddings. The model cards list English.
- Search across a knowledge base of long documents
- Reordering retrieved passages
- Search across policies and technical documentation
- Sizes
- about 33M – 150M
- Hardware
- from: Laptop
Fact-checking and judgesRU2024–2025
NVIDIA · USA
Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data to fine-tune your own model
- Scoring assistant answers in Russian and other languages
- Sizes
- 49B, 70B and 340B
- Hardware
- from: Cluster
TextOllama2023–2025
Meta · USA
The models that started mass open source in AI. A huge ecosystem of fine-tuned versions and tools.
- Assistant for employees
- Summaries of meetings and documents
- Base for industry-specific fine-tuning
- Sizes
- 1B – 405B
- Hardware
- from: Laptop
TextRUGGUF2023–2025
Ilya Gusev (IlyaGusev) · Russia
The best-known Russian community fine-tune: open models (Llama, Mistral, Gemma, YandexGPT) trained to act as a Russian-speaking assistant. A convenient starting point for a Russian chatbot on your own server.
- Russian-language chat assistant
- Answers based on the company knowledge base
- Drafts of emails, descriptions and posts in Russian
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Moderation and safetyOllama2023–2025
Meta · USA
Filter models that check chatbot requests and replies for dangerous topics against a list of categories. Version 4 also checks images. Russian is not officially supported.
- Checking user questions to the bot
- Checking bot replies before sending
- Reporting which rule category was violated
- Sizes
- 1B – 12B
- Hardware
- from: Laptop
Voice assistants2025
Moonshot AI · China
A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.
- Speech recognition
- Detecting emotions and sound events
- Speech-to-speech voice dialogue
- Sizes
- 7B
- Hardware
- from: 1 GPU
Speech to textRUGGUF2022–2025
OpenAI · USA
Speech recognition in 99 languages, including Russian. The de facto standard for transcribing calls and meetings. Hugging Face's faster Distil-Whisper is English only.
- Transcription of calls and video meetings
- Video subtitles
- Voice messages to text
- Sizes
- 39M – 1,5B
- Hardware
- from: Laptop
Avatars2024–2025
Tencent Music (Lyra Lab) · China
Real-time lip sync: matches the mouth in a video to new audio. Suits video translation and live avatars.
- Dubbing video into another language
- Live avatar in a video chat
- Editing speech in a finished video
- Sizes
- under 1B
- Hardware
- from: Laptop
Search and RAGRUOllama2024–2025
Nomic AI · USA
Fully open embeddings, with weights, data and training code. v2 is multilingual on MoE; there are versions for code and for searching PDF pages.
- Search across documents and a knowledge base
- Code search
- Search across scans and PDFs without text recognition
- Sizes
- 137M – 7B
- Hardware
- from: Laptop
Text to speechRUGGUF2023–2025
Rhasspy / Open Home Foundation · USA
Very fast speech synthesis that runs even on a Raspberry Pi. Ready-made voices in 35+ languages, including several Russian ones.
- Voicing notifications and bot replies
- Voice for offline devices
- Voice menus
- Sizes
- about 5M – 30M
- Hardware
- from: Laptop
Text to speechGGUF2025
Canopy Labs · USA
Language-model-based speech synthesis with lively intonation and emotional cues. Responds quickly, suitable for voice assistants. Mainly English.
- Real-time voice for an assistant
- Emotional voiceover
- Voice cloning
- Sizes
- 3B
- Hardware
- from: Laptop
Text to speechGGUF2025
Sesame · USA
A conversational speech model that takes the context of the conversation into account and sounds like a real person. English only.
- Voice for a conversational assistant
- Voicing dialogues
- Voice product prototypes
- Sizes
- 1B
- Hardware
- from: Laptop
Moderation and safetyOllama2024–2025
Google · USA
Gemma-based filters: they check text for dangerous and offensive content, and ShieldGemma 2 checks images. Focused on English.
- Moderating user messages
- Checking bot replies
- Checking generated images before publishing
- Sizes
- 2B – 27B
- Hardware
- from: Laptop
Computer-use agents2024–2025
Salesforce · USA
Salesforce models for function calling and agents: they pick the right tool and fill in its parameters. Strong on benchmarks, but the license is non-commercial.
- Calling APIs and internal services on user request
- Multi-step agents with several tools
- Comparing approaches before choosing a commercial model
- Sizes
- 1B – 8x22B
- Hardware
- from: Laptop
TextOllama2023–2025
Ai2 · USA
Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.
- Employee assistant on your own server
- Math and precise instruction following
- Reference recipe for your own fine-tuning
- Sizes
- 7B – 405B
- Hardware
- from: Laptop
TextGGUF2025
HUMAIN (formerly SDAIA) · Saudi Arabia
A Saudi model for Arabic and English, trained from scratch. One 7B version is openly available.
- An Arabic-language assistant
- Answering questions about documents
- Writing and editing texts in Arabic
- Sizes
- 7B
- Hardware
- from: Laptop
Search and RAGRU2023–2025
Alibaba · China
Alibaba embeddings and rerankers for search: from tiny to 7B based on Qwen2. There is a multilingual mGTE version with long context.
- Semantic search across documents
- Reranking search results
- Clustering and classifying texts
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Search and RAGRUOllama2019–2025
UKP Lab (TU Darmstadt), later Hugging Face · Germany
The classic for meaning-based search: small, fast models that run even on a modest server without a GPU. The multilingual versions understand Russian.
- Search across a knowledge base and FAQ
- Finding similar tickets and duplicates
- Grouping reviews and requests by topic
- Sizes
- about 20M – 470M
- Hardware
- from: Laptop
Visual document search2025
LlamaIndex · USA
A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.
- Search across scans and PDFs without OCR
- Search across invoices, acts and contracts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
Fact-checking and judges2025
Atla · UK
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring chatbot answers against your own criteria
- Comparing two versions of a prompt or model
- Filtering out weak answers before they reach a person
- Sizes
- 8B
- Hardware
- from: 1 GPU
Search and RAGRUOllama2024
Snowflake · USA
Snowflake embeddings built specifically for search. Version 2.0 is multilingual (Russian is on the language list), handles long texts up to 8K tokens and can compress vectors.
- Search across documents and knowledge bases
- Picking passages for RAG
- Search across reports and internal data
- Sizes
- 22M – 568M
- Hardware
- from: Laptop
Finance2023–2024
Du Xiaoman (Duxiaoman-DI) · China
A large Chinese model family for the financial industry: advice, document reading and long texts up to 8k-16k. Not investment advice: decisions are made by a specialist.
- Answering customer questions about banking products
- Working through long financial documents
- An internal assistant for a finance company regulations
- Sizes
- 6B – 176B
- Hardware
- from: 1 GPU
Deepfake detection2024
Meta · USA
An imperceptible mark in synthetic speech plus a fast detector that finds it even inside a fragment of a long recording. The detector errs in both directions: a hit is a reason for a human to check, not proof.
- Marking speech synthesized by your service
- Finding your own mark in third-party publications
- Checking whether synthesis was mixed into a call recording
- Sizes
- a watermark generator and detector, 16-bit message
- Hardware
- from: Laptop
Fact-checking and judges2024
Patronus AI · USA
A small judge: it scores against your criteria and highlights which part of the answer led to that score. The license is non-commercial. The judge itself makes mistakes and does not replace manual review.
- Scoring answers against your criteria with an explanation
- Understanding why a score was lowered
- Bulk review of assistant conversations
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
TextRU2024
MTS AI (MWS AI) · Russia
A lightweight Russian-language model from MTS AI for Russian texts: answers, summaries, drafts. A ready version for CPU without a GPU is available. The larger Cotype Pro is not released openly.
- Drafts of emails and descriptions in Russian
- Short document summaries
- Answers to common customer questions
- Sizes
- 1.5B
- Hardware
- from: Laptop
TranslationRU2024
deepvk (VK) · Russia
Compact Kazakh-Russian translators from VK. At 197M they translate as well as the 600M NLLB and run on a regular CPU.
- Translating requests from Kazakh to Russian
- Translating documents and instructions into Kazakh
- Bilingual customer support
- Sizes
- 197M
- Hardware
- from: Laptop
Cybersecurity2024
cybersectony · not disclosed
A very light classifier for emails and links showing signs of phishing. It errs in both directions, so borderline emails are still reviewed by a person.
- Flagging suspicious incoming emails
- Checking links from correspondence before opening them
- A first-level filter in a mail gateway
- Sizes
- about 66M
- Hardware
- from: Laptop
Fact-checking and judges2024
Flow AI · not disclosed
A small judge model: it checks an answer against your instruction and gives a score with an explanation. Fits on a modest server. The judge itself makes mistakes and does not replace manual review.
- Checking AI assistant answers against the instruction
- Bulk scoring of exported conversations
- Quality control before rolling out changes
- Sizes
- 3.8B (based on Phi-3.5-mini)
- Hardware
- from: Laptop
Fact-checking and judgesOllamaNot maintained2024
UT Austin and Bespoke Labs · USA
Checks whether each claim in an AI answer is supported by the source documents. The small versions are free; the larger 7B is in Ollama but non-commercial.
- Checking RAG bot answers against documents
- Finding unsupported claims in reports and summaries
- Automated quality control of AI answers
- Sizes
- 0.4B – 7B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Stability AI · UK
Small models from Stability AI: StableLM 2 (1.6B) knows 7 European languages, Stable Code (3B) completes code. No updates since 2024.
- A lightweight chatbot on an ordinary PC
- Code autocompletion in the editor
- A base for fine-tuning on your own task
- Sizes
- 1.6B – 12B
- Hardware
- from: Laptop
Text analysisRUNot maintained2020–2024
SberDevices (ai-forever) · Russia
Sber's Russian-language encoders trained on large Russian corpora. A base for classifiers, NER and semantic search in Russian.
- Classifying requests in Russian
- Extracting names, amounts and dates after fine-tuning
- Detecting review sentiment
- Sizes
- about 30M to 430M
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
Vectara · USA
A small model that checks whether an AI answer is grounded in the source text or made up. Runs on a CPU and works well as a filter in RAG systems.
- Checking knowledge base chatbot answers for fabrications
- Quality control of document summaries
- Comparing language models by their tendency to make errors
- Sizes
- 110M
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
BAAI (Beijing Academy of Artificial Intelligence) · China
Rerankers: they take passages found by search and reorder them by how well they actually match the question. v2-m3 is multilingual and lightweight, often paired with bge-m3.
- Refining search results before a chatbot answers
- Sorting knowledge base search results
- Selecting the most relevant clauses of contracts and policies
- Sizes
- 278M – 9B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
Patronus AI · USA
Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.
- Finding invented facts in AI assistant answers
- Checking that answers rest on the attached documents
- Filtering out answers before they go to a customer
- Sizes
- 8B and 70B
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
OpenChat (Tsinghua University) · China
Fine-tunes of Mistral 7B and Llama 3 8B using the C-RLFT method that caught up with ChatGPT-3.5 in 2023–2024 at just 7–8B. A lightweight general-purpose assistant for a modest server.
- Chat assistant on an inexpensive server
- Drafts of emails and replies
- Help with simple code
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
01.AI · China
Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.
- Chat assistant on a single GPU
- Analysis of long documents
- Classification and data extraction from text
- Sizes
- 6B – 34B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
RLHFlow · USA
An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data for model fine-tuning
- Scoring assistant answers across several attributes
- Sizes
- 8B
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
Hugging Face (H4) · USA
Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.
- Lightweight chat assistant
- Reference and starting point for your own fine-tuning
- Drafts of texts and replies
- Sizes
- 7B – 141B-A35B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2024
MyShell and MIT · USA
Instant voice cloning from a short sample with control over emotion and accent; V2 speaks several languages. Use only with the voice owner's consent.
- Voicing videos with the company narrator's voice
- Voice bot with a recognizable brand voice
- Transferring timbre onto existing speech synthesis
- Sizes
- under 1B
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2023–2024
KAIST and LG AI Research (prometheus-eval) · South Korea
An open judge model: it scores other models' answers against your criteria and explains the score. A replacement for paid models in the reviewer role.
- Scoring chatbot answers on your own scale
- Comparing two answer options
- Quality checks before launching an AI service
- Sizes
- 7B – 8x7B
- Hardware
- from: Laptop
Text to speechNot maintained2024
MyShell and MIT · USA
Lightweight multilingual speech synthesis that keeps up in real time on an ordinary CPU. English with accents, Spanish, French, Chinese, Japanese and Korean; no Russian.
- Voicing bot replies in foreign languages
- Voicing training materials
- Reading texts aloud on a server without a GPU
- Sizes
- small, runs in real time on a CPU
- Hardware
- from: Laptop
TextRUNot maintained2023–2024
Sber (ai-forever) · Russia
Sber's Russian text-to-text model, successor to ruT5 (2021). Small and fast: fine-tuned for summarizing, paraphrasing and fixing errors in Russian text; ready-made SAGE spell-checking versions exist.
- Fixing spelling mistakes and typos in Russian text
- Short summaries and paraphrasing
- Normalizing requests and inquiries before processing
- Sizes
- 95M – 1.7B
- Hardware
- from: Laptop
Search and RAGRUNot maintained2022–2024
Microsoft · USA
Proven models for semantic search. The multilingual versions work well with Russian and are still a reliable base for RAG.
- Search across a knowledge base and documents
- Finding answers for a chatbot (RAG)
- Finding similar requests and duplicates
- Sizes
- 33M – 7B
- Hardware
- from: Laptop
Search and RAGRUOllamaNot maintained2024
BAAI · China
A model for meaning-based search in about a hundred languages. The core of RAG: the bot finds the right part of a document before answering.
- Search across a document base
- RAG for a chatbot
- Finding similar requests and duplicates
- Sizes
- 568M
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Intel · USA
A fine-tuned Mistral 7B from Intel that showcased training and running on Intel CPUs and accelerators. Outdated; of interest as an example of optimisation for Intel hardware.
- A simple chat assistant
- Experiments with running on Intel hardware
- A base for fine-tuning
- Sizes
- 7B
- Hardware
- from: Laptop
RerankersNot maintained2023
NetEase Youdao · China
An embedding-plus-reranker pair for knowledge bases. The card lists English, Chinese, Japanese and Korean — Russian is not among the stated languages.
- Search across a knowledge base and reference materials
- Reordering retrieved passages
- Picking answers for a support chatbot
- Sizes
- about 280M
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2023
Google · USA
Google's translator for more than 400 languages under a permissive license. Russian is supported. A good substitute for NLLB when commercial use is needed.
- Translating documents and emails
- Translating catalogs and product descriptions
- Translating into CIS and Asian languages
- Sizes
- 3B – 10B
- Hardware
- from: Laptop
FinanceNot maintained2023
Fudan-DISC, Fudan University · China
A financial assistant made of several fine-tuned experts: advice, calculations, document reading and knowledge-base search. Not investment advice: decisions are made by a specialist.
- In-house advice on financial questions
- Reading financial documents and news
- Prompts for front-office staff
- Sizes
- 13B
- Hardware
- from: 1 GPU
Deepfake detectionNot maintained2022–2023
EURECOM · France
A step beyond AASIST: instead of raw audio it uses the wav2vec 2.0 speech encoder, which helps it hold up on unfamiliar synthesis methods. It errs in both directions - a human reviews the result.
- Spotting synthetic speech in calls
- Checking voice messages and recordings
- Fine-tuning for your own data and codecs
- Sizes
- about 0.3B (wav2vec 2.0 XLS-R encoder)
- Hardware
- from: Laptop
TextOllamaNot maintained2023
LMSYS (Berkeley and partners) · USA
One of the first open chat models (2023): LLaMA fine-tuned on user conversations with ChatGPT. A historical milestone; today it is weaker than any modern model of the same size.
- Experiments and team training
- Simple chat assistant for tests
- Comparison with newer models
- Sizes
- 7B – 33B
- Hardware
- from: Laptop
Voice: speakers and soundNot maintained2022–2023
Hendrik Schröter (University of Erlangen) · Germany
Lightweight real-time speech noise suppression that works even on a regular CPU and low-power devices. Removes hum, street and office noise while keeping the voice.
- Cleaning calls and voice messages of noise
- Preparing recordings before speech recognition
- Noise suppression for video calls
- Sizes
- about 2M
- Hardware
- from: Laptop
Text analysisRUNot maintained2022–2023
David Dale (cointegrated) and the community · Russia
Ready-made tiny rubert-tiny models for Russian text: detect rudeness and insults, sentiment and emotions. They run on a CPU in milliseconds.
- Filtering insults in Russian chats and comments
- Labeling reviews as positive, neutral or negative
- Spotting irritated customers in requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
TextNot maintained2022–2023
Google · USA
Compact input-output models trained to follow instructions. Still used as a cheap base for classification, extraction and short answers.
- Classification of requests and documents
- Extracting fields from text
- Short answers and summaries
- Sizes
- 80M – 20B
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2022–2023
Meta · USA
A translator for 200 languages, including rare and minor ones. Russian is supported. Strong language coverage, but the license prohibits commercial use.
- Translating texts between 200 languages
- Translating into rare languages where no other models exist
- Comparing quality when choosing a translator
- Sizes
- 600M to 3.3B (plus 54B MoE)
- Hardware
- from: Laptop
RerankersRUNot maintained2022
UKP Lab and the Sentence Transformers community · Germany
The most downloaded open rerankers: a tiny model reads a question-passage pair and scores how well they match. The multilingual mMARCO version covers Russian.
- Reordering knowledge base search results
- Selecting passages before a chatbot answers
- Finding duplicates among tickets and product cards
- Sizes
- about 4M – 120M
- Hardware
- from: Laptop
Text analysisRUNot maintained2019–2022
Meta · USA
A classic multilingual encoder for 100 languages, including Russian. The base of many sentiment, NER and embedding models, including BGE-M3.
- Detecting review sentiment in different languages
- Extracting names and organizations after fine-tuning
- Classifying requests
- Sizes
- 270M – 10.7B
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
Microsoft · USA
A time-tested encoder behind many classifiers and NER models (including GLiNER). The multilingual mDeBERTa-v3 understands Russian.
- Classifying review sentiment
- Entity extraction after fine-tuning
- Checking whether a conclusion follows from a text
- Sizes
- 70M – 435M
- Hardware
- from: Laptop
Text analysisRUNot maintained2021
David Dale (cointegrated) · Russia
A very small Russian-English BERT that runs fast on a regular CPU. Ready-made fine-tuned versions exist for sentiment, toxicity and emotions.
- Detecting review sentiment
- Filtering rude chat messages
- Fast classification of requests
- Sizes
- 12M – 29M
- Hardware
- from: Laptop
Deepfake detectionNot maintained2021
NAVER Clova AI Research and EURECOM · South Korea
The baseline open model against voice spoofing: it listens to the raw recording and tells a live person from synthesis or a replay. It errs in both directions - its output is a reason for a human to check, not proof.
- Voice check during phone authentication
- Filtering replays and synthesis in a voice menu
- A baseline when comparing voice detectors
- Sizes
- weight files of 0.4 and 1.3 MB
- Hardware
- from: Laptop
TranslationRUGGUFNot maintained2020–2021
Meta · USA
An early Meta translator that translates directly between 100 languages, without English in the middle. Russian is supported. Old, but light and permissively licensed.
- Translation between any pair of 100 languages
- Quick draft translation on modest hardware
- Base for fine-tuning to your subject area
- Sizes
- 418M – 12B
- Hardware
- from: Laptop