Open AI models for the languages of Russia

Large models know English and Russian, and past that the map is mostly blank: for the languages of Russia there is very little ready to use, and what exists was built by small teams and enthusiasts. This collection gathers those models: speech recognition, synthesis, translation, text analysis. Quality varies, so test on your own recordings, but this is the only way to build a service in a native language without waiting for the big labs to get there.

15 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
TranslationRUGGUF2026

TuvanGemma

Mergen Kungaa · not disclosed

Russian to Tuvan translation that runs on your own machine: ready GGUF builds, a LoRA adapter and a standalone Windows app. The author warns that names, numbers and important text should be double-checked with a native speaker.

  • Translating documents and notices into Tuvan
  • Translating Tuvan texts into Russian
  • Draft translation without sending data outside
Sizes
4B and 12B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRU2025–2026

Лаборатория ИИ СВФУ: якутский язык

AI Laboratory of the North-Eastern Federal University · Russia

The most complete set of models for the Sakha (Yakut) language in one place: language models of several sizes, translation to and from Russian, embeddings and text recognition from images. Quality varies, so test it on your own data.

  • Answering and writing in the Sakha language
  • Translating documents and materials
  • Semantic search over Sakha texts
Sizes
1.1B – 27B
Hardware
from: Laptop
Commercial use with conditionsDetails
Text analysisRU2026

BashkirRoBERTa и набор моделей для башкирского

Tvoy_Tezka (failed09) · not disclosed

A small toolkit for Bashkir text: a fill-mask language model, a Bashkir language detector, fastText embeddings, a Bashkir-Russian pair scorer and a letter restorer. All published in autumn 2026, so check quality on your own data.

  • Spellchecking and word suggestion in Bashkir
  • Separating Bashkir texts from other languages
  • Restoring language-specific letters in text
Sizes
50M for BashkirRoBERTa
Hardware
from: Laptop
Commercial use with conditionsDetails
Speech to textRU2022–2026

Распознавание речи для башкирского, татарского и марийского (AigizK)

Aigiz Kunafin (AigizK), member of the SLONE community · Russia

Years of work by one enthusiast around the Bashkir Common Voice project: Whisper, wav2vec2, w2v-BERT and GigaAM fine-tuned for Bashkir, Tatar and Mari. Quality differs between models, so test them on your own recordings.

  • Transcribing calls and meetings in Bashkir and Tatar
  • Subtitles for video and radio broadcasts
  • Turning archive recordings into text
Sizes
from 220M to 1B across the different models
Hardware
from: Laptop
Commercial use with conditionsDetails
TranslationRU2022–2026

SLONE: перевод для языков народов России

SLONE (David Dale, Aigiz Kunafin and others) · not disclosed

A community that extends NLLB and mBART to low-resource languages. Its models cover Erzya, Tuvan, Bashkir, Tatar, Chuvash, Buryat, Mari, Khakas and Karachay-Balkar. Quality differs by language, so test it on your own texts.

  • Translating from Russian into a national language and back
  • Preparing bilingual materials and signage
  • Aligning parallel texts
Sizes
42M for the compact encoders, 620M – 758M for the translators
Hardware
from: Laptop
Commercial use with conditionsDetails
Text analysisRU2026

Tatar NLP Community

Tatar NLP Community · not disclosed

A recent set of models for Tatar: morphological analysis on several base models, question answering about Tatar and Russian place names, and Mistral and GPT-2 fine-tunes for Tatar.

  • Morphological analysis of Tatar texts
  • Answering questions about place names
  • Generating text in Tatar
Sizes
178M for the morphology models, 7B for the Mistral fine-tune
Hardware
from: Laptop
Commercial use with conditionsDetails
Speech to textRU2024–2025

Распознавание карельской речи (Mihaj)

Mihaj Dolgushin · not disclosed

About two dozen Whisper, wav2vec2 and WavLM fine-tunes for Karelian, including recordings where the speaker switches from Karelian to Russian. The model cards are empty, so quality must be checked on your own recordings.

  • Transcribing field and archive recordings in Karelian
  • Subtitles for Karelian-language material
  • Handling recordings that mix Karelian and Russian
Sizes
from 300M to 1B across the different models
Hardware
from: Laptop
Commercial use with conditionsDetails
Text analysisRUGGUF2025

mmBERT

Johns Hopkins University (JHU CLSP) · USA

A multilingual encoder trained on more than 1800 languages. The list includes Tatar, Bashkir, Chuvash, Udmurt, Buryat, Komi, Ingush and other languages of Russia. A base for classifiers and search over such texts.

  • Classifying requests and documents in national languages
  • Semantic search over texts in rare languages
  • Extracting names, dates and titles from text
Sizes
140M and 307M
Hardware
from: Laptop
Commercial use allowedDetails
Search and RAGRU2023–2025

LaBSE для языков народов России (lingtrain)

lingtrain · not disclosed

A set of LaBSE fine-tunes that turn sentences into vectors: Buryat, Mari, Udmurt, Ingush, Chuvash, Sakha and Kalmyk. Trained on pairs with Russian; quality differs by language.

  • Aligning parallel texts in Russian and a national language
  • Semantic search across bilingual archives
  • Building training sets for translation
Sizes
LaBSE fine-tunes; parameter count is not stated on the cards
Hardware
from: Laptop
Commercial use with conditionsDetails
Text to speechRU2025

F5-TTS для чувашского языка

Misha Yakovlev (Misha24-10) · not disclosed

Chuvash speech synthesis: an F5-TTS fine-tune on top of the Russian version, trained on 24 hours of Common Voice recordings. It can copy a voice from a sample — only with the voice owner’s consent. The author calls it an experiment, so check quality on your own texts.

  • Reading Chuvash texts aloud
  • Voice prompts and announcements
  • Audio versions of learning material
Sizes
parameter count is not stated on the model card
Hardware
from: Laptop
Non-commercial onlyDetails
Text analysisRUNot maintained2024

Zerpal: модели для удмуртского языка

udmurtNLP · not disclosed

A project around the Udmurt language: encoders based on mBERT, rubert-tiny2 and Glot500, part-of-speech taggers, a typo correction model and Udmurt-English embeddings.

  • Parsing Udmurt texts by part of speech
  • Correcting typos and recognition errors
  • Semantic search over Udmurt material
Sizes
33M – 187M
Hardware
from: Laptop
Commercial use with conditionsDetails
Text analysisRUNot maintained2023–2024

GlotLID

CIS, LMU Munich · Germany

Detects which language a text is written in: the third version covers more than two thousand labels, including Tatar, Bashkir, Chuvash, Udmurt, Mari, Erzya, Komi and other languages of Russia.

  • Sorting mixed text archives by language
  • Filtering out noise and foreign languages before training
  • Routing requests to the right operator or model
Sizes
a FastText model; parameter count is not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
TextRUGGUFNot maintained2024

Tweety Tatar

Tweeties (Francois Remy, Ghent University, and co-authors) · Belgium

A language model for Tatar: Mistral 7B converted to a Tatar tokenizer, plus a version for translating between Tatar and a dozen other languages, including Russian. Check quality on your own texts.

  • Generating and continuing Tatar text
  • Translating between Tatar and Russian
  • A base for fine-tuning to your own task
Sizes
7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Speech to textRUNot maintained2023

MMS (Massively Multilingual Speech)

Meta AI · USA

Meta speech models covering more than a thousand languages. The supported list includes Bashkir, Tatar, Chuvash, Sakha, Ossetian, Chechen, Avar, Udmurt, Mari, Erzya, Moksha, Kalmyk and Karelian. Quality varies a lot by language, so test it on your own recordings.

  • Transcribing audio and video in rare languages
  • Reading text aloud in a national language
  • Detecting which language a recording is in
Sizes
300M – 1B
Hardware
from: Laptop
Non-commercial onlyDetails
TextRUGGUFNot maintained2023

mGPT 1.3B для языков народов России

SberDevices (ai-forever) · Russia

Separate mGPT versions, one per language: Bashkir, Buryat, Kalmyk, Mari, Ossetian, Tatar, Tuvan, Chuvash and Sakha. Each card also lists Russian and English.

  • Generating and continuing text in a national language
  • A base for fine-tuning to your own task
  • Experiments with rare languages without training from scratch
Sizes
1.3B
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment