Large models know English and Russian, and past that the map is mostly blank: for the languages of Russia there is very little ready to use, and what exists was built by small teams and enthusiasts. This collection gathers those models: speech recognition, synthesis, translation, text analysis. Quality varies, so test on your own recordings, but this is the only way to build a service in a native language without waiting for the big labs to get there.
Russian to Tuvan translation that runs on your own machine: ready GGUF builds, a LoRA adapter and a standalone Windows app. The author warns that names, numbers and important text should be double-checked with a native speaker.
AI Laboratory of the North-Eastern Federal University · Russia
The most complete set of models for the Sakha (Yakut) language in one place: language models of several sizes, translation to and from Russian, embeddings and text recognition from images. Quality varies, so test it on your own data.
A small toolkit for Bashkir text: a fill-mask language model, a Bashkir language detector, fastText embeddings, a Bashkir-Russian pair scorer and a letter restorer. All published in autumn 2026, so check quality on your own data.
Aigiz Kunafin (AigizK), member of the SLONE community · Russia
Years of work by one enthusiast around the Bashkir Common Voice project: Whisper, wav2vec2, w2v-BERT and GigaAM fine-tuned for Bashkir, Tatar and Mari. Quality differs between models, so test them on your own recordings.
Transcribing calls and meetings in Bashkir and Tatar
SLONE (David Dale, Aigiz Kunafin and others) · not disclosed
A community that extends NLLB and mBART to low-resource languages. Its models cover Erzya, Tuvan, Bashkir, Tatar, Chuvash, Buryat, Mari, Khakas and Karachay-Balkar. Quality differs by language, so test it on your own texts.
Translating from Russian into a national language and back
Preparing bilingual materials and signage
Aligning parallel texts
Sizes
42M for the compact encoders, 620M – 758M for the translators
A recent set of models for Tatar: morphological analysis on several base models, question answering about Tatar and Russian place names, and Mistral and GPT-2 fine-tunes for Tatar.
Morphological analysis of Tatar texts
Answering questions about place names
Generating text in Tatar
Sizes
178M for the morphology models, 7B for the Mistral fine-tune
About two dozen Whisper, wav2vec2 and WavLM fine-tunes for Karelian, including recordings where the speaker switches from Karelian to Russian. The model cards are empty, so quality must be checked on your own recordings.
Transcribing field and archive recordings in Karelian
A multilingual encoder trained on more than 1800 languages. The list includes Tatar, Bashkir, Chuvash, Udmurt, Buryat, Komi, Ingush and other languages of Russia. A base for classifiers and search over such texts.
Classifying requests and documents in national languages
A set of LaBSE fine-tunes that turn sentences into vectors: Buryat, Mari, Udmurt, Ingush, Chuvash, Sakha and Kalmyk. Trained on pairs with Russian; quality differs by language.
Aligning parallel texts in Russian and a national language
Semantic search across bilingual archives
Building training sets for translation
Sizes
LaBSE fine-tunes; parameter count is not stated on the cards
Chuvash speech synthesis: an F5-TTS fine-tune on top of the Russian version, trained on 24 hours of Common Voice recordings. It can copy a voice from a sample — only with the voice owner’s consent. The author calls it an experiment, so check quality on your own texts.
A project around the Udmurt language: encoders based on mBERT, rubert-tiny2 and Glot500, part-of-speech taggers, a typo correction model and Udmurt-English embeddings.
Detects which language a text is written in: the third version covers more than two thousand labels, including Tatar, Bashkir, Chuvash, Udmurt, Mari, Erzya, Komi and other languages of Russia.
Sorting mixed text archives by language
Filtering out noise and foreign languages before training
Routing requests to the right operator or model
Sizes
a FastText model; parameter count is not stated on the model card
Tweeties (Francois Remy, Ghent University, and co-authors) · Belgium
A language model for Tatar: Mistral 7B converted to a Tatar tokenizer, plus a version for translating between Tatar and a dozen other languages, including Russian. Check quality on your own texts.
Meta speech models covering more than a thousand languages. The supported list includes Bashkir, Tatar, Chuvash, Sakha, Ossetian, Chechen, Avar, Udmurt, Mari, Erzya, Moksha, Kalmyk and Karelian. Quality varies a lot by language, so test it on your own recordings.
Separate mGPT versions, one per language: Bashkir, Buryat, Kalmyk, Mari, Ossetian, Tatar, Tuvan, Chuvash and Sakha. Each card also lists Russian and English.
Generating and continuing text in a national language
A base for fine-tuning to your own task
Experiments with rare languages without training from scratch