These models hold a spoken conversation: they listen and reply with speech without a separate recognition, text and synthesis pipeline. They power phone voice bots and in-app assistants. Check response latency, support for your languages, the license and GPU requirements.
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.
Russian-language employee assistant on your own server
Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
High-load agents
Analysis of documents with diagrams and screenshots
Models that listen to speech, sounds and music and answer questions about them. Audio Flamingo Next handles recordings up to 30 minutes. Research use only.
A voice assistant that listens and speaks at the same time, with no delay for recognition and synthesis. Hibiki does simultaneous speech-to-speech translation between several European languages.
A voice conversation partner based on Moshi that listens and speaks at the same time and can be interrupted. The role is set by text, the voice by a sample recording. English only.
Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.
Russian-language assistant on your own PC or server
A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.
Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.