Open-source models for voice assistants

These models hold a spoken conversation: they listen and reply with speech without a separate recognition, text and synthesis pipeline. They power phone voice bots and in-app assistants. Check response latency, support for your languages, the license and GPU requirements.

15 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
TextGGUF2025–2026

MiMo

Xiaomi · China

Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.

  • Logic and calculation tasks
  • Agents with tools
  • Help for developers
Sizes
7B – 1,02T-A42B
Hardware
from: Laptop
Commercial use allowedDetails
TextRU2024–2026

GigaChat

Sber · Russia

Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.

  • Russian-language employee assistant on your own server
  • Customer replies and request handling in Russian
  • Working with contracts and internal policies
Sizes
10B-A1.8B – 702B-A36B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2026

Samsone

Samsung · South Korea

Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.

  • Describing an audio recording in words
  • Answering questions about a sound
  • Identifying the type of sound and the setting
Sizes
99M – 356M
Hardware
from: Laptop
Non-commercial onlyDetails
TextOllama2024–2026

NVIDIA Nemotron

NVIDIA · USA

NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).

  • Agents with tool calling
  • Reasoning and calculation tasks
  • Answers based on long documents
Sizes
4B – 550B-A55B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2025–2026

Liquid LFM

Liquid AI · USA

Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.

  • Offline assistant on a laptop or phone
  • Data extraction from documents
  • Tool calling in apps
Sizes
230M – 24B-A2B
Hardware
from: Laptop
Commercial use with conditionsDetails
TextGGUF2025–2026

LongCat

Meituan · China

Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.

  • Agents for orders and service processes
  • Corporate assistant
  • Analysis of long documents
Sizes
1B – 1.6T-A48B
Hardware
from: Laptop
Commercial use allowedDetails
TextGGUF2025–2026

Step

StepFun · China

StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.

  • High-load agents
  • Analysis of documents with diagrams and screenshots
  • Help for developers
Sizes
8B – 321B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice assistants2024–2026

Audio Flamingo

NVIDIA · USA

Models that listen to speech, sounds and music and answer questions about them. Audio Flamingo Next handles recordings up to 30 minutes. Research use only.

  • Detailed descriptions of audio recordings
  • Questions and answers about a long recording
  • Tagging music and sounds
Sizes
0.5B – 8B
Hardware
from: Laptop
Non-commercial onlyDetails
Voice assistants2024–2026

Moshi / Hibiki

Kyutai · France

A voice assistant that listens and speaks at the same time, with no delay for recognition and synthesis. Hibiki does simultaneous speech-to-speech translation between several European languages.

  • Real-time voice conversation partner
  • Simultaneous speech translation
  • Zero-latency voice interfaces
Sizes
2B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Voice assistantsGGUF2025–2026

MiniCPM-o

OpenBMB (ModelBest, Tsinghua University) · China

A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.

  • Voice assistant on your own server
  • Analyzing videos and documents
  • Voice answers about a camera image
Sizes
8B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistantsGGUF2026

PersonaPlex

NVIDIA · USA

A voice conversation partner based on Moshi that listens and speaks at the same time and can be interrupted. The role is set by text, the voice by a sample recording. English only.

  • A voice assistant with a set role
  • A conversation simulator for staff training
  • Voice interfaces without delay
Sizes
7B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
TextRUGGUF2024–2025

Vikhr

Vikhr Models · Russia

Russian-language fine-tunes of open models (Mistral, Qwen, Llama) by the independent Vikhr team, with compact versions for a regular PC. Borealis is an audio model for recognizing and understanding Russian speech.

  • Russian-language assistant on your own PC or server
  • Knowledge-base answers (RAG)
  • Texts and emails in Russian
Sizes
0.5B – 24B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistantsRUGGUF2025

Qwen Omni

Alibaba (Qwen) · China

Models that understand text, images, audio and video and reply by voice in real time. Qwen3-Omni speaks 10 languages, including Russian.

  • Voice assistant for customers
  • Analyzing calls and videos
  • Voice answers about documents and images
Sizes
3B – 30B-A3B
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2025

Kimi-Audio

Moonshot AI · China

A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.

  • Speech recognition
  • Detecting emotions and sound events
  • Speech-to-speech voice dialogue
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Voice assistantsRUNot maintained2023

SeamlessM4T / Seamless

Meta · USA

Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.

  • Speech-to-speech translation
  • Translating and transcribing recordings
  • Streaming translation
Sizes
281M – 2.3B
Hardware
from: Laptop
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment