Open-weight models released since 2022: text, code, images, video, speech, 3D. For each one — what it does, where it is used, what hardware it needs and whether commercial use is allowed. I can deploy any of them on your server and fine-tune it for your task.
FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China
A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.
Microsoft speech models: long multi-voice dialogue synthesis, fast synthesis for live conversation, and recognition of long recordings split by speaker, including in Russian.
DeepSeek's flagship line: from the first 7B/67B to V4-Pro with 1.6 trillion parameters. Closed-model quality under an open MIT license; V4-Flash-Vision-Exp and V4.1-Flash understand images, context up to 1M tokens.
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
Sber open models with strong Russian language support and local context, from 10B-A1.8B to 702B, all MIT. GigaChat3.1-Audio handles recordings up to two hours; GFusion is a fast diffusion text version.
Russian-language employee assistant on your own server
Yandex models trained from scratch with a focus on the Russian language and Russian context. The new AliceAI-Foundation 80B-A3B (Apache 2.0) is a base model only, with no instruct version: you fine-tune it for your own tasks. The efficient AliceAI-T5 35B-A0.6B is also available.
Multilingual models from Cohere's research arm, covering 23 to 100+ languages. Tiny Aya (2026, 3.3B) runs on a regular PC, but for non-commercial use only.
Translation and correspondence in less common languages
A ready-made model for tables: it takes example rows and immediately predicts for new ones, without lengthy training or tuning. Only v2 is free for business; newer versions are non-commercial.
Amazon's tabular model built into AutoGluon: classification and regression from examples with brief fine-tuning. Mitra-v2 handles more rows and columns.
A tabular model from a Canadian bank's AI lab, trained on real tables rather than only synthetic ones. Version 1.2 Turbo made computation orders of magnitude faster.
The most widely used open toolkit for face detection and recognition. Many identity-preserving image generators are built on it. The pretrained weights are non-commercial.
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
Renmin University of China (GSAI) and Ant Group (inclusionAI) · China
Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.
The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.
An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.
A research testbed for driver assistance systems
Studying driver attention monitoring with an in-cabin camera
Tencent models based on Qwen3.5 for searching scans and PDFs as images. According to the model card, among the top of the ViDoRe leaderboard at release.
A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.
Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.
Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.
An image watermark for arbitrary resolutions built for the Content Authenticity Initiative: it can both apply a mark and remove one. The detector errs in both directions - a human reviews the output.
Marking images on the way out of your own pipeline
Sber embeddings built for Russian: according to the developers, among the best on Russian-language search benchmarks. FRIDA is compact, Giga-Embeddings is more powerful.
Very small and fast speech recognition models for phones, tablets and embedded devices. Version 2 streams, producing text while the person is still speaking.
One of the oldest Chinese open lines: from ChatGLM-6B to GLM-5.3. Strong at agentic tasks and programming; GLM-5.3-Flash understands images and is released under MIT.
Tencent language models: from small 0.5B–7B to Hy4-preview with 770 billion parameters. Since 2026 the line has been renamed Hy, and new versions are released under Apache 2.0.
An Ant Group family: Ling for standard models, Ring for reasoning ones. There are trillion-parameter flagships and the efficient Ling-3.0-tiny, which needs only 1.3 billion active parameters.
Business models: document search with source citations, tool calling, many languages. Command A+ (2026) was the first under Apache 2.0, followed by the North line: code, translation and compact vision.
Knowledge-base answers with source citations
Agents that work with internal systems
Translation and correspondence in different languages
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
Models with a new architecture for on-device use: fast on a regular CPU and on phones. Versions for data extraction, RAG and tools, plus LFM2.5-VL for images and voice LFM2.5-Audio.
Finds the entities you need in text without training: just list what to look for (name, amount, date). GLiNER2 also classifies text. Multilingual versions understand Russian.
Extracting names, amounts and dates from emails and contracts
A new lightweight document parsing model that led the OmniDocBench v1.6 benchmark at release. Handles pages photographed on a phone and crumpled pages well. Languages on the card: Chinese, English, Japanese.
Recognising invoices and delivery notes photographed on a phone
Community: Ultimate Vocal Remover (Anjok07), ZFTurbo, MVSep · International community
A large open collection of models for separating vocals from music and noise: MDX-Net, BS-RoFormer, Mel-RoFormer, SCNet. The quality leaders for vocals among open solutions.
Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.
An Ant Group family for finding elements on screen and completing tasks in phone and computer interfaces. UI-Venus-2 was specifically trained to refuse dangerous actions.
Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.
A fast climate model emulator: simulates the atmosphere years and decades ahead on a single GPU. Coupled with an ocean model (SamudrACE) for long-term scenarios.
Decades-long climate scenarios to assess long-term asset risks
Large-scale what-if runs on temperature and precipitation
Preparing data for crop yield and energy demand models
Google DeepMind's family of global weather models: GraphCast (10-day forecast), GenCast (probabilistic ensemble) and WeatherNext 2 with cyclone forecasting. Since August 2026 the weights are cleared for commercial use.
Medium-range weather forecasts for planning shifts, voyages and deliveries
Probabilistic assessment of extreme weather for insurance portfolios
Tropical cyclone track forecasts for marine and port operations
AlQuraishi Lab (Columbia University) and the OpenFold consortium · USA
A fully open reproduction of AlphaFold 2 and then AlphaFold 3 under Apache 2.0, with training data. OpenFold3 predicts complexes of proteins, nucleic acids and ligands.
Predicting structures of proteins and ligand complexes
Fine-tuning on the company's own data (training code is open)
An in-house structural analysis service without sending data outside
Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.
Auto-labeling camera recordings to train your own driver assistance systems
Analyzing complex road scenes with text explanations
Testing autopilot systems in simulation on rare scenarios
An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.
A perception and planning prototype for autonomous vehicles on closed sites
Answering questions about camera recordings when reviewing incidents
A lightweight detector of generated images, trained on 2.7M samples from nearly 5000 different generators. It errs in both directions: the result is a reason for a human to check, not proof.
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
Checking images from new, unfamiliar generators
A baseline when comparing detectors
Fast rollout of a check without training a large model
European models focused on speed. Mixtral was one of the first open mixture-of-experts models; there are versions for images (Pixtral, Medium 3.5), Lean proofs and moderation (Shieldstral).
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
Short promo videos
Animating product photos
Videos for social media
Sizes
1,3B – 14B
Hardware
from: 1 GPU
Commercial use allowedDetails
No models match these filters. Reset them or describe your task — I will pick one for you.