MAGI
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
- Long videos with continuation
- Video with sound
- Animating images
- Sizes
- 4.5B – 114B-A6B
- Hardware
- from: 1 GPU
A single GPU is the most common business setup: a server with a 16, 24, 48 or 80 GB card covers almost everything except the largest models. This is where live chat, image generation and real-time document processing become practical.
Video generated chunk by chunk in sequence, so a clip can be extended indefinitely. MAGI-2 produces video with sound.
NVIDIA's lightweight, fast video model. Produces 720p clips on a single GPU; a 4-step version enables quick generation.
Text-to-video and image-to-video; the small version runs on a gaming GPU. After 2.2 only applied models are open: editing (VACE), audio-driven talking characters (S2V), dancing to music (Dancer).
A fast video model; with LTX-2 it generates video with sound and speech in one go. Camera and pose control, lightweight versions available.
Animates a character from an image using motion from another video, including complex turns and multiple characters. SCAIL-2 works without an intermediate skeleton and can replace a character in a clip.
An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.
Video models for cinematic scenes with people. Can make videos of unlimited length, extend videos and create talking characters from audio.
Generates video with sound in one pass: lip-synced speech, effects and ambience. A 32B-parameter MoE architecture, with 360p and 720p versions.
Tencent's video model, one of the first open ones on par with closed services. Version 1.5 is lighter (8.3B) and runs on consumer GPUs.
Generates video together with sound and speech from text or an image: two branches (video based on Wan 2.2 and a 5B audio branch) run in sync. Needs 24–32 GB of GPU memory.
A 13.6B video model: from text, from an image and video continuation. Keeps quality on clips several minutes long.
Turns a single image into a controllable game-scene video: the camera moves on keyboard commands. Minimum 24 GB of GPU memory, 80 GB recommended.
Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
A community model retrained from FLUX.1-schnell with a simplified architecture. No style censorship; popular as a base for fine-tuning.
A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.
Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.
Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.
Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.
Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.
Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.
A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.
A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.
An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.
Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.
An open Meta model for agents on affordable hardware: distilled from the closed Muse Spark, understands text and images, trained on 100+ languages.
Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.
Models from Korea's Upstage. Solar Open 2 is built for office document work: 250 billion parameters, 15 billion active; languages are English, Korean and Japanese.
StepFun MoE models built for fast, low-cost work: with 196 billion parameters, Step-3.5/3.7-Flash use about 11 billion per token. Compact Step3-VL-10B for images and voice Step-Audio 2 mini are available.
Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.
OpenAI's first open models since GPT-2. Reasoning and tool calling; the smaller version fits on a single GPU.
An open ByteDance 36B model with up to 512K tokens of context and an adjustable thinking budget. Fits on a single powerful GPU.
Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.
Sber's 13-billion-parameter base Russian model; GigaChat grew out of its fine-tuned version. Continues texts in Russian and English, context only 2048 tokens; today useful as a base for narrow fine-tuning.
Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.
Generates whole 3D worlds and scenes from text or an image that you can walk through. The second version builds a scene from video and photos.
One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.
A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.
Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.
Generates 3D human motion animation from a text description: the skeletal animation is ready for 3D editors and game engines. Understands English and Chinese.
Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.
Tencent's open 3D line: shape and texture from an image, at the level of paid services. Omni adds control of pose and shape, Part splits a model into parts.
Stability AI models that turn a single photo into a textured 3D model in about a second. SPAR3D lets you adjust the shape through a point cloud.
Builds a 3D mesh from a single image in about 10 seconds: first it draws the object from several angles, then assembles the model from them.
A rare open try-on model with a commercial license: mask-free, accepts a photo of the item on a model or a flat lay. Weights are about 2 GB.
Try-on beyond clothing: glasses, earrings, bags, hats, watches and other accessories. Works without a mask. Needs a GPU with 28 GB or more.
Tries on several items at once: top, bottom, shoes, bag. Faster than earlier models thanks to caching. Non-commercial license.
An all-round FLUX-based apparel toolkit: try-on, generating a model wearing a given item, and "taking off" an item from a person into a separate product photo.
Meta's model for virtual try-on and changing a person's pose in a photo. Carefully transfers fine fabric details and lettering. MIT license, but the training data is non-commercial.
Tencent's transformer-based try-on: more accurately reproduces fabric texture, fine prints and garment length. Non-commercial license.
One of the best-known open virtual try-on models: moves a garment from a product photo onto a photo of a person, keeping prints and logos well. Non-commercial license.
An early popular open try-on model: one version for upper-body garments, another for full-length outfits. Non-commercial license.
A research try-on model from CVPR 2024, one of the first built on Stable Diffusion. Now mostly used as a comparison baseline.
A robot control model trained mostly on synthetic data from a world model. It reduces spending on collecting data from real robots.
"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.
A robot control model from Ant Group trained on a large volume of data from real robots. Version 2.0 works with different types of robot arms.
Open robot control models from Xiaomi. Robotics-1 is designed for household and kitchen tasks, U0 combines scene understanding and action.
A fully open robot control model that first "reasons" about space and trajectory, then acts. Its reasoning can be checked.
NVIDIA's foundation model for humanoid robots and robot arms: it sees, understands a command and outputs movements. Built into the Isaac ecosystem.
Robot control models from Physical Intelligence: folding laundry, tidying up, handling objects. π0.5 copes better in unfamiliar settings.
The first large open vision-language-action model: a robot arm carries out commands like "put the apple in the bowl". OFT makes it several times faster.
Audio-driven talking people built on LongCat-Video. Version 1.5 is production-ready: stable long videos in Chinese and English.
A series of audio-driven talking portraits: from short clips to hour-long 4K videos. Hallo-Live is built for real-time use.
Generates video of a talking person with sound in one go: a single transformer processes text, video and audio. Speech in 7 languages; Russian is not among them. Fast distilled versions are available.
Ant Group's talking avatars: the face and, from V2, hand gestures. V3-Flash produces video in 8 steps and fits into 12 GB of GPU memory.
A real-time streaming avatar of unlimited length. Suits live broadcasts and dialogue, but needs powerful server hardware.
Dubbing and talking characters built on Wan: MultiTalk handles dialogue between several people, InfiniteTalk re-dubs videos of any length with facial and body motion.
Talking characters built on HunyuanVideo: conveys emotions from the voice, handles several characters and different styles.
Tencent models based on Qwen3.5 for searching scans and PDFs as images. According to the model card, among the top of the ViDoRe leaderboard at release.
Search across PDF pages and scans as images. The cards list English, Italian, French, German and Spanish — Russian is not among them.
A small model for searching document pages as images, from the team behind a popular RAG framework. The card lists English, Italian, French, German and Spanish.
One vector for text, for an image and for a text-image pair: a single model can find a product by photo, a document page by question and an image by description. The card lists English and Chinese.
Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.
A reranker for document pages as images: after a visual search it reorders the found pages by how well they answer the question. The card does not state the languages.
Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.
An image + video + text model focused on long videos and precise linking of events to timestamps. A Realtime version handles live video streams.
A fine-tuned Qwen3-VL-8B reads resume pages as images and returns a 23-field JSON record. The author states plainly that the model is not meant for automated decisions about candidates; a human decides.
Avito's Russian-language model that understands images: describes photos, answers questions about an image, reads text on it. Based on Qwen2.5-VL, faster in Russian than the original.
An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.
Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.
A fast climate model emulator: simulates the atmosphere years and decades ahead on a single GPU. Coupled with an ocean model (SamudrACE) for long-term scenarios.
Google DeepMind's family of global weather models: GraphCast (10-day forecast), GenCast (probabilistic ensemble) and WeatherNext 2 with cyclone forecasting. Since August 2026 the weights are cleared for commercial use.
A foundation model of Earth's atmosphere: global weather forecasts, plus separate versions for air quality and ocean waves. Computes a forecast in seconds instead of hours on a supercomputer.
NVIDIA's set of weather and climate models: global FourCastNet forecasts, downscaling to kilometers (CorrDiff), regional storm forecasts (StormCast), climate generation (cBottle, Atlas). Run via Earth2Studio.
ECMWF's weather neural network running operationally: a 15-day forecast four times a day, an ensemble version with 51 scenarios, and since version 2, ocean waves.
Models for agentic development: they build their own plan and scaffolding for a task and execute it in the terminal. Fine-tuned from Qwen 3.5 and Gemma 4; work with Claude Code, OpenHands and similar tools.
Kuaishou models for agentic development, trained to solve real tasks in repositories. KAT-Coder-V2.5-Dev (35B, 3B active) is the open version of their closed flagship.
Models for agentic programming: they edit code in a repository on their own. The small XS runs on a Mac with 36 GB of memory; S 2.1 has a 1M-token context.
Mistral models for agentic development: they read the repository, edit files and run commands on their own. The 24B version fits on a single GPU.
Generates full songs with vocals and accompaniment from lyrics and a style description: English, Chinese, Japanese, Korean.
An open model for generating songs with vocals in Chinese, English, Japanese, Korean and Spanish, plus a codec and a lyrics transcription model.
Generates and edits audio for video, text or audio, first "reasoning" about the scene with a multimodal model. PrismAudio is the next version for video-to-audio.
Generates studio-quality (48 kHz) audio for video from the picture and a text prompt: footsteps, impacts, ambience, in sync with the action on screen.
ByteDance's video and photo restoration and upscaling. SeedVR2 does it in a single step, so it is noticeably faster than similar models. Commercial-friendly license.
A FLUX add-on for upscaling small and blurry images with detail reconstruction. Popular, but under the non-commercial FLUX dev license.
Powerful SDXL-based restoration of badly damaged photos: it recreates details rather than just upscaling. Hardware-hungry; non-commercial license.
One of the first Stable Diffusion-based photo upscalers: restores realistic details. Non-commercial license.
Fully open desktop agents: weights, data and training code. They work on Windows, macOS and Linux; the latest Qwen-CUA controls a computer with ordinary clicks and keystrokes.
Meituan's computer-control agent, trained on a large number of simulated tasks in desktop software. It outputs clicks and keyboard input.
One of the first open models for controlling an interface from a screenshot; its successor, AutoGLM-Phone, works in Android smartphone apps.
An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.
A fully open reproduction of AlphaFold 2 and then AlphaFold 3 under Apache 2.0, with training data. OpenFold3 predicts complexes of proteins, nucleic acids and ligands.
DNA language models with context up to a million nucleotides: they assess the impact of mutations, annotate genomes and generate sequences. Evo 2 is trained on genomes from all domains of life.
An open MIT-licensed alternative to AlphaFold 3: predicts structures of protein, DNA and small-molecule complexes; Boltz-2 estimates binding strength, BoltzGen designs new binding proteins.
The reference model for the structure of biomolecules and their complexes. Weights are provided for non-commercial research only; companies need commercial access via Google Cloud or open alternatives (Boltz, OpenFold3).
Models that spell out their reasoning before answering financial questions with numbers and tables. Not investment advice: decisions are made by a specialist.
A large Chinese model family for the financial industry: advice, document reading and long texts up to 8k-16k. Not investment advice: decisions are made by a specialist.
An open set of lightweight add-ons for ordinary language models that work with financial texts and news. Not investment advice: decisions are made by a specialist.
A financial assistant made of several fine-tuned experts: advice, calculations, document reading and knowledge-base search. Not investment advice: decisions are made by a specialist.
A large medical MoE model based on Ling-flash-2.0: 100B parameters with 6B active, so it answers quickly. Does not replace a doctor; decisions are made by a specialist.
An early medical model on Llama 2 70B, trained on dialogues based on medical texts. Now mainly of research interest. Does not replace a doctor; decisions are made by a specialist.
An early general-purpose radiology model: understands 2D and 3D images (CT, MRI) together with text. More of a research base. Does not replace a doctor; decisions are made by a specialist.
Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.
A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.
Turns a scanned page into tagged text with block coordinates, or into markdown. Handy as the first step before parsing a resume. A human makes the decision about a candidate; automatic screening without review must not be used.
Google's large tabular model: classification and regression from examples without training, with numeric and categorical columns. Weights are for non-commercial use only.
A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.
A foundation model for predictions on tables: it classifies and forecasts from a handful of examples, with no separate task-specific training.
FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.
Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.
Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
Finds traces of editing and shows on a map which regions of an image look altered: suitable for scans of contracts, certificates and photos of documents. It errs in both directions - a person decides.
An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
Checks whether a chatbot invented a fact that is not in the source documents. The license is non-commercial. The checking model itself makes mistakes and does not replace manual review on important tasks.
An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.
Moderation by your own rules: you write the policy in plain text, and the model reasons and gives a decision with an explanation. Built on gpt-oss.
An open Ai2 filter: in a single pass it determines whether a request is harmful, whether a reply is harmful, and whether the bot refused needlessly. Works in English.
A voice conversation partner based on Moshi that listens and speaks at the same time and can be interrupted. The role is set by text, the voice by a sample recording. English only.
A general-purpose audio model: speech recognition, answering questions about sounds, detecting emotions and voice dialogue. Trained on 13 million hours of audio; languages are English and Chinese.
Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.
An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.
Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.
Expressive speech and dialogue synthesis with voice cloning, plus recognition models. Version 3 of the synthesis supports about 100 languages, including Russian, but is non-commercial.
Research translators based on Llama 2. The first ALMA covered 5 pairs with English, including Russian; X-ALMA expanded coverage to 50 languages.
A model that cuts the sound you need out of a recording based on a text description, a mark on the video or a time range: a voice, an instrument, noise. Weights are available on request.
Security models from a large Turkish marketplace, published in GGUF format: reviewing alerts and incidents, English and Turkish.
Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.
Plan by video memory, not by parameter count: the weights, the context and the request queue all consume it. A model that fits the weights exactly will fail on a long document. Leave headroom and test on your own scenario rather than a demo prompt.
An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.