MedicineGGUF2023–2026
FreedomIntelligence (The Chinese University of Hong Kong, Shenzhen) · China
A large family of medical models: chat, an imaging version, the reasoning HuatuoGPT-o1 and the new HuatuoGPT-3 on Qwen3. Does not replace a doctor; decisions are made by a specialist.
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Hints for doctors when reviewing images (Vision)
- Sizes
- 7B – 72B
- Hardware
- from: Laptop
TextOllama2023–2026
Shanghai AI Laboratory · China
Models from Shanghai AI Laboratory. The early InternLM line is general-purpose; the new Intern-S1/S2 is scientific: it understands formulas, molecules, charts and images.
- Research assistant: papers, formulas, data
- Analysis of scientific and technical documents
- Corporate chat on small models
- Sizes
- 1.8B – about 1T
- Hardware
- from: Laptop
TextGGUF2025–2026
Xiaomi · China
Xiaomi models for reasoning and agents: from the compact MiMo-7B to MiMo-V2.6-Pro with 1.02 trillion parameters. The larger versions understand text, images, video and audio, with a 1M token context. Languages: English and Chinese.
- Logic and calculation tasks
- Agents with tools
- Help for developers
- Sizes
- 7B – 1,02T-A42B
- Hardware
- from: Laptop
Text2024–2026
MBZUAI, Institute of Foundation Models (IFM, LLM360 project) · UAE
Fully open models from the UAE: data, training code and intermediate checkpoints are published along with the weights. K2-Horizon (2026) spans 0.9B to 375B with context up to 512K tokens.
- Reasoning, maths and technical questions
- Analysing long documents
- Agents and writing code
- Sizes
- 0.9B – 375B-A23B
- Hardware
- from: Laptop
TextGGUF2025–2026
Renmin University of China (GSAI) and Ant Group (inclusionAI) · China
Diffusion language models: text is written in blocks and then refined rather than word by word, which speeds up generation. LLaDA2.2 can edit what it has written and targets agents. LLaDA-Image is a separate product.
- Fast generation of code and text
- Agent scenarios with long context
- Research into alternatives to standard LLMs
- Sizes
- 8B – 100B (MoE)
- Hardware
- from: 1 GPU
3D2024–2026
NAVER LABS Europe · France (NAVER, South Korea)
The family that started "single-pass" 3D reconstruction from a pair or set of photos without camera calibration. MASt3R added point matching and scale; MUSt3R and BLASt3R added video support.
- 3D scene from several photos without calibration
- Point matching between images
- Mapping from video (SLAM)
- Sizes
- 0.57B – 0.69B
- Hardware
- from: Laptop
Music and soundGGUF2022–2026
m-a-p (Multimodal Art Projection) · UK / China
A music encoder: turns a track into a numeric representation used to detect genre, mood, key and rhythm. MERT-v2 handles full songs up to 6 minutes.
- Automatic tagging of a music catalog
- Finding similar tracks
- Detecting genre, mood and tempo
- Sizes
- 95M – 632M
- Hardware
- from: Laptop
Autonomous driving2020–2026
comma.ai · USA
An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.
- A research testbed for driver assistance systems
- Studying driver attention monitoring with an in-cabin camera
- Comparison with your own lane-keeping algorithms
- Sizes
- compact, designed for an in-vehicle device
- Hardware
- from: Laptop
TextRU2026
SberDevices (ai-forever) · Russia
A Russian and English research prototype: the model writes text in blocks at once (diffusion) rather than word by word, which speeds up responses. The authors do not recommend it for production systems.
- Experiments with faster generation
- Fine-tuning small models for your own tasks
- Research
- Sizes
- 0.6B – 4B
- Hardware
- from: Laptop
Documents and OCR2026
Jina AI · Germany
Document parsing in a single model: a whole page becomes Markdown - text in correct reading order, tables and formulas in LaTeX. Built on DeepSeek-OCR, with only 0.6B of its 3.4B parameters active.
- Converting scans and PDFs to Markdown
- Recognizing tables and formulas
- Parsing invoices, acts and reports
- Sizes
- 3.4B-A0.6B
- Hardware
- from: 1 GPU
Voice assistants2026
Samsung · South Korea
Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.
- Describing an audio recording in words
- Answering questions about a sound
- Identifying the type of sound and the setting
- Sizes
- 99M – 356M
- Hardware
- from: Laptop
Computer vision2024–2026
Microsoft Research · USA
Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.
- Measuring rooms and objects from photos
- 3D point cloud from a single shot
- Preparing data for robots and AR
- Sizes
- ViT-S – ViT-G
- Hardware
- from: Laptop
Robotics2025–2026
GigaAI · China
A robot control model trained mostly on synthetic data from a world model. It reduces spending on collecting data from real robots.
- Controlling a robot arm
- Fine-tuning on a small amount of your own data
- Sorting and assembly pilots
- Sizes
- 3.5B
- Hardware
- from: 1 GPU
TextOllama2024–2026
NVIDIA · USA
NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).
- Agents with tool calling
- Reasoning and calculation tasks
- Answers based on long documents
- Sizes
- 4B – 550B-A55B
- Hardware
- from: Laptop
Weather and climate2024–2026
Allen Institute for AI (Ai2) · USA
A fast climate model emulator: simulates the atmosphere years and decades ahead on a single GPU. Coupled with an ocean model (SamudrACE) for long-term scenarios.
- Decades-long climate scenarios to assess long-term asset risks
- Large-scale what-if runs on temperature and precipitation
- Preparing data for crop yield and energy demand models
- Sizes
- checkpoint of about 1.8 GB
- Hardware
- from: 1 GPU
Weather and climate2023–2026
Google DeepMind · UK
Google DeepMind's family of global weather models: GraphCast (10-day forecast), GenCast (probabilistic ensemble) and WeatherNext 2 with cyclone forecasting. Since August 2026 the weights are cleared for commercial use.
- Medium-range weather forecasts for planning shifts, voyages and deliveries
- Probabilistic assessment of extreme weather for insurance portfolios
- Tropical cyclone track forecasts for marine and port operations
- Sizes
- from lightweight 1° versions to full 0.25°
- Hardware
- from: 1 GPU
Biology and chemistry2022–2026
AlQuraishi Lab (Columbia University) and the OpenFold consortium · USA
A fully open reproduction of AlphaFold 2 and then AlphaFold 3 under Apache 2.0, with training data. OpenFold3 predicts complexes of proteins, nucleic acids and ligands.
- Predicting structures of proteins and ligand complexes
- Fine-tuning on the company's own data (training code is open)
- An in-house structural analysis service without sending data outside
- Sizes
- a single set of weights per version
- Hardware
- from: 1 GPU
Autonomous driving2025–2026
NVIDIA · USA
Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.
- Auto-labeling camera recordings to train your own driver assistance systems
- Analyzing complex road scenes with text explanations
- Testing autopilot systems in simulation on rare scenarios
- Sizes
- 10B – 34B
- Hardware
- from: 1 GPU
Autonomous driving2026
Alibaba (Qwen team) · China
An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.
- A perception and planning prototype for autonomous vehicles on closed sites
- Answering questions about camera recordings when reviewing incidents
- Labeling road scenes to train your own models
- Sizes
- 4B
- Hardware
- from: 1 GPU
Deepfake detection2023–2026
University of Wisconsin-Madison · USA
An early and still used approach: a simple classifier trained on top of a frozen CLIP that transfers to unseen generators. It errs in both directions - the output needs a human check.
- Checking images from new, unfamiliar generators
- A baseline when comparing detectors
- Fast rollout of a check without training a large model
- Sizes
- a linear classifier on top of CLIP ViT-L/14
- Hardware
- from: Laptop
Robotics2025–2026
NVIDIA · USA
"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.
- Synthetic video for training robots and self-driving vehicles
- Testing scenarios in simulation
- Robot control (Policy versions)
- Sizes
- 2B – 65B
- Hardware
- from: 1 GPU
Robotics2026
Ant Group (Robbyant) · China
A robot control model from Ant Group trained on a large volume of data from real robots. Version 2.0 works with different types of robot arms.
- Controlling a two-armed robot
- Fine-tuning for your own operation
- Assembly and sorting pilots
- Sizes
- 4B – 6B
- Hardware
- from: 1 GPU
Robotics2026
Xiaomi · China
Open robot control models from Xiaomi. Robotics-1 is designed for household and kitchen tasks, U0 combines scene understanding and action.
- Controlling a robot arm by command
- Household and service scenarios
- Base for fine-tuning
- Sizes
- 4B – 5B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Moonshot AI · China
Very large Moonshot MoE models for agentic work. K3 (2.8 trillion parameters) was the largest open model at release, with up to 1M tokens of context and image understanding; K2.7-Code is built for programming.
- Multi-step agents: search, data collection, reports
- In-depth document analysis
- Help for developers
- Sizes
- 16B-A3B – 2.8T-A104B
- Hardware
- from: 1 GPU
TextOllama2024–2026
LG AI Research · South Korea
Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.
- Corporate assistant
- Working with Korean and English texts
- Analysis of documents and images (4.5)
- Sizes
- 1.2B – 750B-A37B
- Hardware
- from: Laptop
TextGGUF2025–2026
Swiss AI (ETH Zurich, EPFL, CSCS) · Switzerland
Switzerland's public open model: weights, data and recipe are open, with more than 1000 languages in training. Version 1.5 understands images.
- Multilingual assistant
- Answers based on documents
- Analysis of images and scans (v1.5)
- Sizes
- 0.5B – 70B
- Hardware
- from: Laptop
Image + textGGUF2024–2026
Alibaba (AIDC-AI) · China
Vision models from Alibaba's international division with strong text and table reading. The line includes the Ovis2.6 MoE and separate compact OvisOCR models for documents.
- Extracting data from invoices, contracts and delivery notes
- Table recognition
- Answering questions about photos and charts
- Sizes
- 0.9B – 80B-A3B
- Hardware
- from: Laptop
Tabular data2026
LG AI Research · South Korea
LG's small tabular model: with 21M parameters it nearly matches the leaders in classification and regression accuracy. Weights are for non-commercial use only.
- Pilot churn forecasts
- Testing scoring hypotheses
- Exploring customer data
- Sizes
- about 21M
- Hardware
- from: Laptop
Weather and climate2024–2026
Microsoft Research · USA
A foundation model of Earth's atmosphere: global weather forecasts, plus separate versions for air quality and ocean waves. Computes a forecast in seconds instead of hours on a supercomputer.
- Your own forecast of temperature, wind and precipitation for company locations
- Sea state estimates for planning voyages and port operations
- Air pollution forecasts for industrial sites
- Sizes
- about 1.3B (a small test version is available)
- Hardware
- from: 1 GPU
Weather and climate2022–2026
NVIDIA · USA
NVIDIA's set of weather and climate models: global FourCastNet forecasts, downscaling to kilometers (CorrDiff), regional storm forecasts (StormCast), climate generation (cBottle, Atlas). Run via Earth2Studio.
- Global forecasts followed by downscaling to the region you need
- Short-term forecasts of thunderstorms and heavy rain for dispatch services
- Generating many weather scenarios for stress tests
- Sizes
- 98M – 2.5B (FourCastNet 3 — about 711M)
- Hardware
- from: 1 GPU
Image generationGGUF2026
Ideogram · Canada
Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.
- Testing text-in-image generation
- Research and prototypes
- Sizes
- about 9B
- Hardware
- from: 1 GPU
Math and reasoningOllama2025–2026
Open Thoughts (Stanford, Berkeley and other universities) · USA
Fully open reasoning models: both weights and training data are published. Newer OpenThinkerAgent versions can carry out multi-step tasks.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of internal policies
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Tabular data2026
Google Research · USA
Google's large tabular model: classification and regression from examples without training, with numeric and categorical columns. Weights are for non-commercial use only.
- Pilot comparison with current scoring models
- Exploring customer data
- Testing churn hypotheses
- Sizes
- about 1.6B
- Hardware
- from: 1 GPU
Image + text2023–2026
Shanghai AI Lab (OpenGVLab) · China
A family of video models: encoders for search and classification of clips, and chat models that analyze long videos. InternVideo 3 is designed for multi-hour recordings.
- Searching a video archive with a text query
- Action recognition in video
- Answering questions about a long recording
- Sizes
- small encoders – 9B
- Hardware
- from: Laptop
Satellite and geo2025–2026
Allen Institute for AI (Ai2) · USA
Ai2's family of models for Sentinel-1, Sentinel-2 and Landsat imagery, with ready-made fine-tunes for mangroves, deforestation and ecosystem types. The license excludes the extractive industries.
- Monitoring deforestation and forest condition across the supply chain
- Classifying land and crops from image series
- Image embeddings for finding similar plots
- Sizes
- Nano – Large (Base about 114M)
- Hardware
- from: Laptop
MedicineOllama2023–2026
EPFL · Switzerland
Open medical models from Swiss EPFL, fine-tuned on clinical guidelines on top of various base models. Does not replace a doctor; decisions are made by a specialist.
- Answering staff questions based on clinical guidelines
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 2B – 70B
- Hardware
- from: Laptop
Robotics2025–2026
Ai2 (Allen Institute for AI) · USA
A fully open robot control model that first "reasons" about space and trajectory, then acts. Its reasoning can be checked.
- Controlling a robot arm with explainable steps
- Fine-tuning for your own robot
- Research pilots
- Sizes
- 5B – 8B
- Hardware
- from: 1 GPU
Image + textOllama2023–2026
LLaVA / LMMs-Lab (researchers from the USA and China) · USA / China
The open project that started the trend for image-plus-text models. The OneVision line understands photos, documents and video; training data and recipes are open.
- Answering questions about photos and screenshots
- Describing products from a photo
- Frame-by-frame video analysis
- Sizes
- 0.5B – 72B
- Hardware
- from: Laptop
Documents and OCRGGUF2025–2026
Shanghai AI Laboratory (OpenDataLab) · China
A popular open tool for converting PDFs to Markdown with its own small model. MinerU2.5-Pro was improved through data alone, without growing in size. Languages on the card: Chinese and English.
- Converting PDF reports and contracts to Markdown
- Recognising tables and formulas
- Preparing documents for RAG and search
- Sizes
- 0.9B – 1.2B
- Hardware
- from: Laptop
3D2025–2026
Meta and the University of Oxford (VGG) · USA / UK
Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.
- 3D model of a room or object from a photo series
- Camera pose estimation for photogrammetry
- Point cloud for measurements and comparison with the plan
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
Biology and chemistry2022–2026
EvolutionaryScale / Chan Zuckerberg Biohub (ESM-2 — Meta AI) · USA
Protein language models: they understand amino acid sequences, predict structure (ESMFold2) and help with protein design. Since 2026 all open versions are under MIT.
- Protein embeddings for predicting properties (stability, solubility)
- Predicting 3D structures of proteins and complexes
- Screening enzyme and antibody design candidates before lab work
- Sizes
- 8M – 15B (ESM-2), 300M – 6B (ESM C), 1.4B (open ESM3)
- Hardware
- from: Laptop
CodeOllama2025–2026
Essential AI · USA
An 8B model trained from scratch by the company of one of the authors of the transformer architecture. Strong at code and technical tasks; version 1.5 handles context up to 160K tokens.
- Writing and fixing code
- A developer agent on a single GPU
- Solving technical and scientific problems
- Sizes
- 8B
- Hardware
- from: Laptop
Voice assistants2024–2026
NVIDIA · USA
Models that listen to speech, sounds and music and answer questions about them. Audio Flamingo Next handles recordings up to 30 minutes. Research use only.
- Detailed descriptions of audio recordings
- Questions and answers about a long recording
- Tagging music and sounds
- Sizes
- 0.5B – 8B
- Hardware
- from: Laptop
Fact-checking and judgesRU2025–2026
SberDevices (ai-forever) · Russia
Judge models that evaluate other AI models' answers in Russian: they score against a given criterion and explain the score in text.
- Automatic quality checks of Russian chatbot answers
- Comparing several models before choosing one
- Checking answers after fine-tuning
- Sizes
- 4B – 32B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Zhejiang University · China
A medical model for text, images, 3D scans and video: from a light 4B to a large MoE. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images and CT scans
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 4B – 235B-A22B
- Hardware
- from: Laptop
Math and reasoningGGUF2025–2026
Princeton University · USA
Open models for formal proofs in Lean 4 from Princeton. The new Goedel-Code-Prover proves program correctness.
- Formal verification of mathematical workings
- Verifying code correctness
- Training
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Robotics2025–2026
NVIDIA · USA
NVIDIA's foundation model for humanoid robots and robot arms: it sees, understands a command and outputs movements. Built into the Isaac ecosystem.
- Controlling humanoid robots
- Fine-tuning for your own robotic cell
- Training in simulation with transfer to a real robot
- Sizes
- 2B – 3B
- Hardware
- from: 1 GPU
Text to speechRU2024–2026
Fish Audio · USA / China
Speech synthesis with voice cloning and emotion control in 80+ languages, including Russian. Quality is close to paid services, but the weights are for research only.
- Voice cloning
- Emotional voiceover
- Multilingual voiceover
- Sizes
- 0.5B – about 4.5B
- Hardware
- from: Laptop
Image + textGGUF2023–2026
Shanghai AI Laboratory (OpenGVLab) · China
A large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.
- Understanding documents, diagrams and charts
- Answering questions about photos
- Video analysis
- Sizes
- 1B – 241B-A28B
- Hardware
- from: Laptop
Image + text2024–2026
Ai2 (Allen Institute for AI) · USA
Fully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.
- Counting products and objects in photos
- Pointing to where an item is in an image
- Video analysis
- Sizes
- 1B-A7B – 72B
- Hardware
- from: Laptop
Text to speechRU2026
k2-fsa (Next-gen Kaldi) · China
Speech synthesis with voice cloning from a short sample in 646 languages, including Russian and languages of Russia's peoples. A voice can be described in words. Weights are for non-commercial use only.
- Voiceover in rare languages
- Voice cloning from a sample
- Research and prototypes of multilingual voiceover
- Sizes
- 0.6B
- Hardware
- from: Laptop
Video2025–2026
Skywork AI (Kunlun) · China
An interactive "world model": generates video of a game world in real time and responds to keyboard and mouse input. Version 3.0 keeps scene memory for minutes.
- Game world prototypes without an engine
- Interactive demos and simulations
- Generating data to train agents
- Sizes
- 1.8B – 17B
- Hardware
- from: 1 GPU
Tabular data2025–2026
Lexsi Labs · India
Recent open models for tabular data: they predict from a few examples given in the prompt, with no task-specific training.
- Classification and forecasting on tables with no separate training
- Quickly testing models on new datasets
- Assessing features in large tables
- Sizes
- size not stated on the model card
- Hardware
- from: Laptop
Code2026
Ai2 (Allen Institute for AI) · USA
Fully open developer agents from Ai2: weights, data and training recipe are all public. Designed so a company can cheaply fine-tune the agent on its own repository.
- An agent for fixing issues in code
- Fine-tuning the agent on an internal repository
- Automating small edits and tests
- Sizes
- 8B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2026
LM Provers (CMU, Hugging Face, ETH Zurich, Project Numina) · USA, Switzerland, France
A small 4B model on Qwen3 that writes mathematical proofs in plain language almost at the level of large models. Runs on a laptop.
- Checking the logic of reasoning and workings
- Step-by-step explanations of solutions
- Training and olympiad preparation
- Sizes
- 4B
- Hardware
- from: Laptop
Biology and chemistry2024–2026
Arc Institute (with Together AI, Stanford, NVIDIA) · USA
DNA language models with context up to a million nucleotides: they assess the impact of mutations, annotate genomes and generate sequences. Evo 2 is trained on genomes from all domains of life.
- Assessing the likely harmfulness of genetic variants for research
- Annotating genomes of microorganisms and plants
- Finding promising sequences in breeding and synthetic biology
- Sizes
- 1B – 40B
- Hardware
- from: 1 GPU
MedicineOllama2025–2026
Google · USA
Google's medical version of Gemma: reads medical texts and images (X-ray, dermatology, histology). A tool for doctors and developers; does not replace a doctor, decisions are made by a specialist.
- Draft discharge summaries and reports for a doctor to review
- Hints for doctors when reviewing images
- Searching and summarising medical literature
- Sizes
- 4B – 27B
- Hardware
- from: Laptop
Medicine2023–2026
Stanford AIMI · USA
Stanford models for chest X-rays: they describe the image and prepare a draft report. Does not replace a doctor; decisions are made by a specialist.
- A draft X-ray description for the radiologist
- Hints for doctors when reviewing images
- Checking reports for completeness
- Sizes
- 3B – 8B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Alibaba DAMO Academy · China
Alibaba's medical model based on Qwen2.5-VL: understands many types of medical images and medical text, and can reason step by step. Does not replace a doctor; decisions are made by a specialist.
- Hints for doctors when reviewing images
- Draft reports and discharge summaries
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
MedicineGGUF2025–2026
Ant Healthcare (Ant Group) and Zhejiang Provincial Medical Information Center · China
A large medical MoE model based on Ling-flash-2.0: 100B parameters with 6B active, so it answers quickly. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff on clinical questions
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 100B-A6B
- Hardware
- from: 1 GPU
TextOllama2024–2026
Allen Institute for AI (Ai2) · USA
Fully open models: not only the weights but also the data, training code and intermediate checkpoints are published. Useful when transparent provenance matters.
- Assistant and answers based on documents
- Reasoning tasks (Think versions)
- Fine-tuning on your data with a clear model history
- Sizes
- 1B – 32B
- Hardware
- from: Laptop
TextGGUF2024–2026
Prime Intellect · USA
Models trained in a distributed way on GPUs from around the world. INTELLECT-3 (106B) is further trained with reinforcement learning for math, code and agents.
- Reasoning and math tasks
- Programming help
- Agents with tool calling
- Sizes
- 10B – 106B-A12B
- Hardware
- from: 1 GPU
Weather and climate2024–2026
European Centre for Medium-Range Weather Forecasts (ECMWF) · Europe (intergovernmental organization)
ECMWF's weather neural network running operationally: a 15-day forecast four times a day, an ensemble version with 51 scenarios, and since version 2, ocean waves.
- Running your own forecast from open initial data
- Ensemble forecasts to estimate the probability of frost, downpours and storms
- Wave forecasts for marine operations
- Sizes
- checkpoint of about 1 GB
- Hardware
- from: 1 GPU
Computer visionGGUF2024–2025
ByteDance and the University of Hong Kong · China
Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.
- Estimating distances and volumes from a camera
- Depth effects for photo and video
- Navigation for robots and drones
- Sizes
- 25M – 1.4B
- Hardware
- from: Laptop
Speech to textRU2025
Meta · USA
Speech recognition for 1,600+ languages, including Russian and rare languages no system supported before. A new language can be added from a few examples.
- Transcription in rare and local languages
- Digitizing oral archives
- Subtitles in many languages
- Sizes
- 300M – 7B
- Hardware
- from: Laptop
Image + text2023–2025
Zhipu AI (Z.ai) and Tsinghua University · China
Vision models from Zhipu: first CogVLM, then the GLM-V line. GLM-4.6V can call tools based on images and act as an agent operating an interface.
- Answering questions about photos and documents
- An agent that operates an interface from screenshots
- Analysing charts and reports
- Sizes
- 9B – 106B-A12B
- Hardware
- from: Laptop
3D2025
Meta and Carnegie Mellon University · USA
A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.
- 3D reconstruction of an object or room from photos
- Exporting the scene to COLMAP format for further processing
- Depth and camera pose estimation
- Sizes
- about 1.2B
- Hardware
- from: 1 GPU
3D2025
Shanghai AI Lab · China
Reconstructs a 3D scene and camera positions from a set of photos or a video without relying on a "reference" frame. Pi3X gives smoother point clouds and approximate scale in meters.
- 3D scene reconstruction from video
- Camera pose estimation from frames
- Point clouds for research and prototypes
- Sizes
- 0.96B – 1.4B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek's maths models. The first 7B version introduced the GRPO training method; the 685B V2 writes and checks its own olympiad-level proofs.
- Calculations and formula checks
- Checking mathematical workings in reports
- Working through problems step by step
- Sizes
- 7B – 685B
- Hardware
- from: Laptop
3DGGUF2025
Meta · USA
Reconstructs the 3D shape of an object or a human body from one ordinary photo, even when the object is partly hidden. Two models: Objects and Body.
- 3D model of an item from a catalog photo
- Estimating body pose and shape from a photo
- Try-on and AR scenarios
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
Satellite and geo2025
IBM and the European Space Agency (ESA) · USA / Europe
A multimodal Earth model: understands optical and radar imagery, terrain, vegetation index and land use maps, and can generate a missing data type (for example, a "see-through-clouds" image from radar).
- Analyzing fields and forests even in cloudy weather using radar imagery
- Land use maps for assessing plots
- Flood and wildfire assessment (ready-made fine-tunes available)
- Sizes
- tiny – large (checkpoints from ~200 MB to ~3.8 GB)
- Hardware
- from: Laptop
Computer vision2023–2025
Meta · USA
Meta's open reproduction of CLIP with a transparent data collection recipe. MetaCLIP 2 is trained on multilingual data from around the world. Non-commercial license only.
- Image search by text
- Image classification without training
- Search research and prototypes
- Sizes
- 0.15B – 3.6B
- Hardware
- from: Laptop
Documents and OCRGGUF2025
Ai2 (Allen Institute for AI) · USA
A model and toolkit for converting PDFs into clean text at scale, preserving reading order, tables and formulas. Built to process millions of pages.
- Bulk digitisation of a PDF archive
- Converting contracts and reports into text
- Preparing documents for search and RAG
- Sizes
- 7B
- Hardware
- from: 1 GPU
Biology and chemistry2024–2025
MIT (Jameel Clinic) and Recursion · USA
An open MIT-licensed alternative to AlphaFold 3: predicts structures of protein, DNA and small-molecule complexes; Boltz-2 estimates binding strength, BoltzGen designs new binding proteins.
- Predicting how a candidate molecule binds to a target protein
- Ranking compounds by predicted binding strength before synthesis
- Designing binder proteins for a given target
- Sizes
- checkpoints of about 2 GB
- Hardware
- from: 1 GPU
Deepfake detection2025
National Institute of Informatics, Yamagishi Lab · Japan
Seven speech encoders (wav2vec 2.0, XLS-R, MMS, HuBERT) post-trained to tell live speech from synthetic. The authors note themselves that quality depends heavily on the dataset; a human reviews the output.
- Checking audio recordings for synthesis
- Fine-tuning for your own language and recording channel
- Comparing several encoders on your own data
- Sizes
- 0,3B – 2B
- Hardware
- from: Laptop
Robotics2025
Physical Intelligence · USA
Robot control models from Physical Intelligence: folding laundry, tidying up, handling objects. π0.5 copes better in unfamiliar settings.
- Controlling robot arms and two-armed robots
- Fine-tuning for your own operations
- Pilots for automating manual work
- Sizes
- about 3B
- Hardware
- from: 1 GPU
Satellite and geo2023–2025
IBM and NASA · USA
Foundation models for Landsat and Sentinel-2 satellite imagery that account for image time series. Ready-made fine-tunes for floods, burn scars and crop types, plus a separate weather model, WxC.
- Mapping crops and field condition over the season
- Assessing flood zones and burn scars after natural disasters
- Monitoring changes in buildings and land use
- Sizes
- tiny – 600M (imagery), 2.3B (Prithvi WxC weather)
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Moonshot AI and Project Numina · China, France
Models for formal proofs in Lean 4 from Moonshot AI (Kimi) and Numina. Small versions from 0.6B run on a laptop.
- Formal verification of mathematical workings
- Translating a problem from plain language into Lean
- Training and olympiad preparation
- Sizes
- 0.6B – 72B
- Hardware
- from: Laptop
Computer vision2023–2025
Meta · USA
Foundation models that turn an image into a numeric "fingerprint". They are used to build similar-image search, classification and segmentation without large labeled datasets.
- Finding similar products and photos
- Image classification on small datasets
- Base for your own quality-control models
- Sizes
- 21M – 7B
- Hardware
- from: Laptop
Text2024–2025
xAI · USA
xAI publishes the weights of previous Grok generations. The models are very large and need a GPU cluster, so in practice they are rarely run.
- Research on large models
- Assistant on your own infrastructure
- Text generation and analysis
- Sizes
- 314B (Grok-1), Grok-2 is larger
- Hardware
- from: Cluster
Faces2025
ByteDance · China
Transformer-based face recognition from ByteDance, one of the most accurate open models on benchmarks. Weights are published in ONNX format but are non-commercial.
- Comparing faces and searching a photo database
- Research on recognition accuracy
- Access control prototypes
- Sizes
- ViT-T – ViT-L
- Hardware
- from: Laptop
Medicine2025
Google · USA
Google's lightweight encoder for medical images and text, the same one inside MedGemma. Sorts images and finds similar ones. Does not replace a doctor; decisions are made by a specialist.
- Finding similar images in a clinic's archive
- Pre-sorting images for a doctor
- A base for your own image classifiers
- Sizes
- about 0.9B
- Hardware
- from: Laptop
MedicineGGUF2025
Intelligent Internet · UK
Reasoning medical models on Qwen3, designed to run on an ordinary computer. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff with the reasoning shown
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Math and reasoningOllama2025
Agentica (Berkeley, Sky Computing Lab) and Together AI · USA
Small models fine-tuned with reinforcement learning: DeepScaleR (1.5B) solves olympiad maths, DeepCoder writes code, DeepSWE works as a developer agent. Recipes and data are open.
- Solving maths problems with step-by-step working
- Generating and checking code
- An agent for fixing bugs in a repository
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Fact-checking and judges2024–2025
OpenCompass (Shanghai AI Laboratory) · China
A line of judges from the team behind open model benchmarks: they score answers and check them against a reference. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring model answers against set criteria
- Checking an answer against a reference solution
- Comparing several models on your own data
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoning2025
NVIDIA · USA
NVIDIA models for maths and reasoning based on Qwen. AceReason was fine-tuned with reinforcement learning first on maths, then on code.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Working through programming problems
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Robotics2025
Hugging Face · USA
A small robot control model that runs on a regular laptop. Trained on open data from the LeRobot community, suited to low-cost robot arms.
- Controlling a low-cost robot arm
- Quick robotization pilots and demos
- Training staff and students
- Sizes
- 450M
- Hardware
- from: Laptop
Image + textGGUF2025
Moonshot AI · China
An efficient MoE vision model (16B, 3B active) with a long context and a reasoning version. Handles long documents and video well.
- Analysing long PDFs and presentations
- Answering questions about video
- Operating interfaces from screenshots
- Sizes
- 16B-A3B
- Hardware
- from: 1 GPU
Satellite and geo2025
MBZUAI · UAE
A compact research model for satellite imagery, trained on both optical (Sentinel-2) and radar (Sentinel-1) data. Narrower in scope and community than Prithvi and TerraMind.
- Classification and segmentation of satellite imagery after fine-tuning
- A base for a land monitoring prototype
- Sizes
- TerraFM-B (ViT-Base)
- Hardware
- from: Laptop
Deepfake detection2024–2025
Xiaohongshu, USTC and Shanghai Jiao Tong University · China
An AI-image detector made of several experts: some look at visual artifacts, others at noise. The hard Chameleon benchmark was released with it. It errs in both directions - a human reviews the result.
- Checking realistic AI images without obvious artifacts
- Comparing detectors on hard examples
- Fine-tuning for your own type of content
- Sizes
- several experts based on ConvNeXt and CLIP
- Hardware
- from: 1 GPU
TextOllama2025
DeepSeek · China
A reasoning model that thinks step by step before answering. Strong at calculations, logic and code; compact distilled versions are available.
- Complex calculations and logic checks
- Analysis of contracts and internal policies
- Help for developers
- Sizes
- 1,5B – 671B
- Hardware
- from: Laptop
Image generationGGUF2025
ByteDance Seed · China
A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.
- Photo editing in a conversation
- Answering questions about an image
- Image generation with explanations
- Sizes
- 14B-A7B
- Hardware
- from: 1 GPU
Finance2025
The Fin AI · international project
Models that spell out their reasoning before answering financial questions with numbers and tables. Not investment advice: decisions are made by a specialist.
- Calculation questions on statements with the working shown
- Working through tasks with tables and numbers from documents
- Checking calculations made by hand
- Sizes
- 8B и 14B
- Hardware
- from: 1 GPU
Fact-checking and judgesRU2024–2025
NVIDIA · USA
Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data to fine-tune your own model
- Scoring assistant answers in Russian and other languages
- Sizes
- 49B, 70B and 340B
- Hardware
- from: Cluster
Math and reasoningGGUF2024–2025
DeepSeek · China
DeepSeek models for formal proofs in Lean 4: the proof is checked by a program, not a person. A narrow tool for mathematicians and engineers.
- Formal verification of mathematical workings
- Verifying algorithm correctness
- Training and olympiad preparation
- Sizes
- 7B – 671B
- Hardware
- from: Laptop
Video2024–2025
HPC-AI Tech · Singapore
A fully open video generation project: weights, code and training recipe. Version 2.0 at 11B makes video from text and from an image.
- Video from a text description
- Animating images
- Training your own video model
- Sizes
- up to 11B
- Hardware
- from: 1 GPU
Video2025
StepFun · China
A large 30B video model producing clips of up to 204 frames. Needs server hardware, but is open under MIT.
- Video from a description
- Animating images
- Sizes
- 30B
- Hardware
- from: Cluster
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Qwen's first open reasoning model: it thinks step by step before answering and comes close to DeepSeek-R1 on maths tasks with only 32B parameters.
- Calculations and formula checks
- Complex analytics with step-by-step breakdowns
- Checking the logic of contracts and internal policies
- Sizes
- 32B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
Stanford University · USA
A reasoning model trained on just a thousand problems. It can be told to think longer to answer a hard question more accurately.
- Calculations and formula checks
- Working through complex problems step by step
- Training
- Sizes
- 1.5B – 32B
- Hardware
- from: Laptop
Math and reasoningGGUF2025
Qihoo 360 · China
Reasoning models from Qihoo 360: a standard Qwen2.5 was fine-tuned for long reasoning using an open recipe; data and code are published.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Text to speechGGUF2024–2025
Shanghai Jiao Tong University and partners · China
A voice cloning model that needs only a few seconds of a sample, in English and Chinese. The community has released many fine-tuned versions for other languages, including Russian.
- Voice cloning
- Voicing audiobooks and videos
- Research and prototypes
- Sizes
- about 340M
- Hardware
- from: Laptop
Visual document search2025
Nomic AI · USA
Search across PDF pages and scans as images. The cards list English, Italian, French, German and Spanish — Russian is not among them.
- Search across an archive of scans and PDFs
- Search across tables and diagrams inside documents
- Picking pages for an AI assistant answer
- Sizes
- 3B and 7B
- Hardware
- from: 1 GPU
Math and reasoningGGUF2025
NovaSky (Sky Computing Lab, Berkeley) · USA
A Berkeley reasoning model trained for under 450 dollars. It showed that o1-preview-level reasoning can be reproduced with modest resources.
- Calculations and formula checks
- Working through problems step by step
- A base for your own reasoning fine-tuning
- Sizes
- 7B – 32B
- Hardware
- from: Laptop
Robotics2024–2025
Stanford, Berkeley and partners · USA
The first large open vision-language-action model: a robot arm carries out commands like "put the apple in the bowl". OFT makes it several times faster.
- Controlling a robot arm by text command
- Pilots for robotizing simple operations
- Base for fine-tuning to your own robot
- Sizes
- 7B
- Hardware
- from: 1 GPU
TextOllama2023–2025
Ai2 · USA
Ai2 fine-tunes of Llama with a fully open recipe: data, code and all intermediate stages. Tulu 3 405B is one of the largest openly fine-tuned models; OLMo chat versions use the same recipe.
- Employee assistant on your own server
- Math and precise instruction following
- Reference recipe for your own fine-tuning
- Sizes
- 7B – 405B
- Hardware
- from: Laptop
Computer-use agentsGGUF2025
Microsoft Research · USA
An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.
- Pilots in interface control
- Research projects spanning screens and robotics
- Analyzing screenshots with an action plan
- Sizes
- 8B
- Hardware
- from: 1 GPU
Medicine2024–2025
Bioptimus · France
Pathology foundation models from France's Bioptimus with 1.1B parameters, plus the compact H0-mini. The first version is open under Apache 2.0. Does not replace a doctor; decisions are made by a specialist.
- Histology slide patch features for research models
- A prototype for sorting slides by tissue type
- Research projects linking morphology and molecular data
- Sizes
- 86M (H0-mini) – 1,1B
- Hardware
- from: Laptop
Math and reasoningOllama2024–2025
Qwen (Alibaba) · China
Maths versions of Qwen: they solve problems step by step and can calculate via code. Includes reward models that check each step of a solution.
- Calculations and formula checks
- Checking calculations in estimates and reports
- Working through problems step by step
- Sizes
- 1.5B – 72B
- Hardware
- from: Laptop
Medicine2024–2025
Mahmood Lab (Mass General Brigham, Harvard) · USA
A foundation model for histology slides: turns patches of digital slides into features for tissue classification. Does not replace a doctor; decisions are made by a specialist.
- Research classifiers of tissue types from digital slides
- Finding similar cases in a slide archive
- Preparing features for research prognosis models
- Sizes
- about 300M (UNI) – about 680M (UNI2-h)
- Hardware
- from: Laptop
Fact-checking and judges2025
Atla · UK
An 8B judge model: it scores another model answer against your criteria and writes a rationale. The judge itself makes mistakes and does not replace manual review on important tasks.
- Scoring chatbot answers against your own criteria
- Comparing two versions of a prompt or model
- Filtering out weak answers before they reach a person
- Sizes
- 8B
- Hardware
- from: 1 GPU
Music and sound2024
Tencent AI Lab · China
The MuQ music encoder and the MuQ-MuLan model, which matches music and text: you can search for tracks by a description in English or Chinese.
- Searching music by text description
- Tagging tracks by genre and mood
- Finding similar music
- Sizes
- 300M – 700M
- Hardware
- from: Laptop
Medicine2024
Mahmood Lab (Mass General Brigham, Harvard) · USA
Image-plus-text models for pathology: search slides by an English description, classify without fine-tuning; TITAN describes a whole slide. Does not replace a doctor; decisions are made by a specialist.
- Text-query search across a slide archive for research
- Draft slide descriptions for research projects
- Tissue classification without labels at the start of a study
- Sizes
- about 160M – 300M
- Hardware
- from: Laptop
Video2024
Zhipu AI (Z.ai) and Tsinghua University · China
A 2–5B video model that runs on a single gaming GPU. A popular base for research and add-ons.
- Short clips from text
- Animating images
- Video fine-tuning experiments
- Sizes
- 2B – 5B
- Hardware
- from: 1 GPU
Music and sound2023–2024
Meta · USA
Generates instrumental music from a text description or a sample melody. One of the first open models of its kind.
- Draft music sketches
- Music for video prototypes
- Research
- Sizes
- 300M – 3.3B
- Hardware
- from: Laptop
Image + textGGUF2024
Google · USA
Google's vision model built on Gemma, designed as a base for fine-tuning on a narrow task: captions, object detection, reading text.
- Fine-tuning for your own recognition task
- Finding objects in photos
- Reading text in images
- Sizes
- 3B – 28B
- Hardware
- from: Laptop
TextOllama2024
Nexusflow · USA
Fine-tuned Llama 3 and Qwen 2.5 models from Nexusflow. Athene-V2-Agent is specially trained for function calling and agent scenarios. Commercial use is prohibited.
- Research on agents and function calling
- Comparison with commercial models
- Experiments with a chat assistant
- Sizes
- 70B – 72B
- Hardware
- from: 1 GPU
Biology and chemistry2024
Google DeepMind and Isomorphic Labs · UK
The reference model for the structure of biomolecules and their complexes. Weights are provided for non-commercial research only; companies need commercial access via Google Cloud or open alternatives (Boltz, OpenFold3).
- Academic research on protein and complex structures
- Benchmarking open alternatives against the reference on your own targets
- Sizes
- a single set of weights
- Hardware
- from: 1 GPU
Finance2023–2024
The Fin AI / ChanceFocus · international project
One of the first open model families for financial text: reading statements, news and questions about numbers. Not investment advice: decisions are made by a specialist.
- Reading financial statements and press releases
- Answering questions about numeric data in documents
- Classifying financial texts
- Sizes
- 0.5B – 30B
- Hardware
- from: Laptop
Finance2023–2024
AI4Finance Foundation · USA
An open set of lightweight add-ons for ordinary language models that work with financial texts and news. Not investment advice: decisions are made by a specialist.
- Assessing the tone of financial news and reports
- Tagging mentions of companies and instruments in text
- Preparing digests from a stream of business news
- Sizes
- adapters for 6B - 20B base models
- Hardware
- from: 1 GPU
Visual document search2024
TIGER-Lab · Canada
Turns an image-plus-text model into an embedding model: one vector for a page, a diagram or a captioned photo. The card states English.
- Search across a mixed archive of texts and images
- Search across document pages as images
- Finding similar cards and illustrations
- Sizes
- about 4B (based on Phi-3.5-V)
- Hardware
- from: 1 GPU
Visual document search2024
University of Waterloo, Tevatron project · Canada
Searches page screenshots: the page is not OCRed but turned into a single vector, so the index is more compact than with late-interaction models. The card lists English and French.
- Search across scans and PDFs without OCR
- Search across presentations and reports with complex layouts
- Picking pages for an AI assistant answer
- Sizes
- 2B (based on Qwen2-VL)
- Hardware
- from: 1 GPU
TranslationRUNot maintained2023–2024
Johns Hopkins University and Microsoft · USA
Research translators based on Llama 2. The first ALMA covered 5 pairs with English, including Russian; X-ALMA expanded coverage to 50 languages.
- Translation between English and Russian
- Experiments with LLM-based translation
- Base for fine-tuning a translator
- Sizes
- 7B – 13B
- Hardware
- from: 1 GPU
Image + textNot maintained2023–2024
Hugging Face · France / USA
Open vision models from Hugging Face that reproduced the closed Flamingo. Idefics3 became the basis for the compact SmolVLM line.
- Answering questions about images
- Analysing documents and screenshots
- A base for fine-tuning
- Sizes
- 8B – 80B
- Hardware
- from: 1 GPU
MedicineNot maintained2024
Paige · USA
Paige's pathology foundation model, trained on millions of digital slides. The first version is Apache 2.0, the second is for research only. Does not replace a doctor; decisions are made by a specialist.
- Slide patch features for research classifiers
- Selecting slides for re-review in research projects
- Comparison with other pathology models on your own archive
- Sizes
- about 632M
- Hardware
- from: Laptop
MedicineGGUFNot maintained2023–2024
M42 Health · UAE
Clinical models from Abu Dhabi-based M42, built on Llama and tuned to answer medical questions. Does not replace a doctor; decisions are made by a specialist.
- Reference answers to staff on clinical questions
- Draft discharge summaries and letters for a doctor to review
- Searching medical literature
- Sizes
- 8B – 70B
- Hardware
- from: Laptop
Tabular dataNot maintained2024
ML Foundations · USA
A foundation model for predictions on tables: it classifies and forecasts from a handful of examples, with no separate task-specific training.
- Classifying table rows from a few examples
- Predicting a value from a data row
- Quickly testing hypotheses on new datasets
- Sizes
- 8B
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
01.AI · China
Bilingual (English and Chinese) 01.AI models of 6–34B, with versions supporting up to 200K tokens of context. No new open releases since 2024.
- Chat assistant on a single GPU
- Analysis of long documents
- Classification and data extraction from text
- Sizes
- 6B – 34B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2024
UC Santa Barbara and co-authors · USA
A Longformer-based AI-text detector: it holds a long document whole and was trained on texts from many different language models. It errs in both directions; its output is a reason for a human to check.
- Checking long articles and reports as a whole
- Filtering machine text in a publication flow
- Comparing detectors on your own data
- Sizes
- about 150M (Longformer-base)
- Hardware
- from: Laptop
Fact-checking and judgesNot maintained2024
RLHFlow · USA
An answer scorer that returns a breakdown across several attributes rather than a single overall score. The scorer itself makes mistakes and does not replace manual review on important tasks.
- Choosing the best of several candidate answers
- Preparing data for model fine-tuning
- Scoring assistant answers across several attributes
- Sizes
- 8B
- Hardware
- from: 1 GPU
MedicineGGUFNot maintained2024
Avignon University and Nantes University · France
Mistral 7B fine-tuned on PubMed Central papers, plus several merges with the general model. Compact and easy to run. Does not replace a doctor; decisions are made by a specialist.
- Searching and summarising medical papers
- Draft reference materials for staff
- Explaining medical terminology
- Sizes
- 7B
- Hardware
- from: Laptop
MedicineGGUFNot maintained2024
Saama AI Research · India
Llama 3 fine-tuned on medical and biological data. One of the first strong open medical models of 2024. Does not replace a doctor; decisions are made by a specialist.
- Extracting data from medical documents
- Draft discharge summaries for a doctor to review
- Searching medical literature
- Sizes
- 8B, 70B
- Hardware
- from: Laptop
TextOllamaNot maintained2023–2024
Hugging Face (H4) · USA
Hugging Face educational chat models based on Mistral, Gemma and Mixtral with an open fine-tuning recipe. Zephyr 7B Beta showed a small model can be trained to large-model level without human labeling.
- Lightweight chat assistant
- Reference and starting point for your own fine-tuning
- Drafts of texts and replies
- Sizes
- 7B – 141B-A35B
- Hardware
- from: Laptop
Virtual try-onNot maintained2024
KAIST and OMNIOUS.AI · South Korea
One of the best-known open virtual try-on models: moves a garment from a product photo onto a photo of a person, keeping prints and logos well. Non-commercial license.
- Pilot of a fitting room on a store website
- Prototype product cards on a model without a photo shoot
- Comparison with commercial try-on services
- Sizes
- based on SDXL
- Hardware
- from: 1 GPU
Text analysisNot maintained2024
University of Southern Denmark · Denmark
Turns a free-form occupation description into a standard HISCO code in 13 languages. Built for historical archives, but also useful for cleaning up job title reference lists. A human makes the decision about a candidate; automatic screening without review must not be used.
- Mapping mixed occupation names onto a single code
- Processing archives of HR and statistical data
- Preparing data for reporting
- Sizes
- based on CANINE-s, size not stated on the model card
- Hardware
- from: Laptop
Virtual try-onNot maintained2024
Xiao-i Research · China
An early popular open try-on model: one version for upper-body garments, another for full-length outfits. Non-commercial license.
- Trying tops on a model photo
- Full-length try-on: tops, bottoms, dresses
- Fitting room prototype for testing
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
TinyLlama (SUTD researchers) · Singapore
A 1.1B model with the Llama 2 architecture, trained on 3 trillion tokens. Now behind newer small models, but still a popular base for experiments and fine-tuning.
- Simple chatbots on low-end hardware
- Experiments and team training
- A base for fine-tuning on a narrow task
- Sizes
- 1.1B
- Hardware
- from: Laptop
Deepfake detectionNot maintained2023–2024
Sichuan University and co-authors · China
An open model for finding forgeries in images: it outputs a pixel-level mask of altered regions. It errs in both directions; its map is a hint for an expert, not proof of a forgery.
- Finding pasted and erased fragments in photos
- Checking document scans for edits
- A baseline when comparing manipulation-localization models
- Sizes
- a Vision Transformer based model
- Hardware
- from: 1 GPU
ForecastingNot maintained2024
ServiceNow, Mila and partners · Canada
One of the first open out-of-the-box forecasting models. Tiny, gives a probabilistic forecast, now behind Chronos and TimesFM.
- Probabilistic sales forecast
- Quick forecasting pilots
- Baseline model for comparison
- Sizes
- 2.4M
- Hardware
- from: Laptop
Virtual try-onNot maintained2024
KAIST · South Korea
A research try-on model from CVPR 2024, one of the first built on Stable Diffusion. Now mostly used as a comparison baseline.
- Pilot of upper-body garment try-on
- Comparing quality of different try-on models
- Training your own try-on on the open code
- Sizes
- based on Stable Diffusion
- Hardware
- from: 1 GPU
Photo editingNot maintained2024
XPixel Group (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shanghai AI Lab and others) · China
Powerful SDXL-based restoration of badly damaged photos: it recreates details rather than just upscaling. Hardware-hungry; non-commercial license.
- Restoring old and blurry photos
- Upscaling with detail reconstruction
- Archive restoration pilots
- Sizes
- based on SDXL, plus LLaVA 13B for captions
- Hardware
- from: 1 GPU
TextOllamaNot maintained2023–2024
WizardLM (Microsoft and Peking University) · USA / China
Fine-tunes of Llama, Mistral and StarCoder using Evol-Instruct, which automatically makes instructions more complex. WizardLM-2 was released in April 2024 and removed almost immediately, so only the 2023 versions are relevant.
- Complex multi-step instructions
- Help for developers
- Solving math problems
- Sizes
- 7B – 70B
- Hardware
- from: Laptop
Satellite and geoNot maintained2023–2024
Allen Institute for AI (Ai2) · USA
Pretrained models from the Satlas project for Sentinel-2, Landsat and high-resolution aerial imagery. The predecessor of OlmoEarth, still used in TorchGeo.
- Detecting objects in imagery: solar farms, wind turbines, ships
- Mapping roads and buildings from aerial photos
- A starting point for fine-tuning your own geo model
- Sizes
- Swin-v2 and ResNet backbones (Base)
- Hardware
- from: Laptop
RerankersGGUFNot maintained2023–2024
University of Waterloo, Castorini group · Canada
Rerankers that are language models: they receive the whole list of retrieved passages and reorder it as a list, instead of scoring passages one by one. Heavier than ordinary rerankers.
- Reordering a long list of search results
- Selecting sources for an AI assistant answer
- Research comparisons of retrieval approaches
- Sizes
- 7B – 13B
- Hardware
- from: 1 GPU
Voice assistantsRUNot maintained2023
Meta · USA
Speech and text translation across roughly a hundred languages, including Russian: speech to text, speech to speech, and streaming translation that keeps intonation.
- Speech-to-speech translation
- Translating and transcribing recordings
- Streaming translation
- Sizes
- 281M – 2.3B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
Microsoft Research · USA
Microsoft research models based on Llama 2, trained to choose a reasoning approach for each task. The orca-mini model in Ollama is a different project by independent developer Pankaj Mathur.
- Research on reasoning methods
- Comparison with modern small models
- Training specialists
- Sizes
- 7B – 13B
- Hardware
- from: Laptop
Tabular dataNot maintained2023
OSU NLP Group, Ohio State University · USA
A general-purpose model for tables: filling gaps, finding rows, matching columns and answering questions about the data.
- Answering questions about tables inside documents
- Matching columns across different tables
- Finding and completing records in reference books
- Sizes
- 7B
- Hardware
- from: Laptop
Text to speechRUNot maintained2023
Coqui · Germany
A popular model for cloning a voice from a short sample in 17 languages, including Russian. Coqui has shut down and development has stopped.
- Voice cloning from a sample
- Multilingual voiceover
- Research and prototypes
- Sizes
- about 470M
- Hardware
- from: Laptop
Documents and OCRNot maintained2023
Meta · USA
An early model that converts scientific PDFs into text with formulas. Now outdated and outperformed by almost all modern OCR models.
- Converting scientific papers from PDF into text with formulas
- Digitising technical documentation
- Sizes
- 250M – 350M
- Hardware
- from: Laptop
MedicineNot maintained2023
Bo Wang's lab (University of Toronto, Vector Institute) · Canada
An early medical model on Llama 2 70B, trained on dialogues based on medical texts. Now mainly of research interest. Does not replace a doctor; decisions are made by a specialist.
- Research pilots on medical dialogue
- Training materials for staff
- Comparison with newer medical models
- Sizes
- 70B
- Hardware
- from: 1 GPU
MedicineNot maintained2023
Shanghai Jiao Tong University and Shanghai AI Lab · China
An early general-purpose radiology model: understands 2D and 3D images (CT, MRI) together with text. More of a research base. Does not replace a doctor; decisions are made by a specialist.
- Research pilots on CT and MRI analysis
- Hints for doctors when reviewing images
- A base for fine-tuning on the clinic's own images
- Sizes
- size not stated on the model card
- Hardware
- from: 1 GPU
TextRUGGUFNot maintained2022–2023
Sber (ai-forever) · Russia
Sber's multilingual model covering 61 languages, including languages of the peoples of Russia and the CIS. Separate fine-tunes exist for Buryat, Yakut, Tatar, Bashkir, Kazakh and others, rare for open models.
- Texts in languages of the peoples of Russia and the CIS
- Base for fine-tuning on a less common language
- Drafts and templates in several languages
- Sizes
- 1.3B – 13B
- Hardware
- from: Laptop
TextOllamaNot maintained2023
LMSYS (Berkeley and partners) · USA
One of the first open chat models (2023): LLaMA fine-tuned on user conversations with ChatGPT. A historical milestone; today it is weaker than any modern model of the same size.
- Experiments and team training
- Simple chat assistant for tests
- Comparison with newer models
- Sizes
- 7B – 33B
- Hardware
- from: Laptop
CybersecurityNot maintained2022–2023
Ehsan Aghaei and co-authors · USA
A compact language encoder trained on cybersecurity texts: tagging threat reports, finding entities and classification.
- Tagging threat reports and vulnerability bulletins
- Extracting entities from security texts
- Classifying and searching an internal incident base
- Sizes
- около 125M
- Hardware
- from: Laptop
Weather and climateNot maintained2023
Huawei Cloud · China
One of the first weather neural networks, published in Nature and added to ECMWF charts. The weights are open for research only; commercial use is prohibited.
- Research weather forecasts a week ahead
- Comparison with other weather models on your own data
- Training courses on weather neural networks
- Sizes
- 4 models of ~1.1 GB each (1, 3, 6 and 24-hour steps)
- Hardware
- from: Laptop
Photo editingNot maintained2023
S-Lab, Nanyang Technological University · Singapore
One of the first Stable Diffusion-based photo upscalers: restores realistic details. Non-commercial license.
- Upscaling photos with detail reconstruction
- Research restoration pilots
- Comparison with classic upscalers
- Sizes
- based on SD 2.1
- Hardware
- from: 1 GPU
Image generationNot maintained2023
DeepFloyd (Stability AI) · UK
An early model that was among the first to render text on images accurately. Today it is mostly of historical interest; development has stopped.
- Research experiments
- Prototype images with captions
- Sizes
- 0.4B – 4.3B
- Hardware
- from: 1 GPU
TextNot maintained2022–2023
EleutherAI · USA
Fully open models from the non-profit lab EleutherAI: GPT-NeoX-20B and the Pythia series with published intermediate training checkpoints.
- Base model for fine-tuning
- Research into model behavior
- Simple text generation and completion
- Sizes
- 70M – 20B
- Hardware
- from: Laptop
TextNot maintained2022
Meta · USA
An early open Meta series matching GPT-3 in size. Outdated; useful for research and comparison.
- Research experiments
- Training specialists
- Comparison with modern models
- Sizes
- 125M – 66B (175B on request)
- Hardware
- from: Laptop
TextGGUFNot maintained2022
BigScience (Hugging Face and the community) · France
One of the first large open models, trained by a community of hundreds of researchers in 46 languages. Today it is interesting mostly as a historical milestone.
- Text generation and translation in many languages
- Experiments and team training
- Base model for fine-tuning on a narrow task
- Sizes
- 560M – 176B
- Hardware
- from: Laptop