AI models for manufacturing and warehousing

In manufacturing and warehousing, open models inspect product quality via cameras, count items, forecast inventory and control robots. Many run on inexpensive hardware right next to the line. Look at speed, whether you can fine-tune on your own data, and the license terms.

67 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Tabular data2025–2026

LimiX

Stable AI (Beijing, with Tsinghua University) · China

A table model that alone can classify, predict numbers and fill in missing data. The lightweight LimiX-2M runs on an ordinary computer.

  • Filling gaps in 1C and CRM exports
  • Churn prediction and scoring
  • Classifying customers and products
Sizes
2M – 16M and LimiX-2
Hardware
from: Laptop
Commercial use with conditionsDetails
Autonomous driving2020–2026

comma.ai openpilot (supercombo)

comma.ai · USA

An open driver assistance system: a neural network keeps the lane and controls speed from a camera, plus a driver attention monitoring model. The models live right in the repository and are updated constantly.

  • A research testbed for driver assistance systems
  • Studying driver attention monitoring with an in-cabin camera
  • Comparison with your own lane-keeping algorithms
Sizes
compact, designed for an in-vehicle device
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistants2026

Samsone

Samsung · South Korea

Tiny audio-understanding models for smartphones: they listen to speech, music and ambient sounds and answer in text - describing a recording and answering questions about it. They run on the device itself; prompts and answers are in English - no other languages are present in the training data.

  • Describing an audio recording in words
  • Answering questions about a sound
  • Identifying the type of sound and the setting
Sizes
99M – 356M
Hardware
from: Laptop
Non-commercial onlyDetails
Computer vision2024–2026

MoGe

Microsoft Research · USA

Reconstructs the 3D geometry of a scene from one photo: depth in meters, a point cloud and surface normals.

  • Measuring rooms and objects from photos
  • 3D point cloud from a single shot
  • Preparing data for robots and AR
Sizes
ViT-S – ViT-G
Hardware
from: Laptop
Commercial use allowedDetails
Forecasting2024–2026

TimesFM

Google · USA

A ready-made Google forecasting model: forecasts any time series without training on your data.

  • Sales and demand forecasting
  • Purchase and inventory planning
  • Load and traffic forecasting
Sizes
200M – 500M
Hardware
from: Laptop
Commercial use with conditionsDetails
Robotics2025–2026

GigaBrain

GigaAI · China

A robot control model trained mostly on synthetic data from a world model. It reduces spending on collecting data from real robots.

  • Controlling a robot arm
  • Fine-tuning on a small amount of your own data
  • Sorting and assembly pilots
Sizes
3.5B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextOllama2024–2026

NVIDIA Nemotron

NVIDIA · USA

NVIDIA models for agents and reasoning, optimized to run fast on its GPUs. Nemotron 3 is a Mamba and MoE hybrid from 4B to 550B; Nano Omni handles video, audio and images (English only).

  • Agents with tool calling
  • Reasoning and calculation tasks
  • Answers based on long documents
Sizes
4B – 550B-A55B
Hardware
from: Laptop
Commercial use with conditionsDetails
Weather and climate2023–2026

Google DeepMind GraphCast / GenCast / WeatherNext 2

Google DeepMind · UK

Google DeepMind's family of global weather models: GraphCast (10-day forecast), GenCast (probabilistic ensemble) and WeatherNext 2 with cyclone forecasting. Since August 2026 the weights are cleared for commercial use.

  • Medium-range weather forecasts for planning shifts, voyages and deliveries
  • Probabilistic assessment of extreme weather for insurance portfolios
  • Tropical cyclone track forecasts for marine and port operations
Sizes
from lightweight 1° versions to full 0.25°
Hardware
from: 1 GPU
Commercial use allowedDetails
Autonomous driving2025–2026

NVIDIA Alpamayo

NVIDIA · USA

Vision-language-action models for self-driving vehicles: they plan a trajectory from camera video and explain the decision in text. Used to develop and test autopilot systems, not as a ready-made autopilot.

  • Auto-labeling camera recordings to train your own driver assistance systems
  • Analyzing complex road scenes with text explanations
  • Testing autopilot systems in simulation on rare scenarios
Sizes
10B – 34B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Autonomous driving2026

Qwen-Drive

Alibaba (Qwen team) · China

An autonomous driving model based on Qwen3.5-4B: 3D detection of objects around the vehicle, answers to questions about the road scene and trajectory planning in one model.

  • A perception and planning prototype for autonomous vehicles on closed sites
  • Answering questions about camera recordings when reviewing incidents
  • Labeling road scenes to train your own models
Sizes
4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2025–2026

NVIDIA Cosmos

NVIDIA · USA

"World" models for robots and self-driving vehicles: they generate realistic video of physical scenes and predict actions. Cosmos 3 combines understanding, generation and control.

  • Synthetic video for training robots and self-driving vehicles
  • Testing scenarios in simulation
  • Robot control (Policy versions)
Sizes
2B – 65B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2026

LingBot-VLA

Ant Group (Robbyant) · China

A robot control model from Ant Group trained on a large volume of data from real robots. Version 2.0 works with different types of robot arms.

  • Controlling a two-armed robot
  • Fine-tuning for your own operation
  • Assembly and sorting pilots
Sizes
4B – 6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Robotics2026

Xiaomi Robotics

Xiaomi · China

Open robot control models from Xiaomi. Robotics-1 is designed for household and kitchen tasks, U0 combines scene understanding and action.

  • Controlling a robot arm by command
  • Household and service scenarios
  • Base for fine-tuning
Sizes
4B – 5B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextGGUF2025–2026

LongCat

Meituan · China

Models from Meituan, China's largest delivery service. LongCat-Flash adjusts compute to query complexity; LongCat-2.0 has 1.6 trillion parameters under MIT. Omni models (Flash-Omni, Next) and AudioDiT speech synthesis too.

  • Agents for orders and service processes
  • Corporate assistant
  • Analysis of long documents
Sizes
1B – 1.6T-A48B
Hardware
from: Laptop
Commercial use allowedDetails
TextOllama2024–2026

EXAONE

LG AI Research · South Korea

Korean-English models from LG. Most of the line is non-commercial, but the flagship K-EXAONE 2.0 with 750 billion parameters is released under Apache 2.0.

  • Corporate assistant
  • Working with Korean and English texts
  • Analysis of documents and images (4.5)
Sizes
1.2B – 750B-A37B
Hardware
from: Laptop
Commercial use with conditionsDetails
Tabular data2026

EXAONE Tabular

LG AI Research · South Korea

LG's small tabular model: with 21M parameters it nearly matches the leaders in classification and regression accuracy. Weights are for non-commercial use only.

  • Pilot churn forecasts
  • Testing scoring hypotheses
  • Exploring customer data
Sizes
about 21M
Hardware
from: Laptop
Non-commercial onlyDetails
Weather and climate2024–2026

Microsoft Aurora

Microsoft Research · USA

A foundation model of Earth's atmosphere: global weather forecasts, plus separate versions for air quality and ocean waves. Computes a forecast in seconds instead of hours on a supercomputer.

  • Your own forecast of temperature, wind and precipitation for company locations
  • Sea state estimates for planning voyages and port operations
  • Air pollution forecasts for industrial sites
Sizes
about 1.3B (a small test version is available)
Hardware
from: 1 GPU
Commercial use allowedDetails
Weather and climate2022–2026

NVIDIA FourCastNet и модели Earth-2

NVIDIA · USA

NVIDIA's set of weather and climate models: global FourCastNet forecasts, downscaling to kilometers (CorrDiff), regional storm forecasts (StormCast), climate generation (cBottle, Atlas). Run via Earth2Studio.

  • Global forecasts followed by downscaling to the region you need
  • Short-term forecasts of thunderstorms and heavy rain for dispatch services
  • Generating many weather scenarios for stress tests
Sizes
98M – 2.5B (FourCastNet 3 — about 711M)
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Computer vision2025–2026

RF-DETR

Roboflow · USA

Real-time object detector, an open alternative to YOLO without AGPL. Supports segmentation (object outlines) and, since 2026, keypoints.

  • Object detection in video and photos
  • Precise outlines of parts and defects
  • Fine-tuning for your own object classes
Sizes
Nano – 2XL
Hardware
from: Laptop
Commercial use allowedDetails
Forecasting2025–2026

TiRex

NXAI · Austria

A compact forecasting model on the xLSTM architecture, a leader in open benchmarks despite its small size. Runs fast on a regular CPU.

  • Demand and sales forecasting
  • Energy consumption forecasting
  • Forecasts on modest hardware and on site
Sizes
about 35M to 82M
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2024–2026

Moondream

Moondream (M87 Labs) · USA

A small, fast vision model for product use cases: answering questions, finding and pointing to objects, captions. Moondream 3.1 is a 9B MoE with 2B active.

  • Finding and counting objects in photos
  • Checking photos from field reports
  • Captions and tags for a catalogue
Sizes
2B – 9B-A2B
Hardware
from: Laptop
Commercial use with conditionsDetails
Documents and OCRRU2022–2026

PP-OCR и PP-DocLayout (классический PaddleOCR)

Baidu (PaddlePaddle) · China

Classic lightweight PaddleOCR models: detecting and recognizing lines of text plus page layout. They run on CPUs and phones; there is a separate model for East Slavic languages, including Russian.

  • Recognizing text on scans, photos and screens
  • Reading labels, displays and markings in production and warehouses
  • Page layout: tables, formulas, stamps, headings
Sizes
from 1.5M to tens of millions of parameters
Hardware
from: Laptop
Commercial use allowedDetails
Satellite and geo2025–2026

Ai2 OlmoEarth

Allen Institute for AI (Ai2) · USA

Ai2's family of models for Sentinel-1, Sentinel-2 and Landsat imagery, with ready-made fine-tunes for mangroves, deforestation and ecosystem types. The license excludes the extractive industries.

  • Monitoring deforestation and forest condition across the supply chain
  • Classifying land and crops from image series
  • Image embeddings for finding similar plots
Sizes
Nano – Large (Base about 114M)
Hardware
from: Laptop
Commercial use with conditionsDetails
3D2024–2026

TripoSR / TripoSG

VAST (TripoSR together with Stability AI) · China

VAST family: a 3D model from a single photo. TripoSR runs in under a second, TripoSG gives cleaner geometry, TripoSplat builds a scene from Gaussian points.

  • 3D product model from a photo
  • Object assets for games and AR
  • Quick 3D prototype for printing
Sizes
up to 1.5B
Hardware
from: Laptop
Commercial use allowedDetails
Robotics2025–2026

MolmoAct

Ai2 (Allen Institute for AI) · USA

A fully open robot control model that first "reasons" about space and trajectory, then acts. Its reasoning can be checked.

  • Controlling a robot arm with explainable steps
  • Fine-tuning for your own robot
  • Research pilots
Sizes
5B – 8B
Hardware
from: 1 GPU
Commercial use allowedDetails
3D2025–2026

VGGT

Meta and the University of Oxford (VGG) · USA / UK

Reconstructs a 3D scene from one, several or hundreds of photos in seconds: camera positions, depth and a point cloud. Best Paper at CVPR 2025.

  • 3D model of a room or object from a photo series
  • Camera pose estimation for photogrammetry
  • Point cloud for measurements and comparison with the plan
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Rerankers2025–2026

Llama Nemotron Rerank

NVIDIA · USA

A small 1B reranker from NVIDIA. The vl version also takes document pages as images, not just text. The card states multilingual support without listing the languages.

  • Reordering passages before an AI assistant answers
  • Sorting retrieved scan and PDF pages
  • Search across internal policies and instructions
Sizes
1B
Hardware
from: Laptop
Commercial use with conditionsDetails
Forecasting2024–2026

Timer (THUML)

THUML, Tsinghua University · China

A compact forecasting foundation model from the Tsinghua lab: trained on a large set of diverse series and fine-tunable on your own data.

  • Forecasting demand and load
  • Forecasting sensor readings on the shop floor
  • Fine-tuning forecasts on your own history
Sizes
84M (timer-base)
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2026

Segment Anything (SAM)

Meta · USA

Selects any object in photos and videos with a click or a box. The basis for background removal and object counting.

  • Background removal from product photos
  • Counting objects in photos
  • Data labeling for training
Sizes
91M – ~0,85B
Hardware
from: Laptop
Commercial use with conditionsDetails
Robotics2025–2026

NVIDIA Isaac GR00T

NVIDIA · USA

NVIDIA's foundation model for humanoid robots and robot arms: it sees, understands a command and outputs movements. Built into the Isaac ecosystem.

  • Controlling humanoid robots
  • Fine-tuning for your own robotic cell
  • Training in simulation with transfer to a real robot
Sizes
2B – 3B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Text2025–2026

Reka Flash и Reka Edge

Reka AI · USA

Compact Reka models: Flash 3 (21B) for reasoning and Reka Edge (7B), which quickly analyzes images and video on-device.

  • Photo and video analysis (Edge)
  • Object detection in images
  • Reasoning tasks (Flash)
Sizes
7B – 21B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image + textGGUF2023–2026

InternVL

Shanghai AI Laboratory (OpenGVLab) · China

A large family of Chinese vision models sized from 1B to 241B. InternVL-U (4B) combines image understanding, generation and editing.

  • Understanding documents, diagrams and charts
  • Answering questions about photos
  • Video analysis
Sizes
1B – 241B-A28B
Hardware
from: Laptop
Commercial use allowedDetails
Image + text2024–2026

Molmo

Ai2 (Allen Institute for AI) · USA

Fully open vision models from Ai2 (weights and data). They can point to a spot in an image and count objects; Molmo2 understands video, MolmoWeb controls a browser.

  • Counting products and objects in photos
  • Pointing to where an item is in an image
  • Video analysis
Sizes
1B-A7B – 72B
Hardware
from: Laptop
Commercial use allowedDetails
Tabular data2025–2026

Orion-MSP / Orion-BiX

Lexsi Labs · India

Recent open models for tabular data: they predict from a few examples given in the prompt, with no task-specific training.

  • Classification and forecasting on tables with no separate training
  • Quickly testing models on new datasets
  • Assessing features in large tables
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Voice assistantsGGUF2025–2026

MiniCPM-o

OpenBMB (ModelBest, Tsinghua University) · China

A small model that sees, hears and replies by voice in real time, and can clone a voice. Voice dialogue in English and Chinese, text in 30+ languages.

  • Voice assistant on your own server
  • Analyzing videos and documents
  • Voice answers about a camera image
Sizes
8B – 9B
Hardware
from: Laptop
Commercial use allowedDetails
Tabular data2025–2026

TabICL

Inria (Soda team) · France

An open tabular model from the creators of scikit-learn: classifies and predicts from examples without training and handles tables of up to hundreds of thousands of rows. The license allows business use.

  • Predicting customer churn
  • Scoring applications and deals
  • Classifying customers from 1C and CRM data
Sizes
about 25–30M
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2026

YOLO (Ultralytics)

Ultralytics · USA

The most widely used real-time object detector: finds and marks items in video even on modest hardware. YOLOv5 came out back in 2020; the catalog starts from YOLOv8.

  • Counting people, cars and goods on video
  • Checking hard hats and workwear
  • Spotting defects on the production line
Sizes
2.4M – 68M
Hardware
from: Laptop
Commercial use with conditionsDetails
TextRUOllama2023–2026

Falcon

Technology Innovation Institute (TII) · UAE

A family from Abu Dhabi: from the early Falcon 40B and 180B to hybrid Falcon-H1 and tiny Falcon-H1-Tiny models of 90–600M parameters for devices.

  • Assistant and answers based on documents
  • Running on low-end hardware and devices
  • Tool calling in simple agents
Sizes
90M – 180B
Hardware
from: Laptop
Commercial use with conditionsDetails
Weather and climate2024–2026

ECMWF AIFS

European Centre for Medium-Range Weather Forecasts (ECMWF) · Europe (intergovernmental organization)

ECMWF's weather neural network running operationally: a 15-day forecast four times a day, an ensemble version with 51 scenarios, and since version 2, ocean waves.

  • Running your own forecast from open initial data
  • Ensemble forecasts to estimate the probability of frost, downpours and storms
  • Wave forecasts for marine operations
Sizes
checkpoint of about 1 GB
Hardware
from: 1 GPU
Commercial use allowedDetails
3DGGUF2024–2025

TRELLIS

Microsoft · USA

One of the strongest open 3D models: from an image or text it produces a textured mesh or a Gaussian scene. TRELLIS.2 is noticeably more detailed than the first version.

  • 3D models of products and interiors from photos
  • Assets for games and AR/VR
  • Prototypes for 3D printing
Sizes
up to 4B (TRELLIS.2)
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer visionGGUF2024–2025

Depth Anything

ByteDance and the University of Hong Kong · China

Estimates depth, the distance to every point, from one ordinary photo or video. DA3 reconstructs scene geometry from several frames.

  • Estimating distances and volumes from a camera
  • Depth effects for photo and video
  • Navigation for robots and drones
Sizes
25M – 1.4B
Hardware
from: Laptop
Commercial use with conditionsDetails
3D2025

MapAnything

Meta and Carnegie Mellon University · USA

A single model builds a metric 3D reconstruction from photos, and uses camera, depth or pose data when available. One weights variant is under Apache 2.0.

  • 3D reconstruction of an object or room from photos
  • Exporting the scene to COLMAP format for further processing
  • Depth and camera pose estimation
Sizes
about 1.2B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer vision2025

Perception Encoder (PE)

Meta · USA

Meta's family of encoders for images and video, and with PE-AV also for audio. PE-Core searches by text more accurately than SigLIP 2 (per Meta); small versions are available.

  • Search photos and videos by description
  • Catalog labeling and tagging
  • Search across audio and video (PE-AV)
Sizes
size not stated on the model card
Hardware
from: Laptop
Commercial use allowedDetails
Tabular dataGGUF2024–2025

TableGPT2 / TableGPT-R1

Zhejiang University · China

A family for working with tables and databases: it understands data structure, writes parsing code and answers questions about exports.

  • Answering questions about tables and data exports
  • Automated data analysis with generated code
  • A helper for BI and internal reporting
Sizes
7B – 72B
Hardware
from: 1 GPU
Commercial use allowedDetails
Satellite and geo2025

IBM–ESA TerraMind

IBM and the European Space Agency (ESA) · USA / Europe

A multimodal Earth model: understands optical and radar imagery, terrain, vegetation index and land use maps, and can generate a missing data type (for example, a "see-through-clouds" image from radar).

  • Analyzing fields and forests even in cloudy weather using radar imagery
  • Land use maps for assessing plots
  • Flood and wildfire assessment (ready-made fine-tunes available)
Sizes
tiny – large (checkpoints from ~200 MB to ~3.8 GB)
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2025

Grounding DINO / Rex-Omni

IDEA Research · China

Finds any objects in an image from a text description, without training on your data: "red box", "person without a hard hat". Rex-Omni is the new VLM-based generation.

  • Finding objects by description without labeling
  • Automatic data labeling for training
  • Checking photos against requirements
Sizes
172M – 3B
Hardware
from: Laptop
Commercial use with conditionsDetails
ForecastingGGUF2024–2025

Chronos

Amazon · USA

Amazon forecasting models, among the most downloaded. Chronos-2 takes external factors into account: prices, promotions, weather.

  • Demand forecasting with promotions and prices
  • Inventory planning
  • Forecasting revenue and customer flow
Sizes
8M – 710M
Hardware
from: Laptop
Commercial use allowedDetails
Image + textOllama2023–2025

Qwen-VL

Alibaba (Qwen team) · China

One of the strongest open vision models: reads documents, tables, charts and video, and works with user interfaces. Since Qwen3.5, vision is built directly into the main Qwen model.

  • Extracting data from scanned invoices and delivery notes
  • Analysing photos of products and shelves
  • Analysing video and camera footage
Sizes
2B – 235B-A22B
Hardware
from: Laptop
Commercial use allowedDetails
Robotics2025

π0 / π0.5 (openpi)

Physical Intelligence · USA

Robot control models from Physical Intelligence: folding laundry, tidying up, handling objects. π0.5 copes better in unfamiliar settings.

  • Controlling robot arms and two-armed robots
  • Fine-tuning for your own operations
  • Pilots for automating manual work
Sizes
about 3B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Satellite and geo2023–2025

IBM–NASA Prithvi

IBM and NASA · USA

Foundation models for Landsat and Sentinel-2 satellite imagery that account for image time series. Ready-made fine-tunes for floods, burn scars and crop types, plus a separate weather model, WxC.

  • Mapping crops and field condition over the season
  • Assessing flood zones and burn scars after natural disasters
  • Monitoring changes in buildings and land use
Sizes
tiny – 600M (imagery), 2.3B (Prithvi WxC weather)
Hardware
from: Laptop
Commercial use allowedDetails
Computer vision2023–2025

DINOv2 / DINOv3

Meta · USA

Foundation models that turn an image into a numeric "fingerprint". They are used to build similar-image search, classification and segmentation without large labeled datasets.

  • Finding similar products and photos
  • Image classification on small datasets
  • Base for your own quality-control models
Sizes
21M – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
ForecastingGGUF2024–2025

Moirai

Salesforce · USA

Salesforce's universal forecasting model for series with different frequencies and many variables. Weights are open for research only.

  • Research forecasting pilots
  • Comparison with other forecasting models
  • Forecasts across many related series
Sizes
11M – 311M
Hardware
from: Laptop
Non-commercial onlyDetails
Robotics2025

SmolVLA

Hugging Face · USA

A small robot control model that runs on a regular laptop. Trained on open data from the LeRobot community, suited to low-cost robot arms.

  • Controlling a low-cost robot arm
  • Quick robotization pilots and demos
  • Training staff and students
Sizes
450M
Hardware
from: Laptop
Commercial use allowedDetails
Satellite and geo2025

MBZUAI TerraFM

MBZUAI · UAE

A compact research model for satellite imagery, trained on both optical (Sentinel-2) and radar (Sentinel-1) data. Narrower in scope and community than Prithvi and TerraMind.

  • Classification and segmentation of satellite imagery after fine-tuning
  • A base for a land monitoring prototype
Sizes
TerraFM-B (ViT-Base)
Hardware
from: Laptop
Commercial use allowedDetails
Forecasting2025

Sundial

THUML, Tsinghua University · China

A forecasting model that returns a set of possible scenarios rather than a single line — useful when you need a range for demand or load, not one number.

  • Forecasting demand with a range of values
  • Planning stock while accounting for spread
  • Forecasting load on services and staff
Sizes
128M (sundial-base)
Hardware
from: Laptop
Commercial use allowedDetails
Robotics2024–2025

OpenVLA

Stanford, Berkeley and partners · USA

The first large open vision-language-action model: a robot arm carries out commands like "put the apple in the bowl". OFT makes it several times faster.

  • Controlling a robot arm by text command
  • Pilots for robotizing simple operations
  • Base for fine-tuning to your own robot
Sizes
7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer-use agentsGGUF2025

Magma

Microsoft Research · USA

An agent model that plans actions both in an interface (buttons on screen) and for a robot (arm movements). For now more of a research base than a finished product.

  • Pilots in interface control
  • Research projects spanning screens and robotics
  • Analyzing screenshots with an action plan
Sizes
8B
Hardware
from: 1 GPU
Commercial use allowedDetails
Computer vision2022–2025

ViTPose

University of Sydney and JD Explore Academy · Australia / China

A simple, accurate model for human pose estimation via keypoints. ViTPose++ handles human, animal and whole-body poses; built into the Transformers library.

  • Body keypoints in photos and video
  • Motion analysis in sports and rehabilitation
  • Monitoring work postures and safety practices
Sizes
33M – about 1B
Hardware
from: Laptop
Commercial use allowedDetails
Image + textGGUF2024

PaliGemma

Google · USA

Google's vision model built on Gemma, designed as a base for fine-tuning on a narrow task: captions, object detection, reading text.

  • Fine-tuning for your own recognition task
  • Finding objects in photos
  • Reading text in images
Sizes
3B – 28B
Hardware
from: Laptop
Commercial use with conditionsDetails
Forecasting2024

MOMENT

Auton Lab, Carnegie Mellon University · USA

A foundation model for numeric series: one engine is used for forecasting, anomaly detection, filling gaps and classification.

  • Forecasting demand and load
  • Detecting anomalies in sensor readings and metrics
  • Filling gaps in historical data
Sizes
about 40M – 385M
Hardware
from: Laptop
Commercial use allowedDetails
Forecasting2024

Granite Time Series (TinyTimeMixers, PatchTST)

IBM Research · USA

Tiny forecasting models from IBM: they run on an ordinary CPU and sit next to the business system without a separate GPU server.

  • Forecasting sales and warehouse stock
  • Forecasting energy use and equipment load
  • Fast forecasts right on the company server
Sizes
very small: TinyTimeMixers have about 1M parameters
Hardware
from: Laptop
Commercial use allowedDetails
Tabular dataGGUF2024

TableLLM

RUCKBReasoning, Renmin University of China · China

A model for office work with tables: for a given question it returns either a direct answer or code to process the data in a table or document.

  • Processing tables from Excel and documents from a text instruction
  • Generating code for recalculations and selections
  • Answering questions about data in reports
Sizes
7B и 13B
Hardware
from: Laptop
Commercial use with conditionsDetails
Forecasting2024

Time-MoE

The Time-MoE team · not disclosed

A forecasting model with a sparse architecture: only part of the network runs at each step, so it stays fast at a small size.

  • Forecasting sales and stock levels
  • Forecasting load on services and staff
  • Planning purchases from history
Sizes
50M and 200M
Hardware
from: Laptop
Commercial use allowedDetails
Tabular dataNot maintained2024

TabuLa-8B

ML Foundations · USA

A foundation model for predictions on tables: it classifies and forecasts from a handful of examples, with no separate task-specific training.

  • Classifying table rows from a few examples
  • Predicting a value from a data row
  • Quickly testing hypotheses on new datasets
Sizes
8B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image + textNot maintained2024

Florence-2

Microsoft · USA

A very small vision model: captions, object detection, segmentation and text reading from a single prompt. Runs even on a CPU.

  • Reading text in photos
  • Finding and highlighting objects
  • Automatic photo captions
Sizes
0.23B – 0.77B
Hardware
from: Laptop
Commercial use allowedDetails
Satellite and geoNot maintained2023–2024

Ai2 SatlasPretrain

Allen Institute for AI (Ai2) · USA

Pretrained models from the Satlas project for Sentinel-2, Landsat and high-resolution aerial imagery. The predecessor of OlmoEarth, still used in TorchGeo.

  • Detecting objects in imagery: solar farms, wind turbines, ships
  • Mapping roads and buildings from aerial photos
  • A starting point for fine-tuning your own geo model
Sizes
Swin-v2 and ResNet backbones (Base)
Hardware
from: Laptop
Commercial use allowedDetails
Music and soundNot maintained2022

AST (Audio Spectrogram Transformer)

MIT · USA

A classic 2021 sound recognition model: detects 527 AudioSet event classes (siren, barking, breaking glass, music). Lightweight, runs without a GPU, in Transformers since 2022.

  • Sound event recognition
  • Tagging an audio archive
  • Detecting alarm sounds
Sizes
about 87M
Hardware
from: Laptop
Commercial use allowedDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment