What needs several GPUs

Several GPUs are needed by flagship models, the ones comparable to the best cloud services. It is an expensive setup, and it makes sense only after a smaller model has been tested on your tasks and fell short.

7 familiesUpdated 22 Sep 2026Calculate video memory for your model size

What this means in practice

A 2-4 GPU serverLarge language models and sparse architectures. Usually rented rather than bought.
An 8 GPU serverThe largest open models and long video generation. Data-centre territory.
Hourly rentalThe sensible first step: rent a cluster while you test the idea instead of buying hardware for an unproven task.

What runs on it

Text 3

TextGGUF2026

Inkling

Thinking Machines Lab · USA

Flagship open models from Mira Murati's lab: they take text, images and audio. Large MoE models that need several GPUs.

  • Flagship-level corporate assistant
  • Analysis of documents, images and audio
  • Programming help
Sizes
276B-A12B, 975B-A41B
Hardware
from: Cluster
Commercial use allowedDetails
TextRUGGUF2025–2026

MiniMax

MiniMax · China

Large MoE models with very long context (up to 1M tokens for Text-01 and M3). M3 is multimodal and understands images. Licenses differ greatly from version to version.

  • Analysis of large document archives in a single request
  • Agents with tools
  • Help for developers
Sizes
230B-A10B – 456B-A46B
Hardware
from: Cluster
Commercial use with conditionsDetails
Text2024–2025

Grok (открытые веса)

xAI · USA

xAI publishes the weights of previous Grok generations. The models are very large and need a GPU cluster, so in practice they are rarely run.

  • Research on large models
  • Assistant on your own infrastructure
  • Text generation and analysis
Sizes
314B (Grok-1), Grok-2 is larger
Hardware
from: Cluster
Commercial use with conditionsDetails

Video 2

Video2026

MiniMax H3 (Hailuo)

MiniMax · China

Open weights of MiniMax's Hailuo video model. A large 33B model that makes video from text and images, but needs several server GPUs.

  • Cinematic ad videos
  • Video from text and images
  • Complex scenes with motion
Sizes
33B + 32B encoder
Hardware
from: Cluster
Commercial use with conditionsDetails
Video2025

Step-Video

StepFun · China

A large 30B video model producing clips of up to 204 frames. Needs server hardware, but is open under MIT.

  • Video from a description
  • Animating images
Sizes
30B
Hardware
from: Cluster
Commercial use allowedDetails

Finance 1

FinanceNot maintained2024

Palmyra-Fin

Writer · USA

A large model for financial documents with a context window of about 131k tokens: it holds long reports whole. Not investment advice: decisions are made by a specialist.

  • Working with long annual reports and prospectuses
  • Summaries and digests of financial documents
  • Finding answers inside a large document pack
Sizes
70B (the model card states 72 billion parameters)
Hardware
from: Cluster
Commercial use with conditionsDetails

Fact-checking and judges 1

Fact-checking and judgesRU2024–2025

Nemotron Reward / GenRM

NVIDIA · USA

Large NVIDIA scorers for selecting and fine-tuning answers. The multilingual GenRM version lists Russian among its languages. The scorer itself makes mistakes and does not replace manual review on important tasks.

  • Choosing the best of several candidate answers
  • Preparing data to fine-tune your own model
  • Scoring assistant answers in Russian and other languages
Sizes
49B, 70B and 340B
Hardware
from: Cluster
Commercial use with conditionsDetails

Where people usually hit the wall

The usual mistake is to start with the flagship. In practice it goes the other way: a small model on one GPU, quality measured on your own examples, and only then a move up. A cluster adds more than hardware cost: networking, cooling, updates and someone on call.

By hardware

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment