Open-source AI models for image generation

Generative models create images from text prompts: banners, illustrations, social media and product visuals. Hosting your own model gives predictable costs and control over style. Read the license carefully, since some models restrict commercial use, and account for GPU memory requirements.

27 open model families in this collection.Updated 22 Sep 2026Open the full catalog with filters
Image generationGGUF2025–2026

Qwen-Image

Alibaba · China

Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.

  • Infographics for product cards
  • Photo editing by text command
  • Ad creatives
Sizes
7B – 20B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationRU2022–2026

Kandinsky

Sber (Kandinsky Lab) · Russia

Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.

  • Images from Russian-language descriptions
  • Short promo videos from text or a photo
  • Instruction-based image editing
Sizes
2B – 19B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025–2026

Chroma

lodestones (independent developer) · not disclosed

A community model retrained from FLUX.1-schnell with a simplified architecture. No style censorship; popular as a base for fine-tuning.

  • Base for fine-tuning your own styles
  • Artistic illustrations
  • Images for games
Sizes
4B – 8.9B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2026

Krea 2

Krea · USA

A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.

  • Realistic photos for advertising
  • High-resolution images
  • Style fine-tuning
Sizes
12B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2026

Ideogram 4

Ideogram · Canada

Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.

  • Testing text-in-image generation
  • Research and prototypes
Sizes
about 9B
Hardware
from: 1 GPU
Non-commercial onlyDetails
Image generationGGUF2025–2026

HiDream

HiDream.ai · China

Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.

  • Image generation from descriptions
  • Editing images with words
  • Variations of product photos
Sizes
about 9B – 17B
Hardware
from: 1 GPU
Commercial use allowedDetails
TextGGUF2025–2026

ERNIE 4.5

Baidu · China

Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.

  • Corporate assistant
  • Analysis of documents and images
  • Customer request classification
Sizes
0.3B – 424B-A47B
Hardware
from: Laptop
Commercial use allowedDetails
Image generationGGUF2025–2026

LongCat-Image

Meituan · China

Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.

  • Image generation from descriptions
  • Instruction-based photo editing
  • Visuals for product cards
Sizes
6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2024–2026

FLUX

Black Forest Labs · Germany

Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.

  • Images for product cards
  • Banners and covers
  • Photo editing by description (Kontext)
Sizes
4B – 32B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2024–2026

Hunyuan Image

Tencent · China

Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.

  • Complex scenes from long descriptions
  • Images with Chinese and English text
  • Instruction-based image editing
Sizes
1.5B – 80B-A13B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationGGUF2025–2026

Z-Image

Alibaba (Tongyi-MAI) · China

A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.

  • Photorealistic ad images
  • Images with English and Chinese text
  • Bulk visual generation
Sizes
6B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generation2026

GLM-Image

Zhipu AI (Z.ai) · China

A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.

  • Posters and banners with text
  • Infographics
  • Illustrations for presentations
Sizes
9B + 7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2024–2025

OmniGen

BAAI (Beijing Academy of Artificial Intelligence) · China

An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.

  • Placing a product or person into a new scene
  • Instruction-based photo editing
  • Generation from multiple references
Sizes
about 4B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2025

BAGEL

ByteDance Seed · China

A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.

  • Photo editing in a conversation
  • Answering questions about an image
  • Image generation with explanations
Sizes
14B-A7B
Hardware
from: 1 GPU
Commercial use allowedDetails
Image generationGGUF2024–2025

SANA

NVIDIA · USA

NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.

  • Bulk image generation
  • High-resolution visuals
  • Real-time generation inside apps
Sizes
0.6B – 4.8B
Hardware
from: Laptop
Commercial use allowedDetails
Faces2025

InfiniteYou

ByteDance · China

FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.

  • Portraits from one photo with a precise scene description
  • Testing characters for advertising
  • Comparing face-preservation methods
Sizes
adapter for FLUX.1-dev
Hardware
from: 1 GPU
Non-commercial onlyDetails
Image + textGGUF2024–2025

Janus

DeepSeek · China

A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.

  • Answering questions about images
  • Draft illustrations from a description
  • Experiments with a unified vision and generation model
Sizes
1B – 7B
Hardware
from: Laptop
Commercial use with conditionsDetails
Image generationGGUF2022–2024

Stable Diffusion

Stability AI · UK

The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.

  • Illustrations and banners for advertising
  • Backgrounds and scenes for product cards
  • Fine-tuning to a brand style
Sizes
0.9B – 8B
Hardware
from: Laptop
Commercial use with conditionsDetails
Faces2024

PuLID

ByteDance · China

Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.

  • Portraits and avatars from one photo
  • Ad characters with a recognizable face
  • Photoshoot prototypes
Sizes
adapters for SDXL and FLUX.1-dev
Hardware
from: 1 GPU
Commercial use with conditionsDetails
Image generationNot maintained2023–2024

ControlNet

Lvmin Zhang (Stanford) and the community · USA

An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.

  • Image from a sketch or outline
  • Keeping pose and composition
  • Interior visualization from a floor plan
Sizes
0.4B – 1.3B
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2023–2024

PowerPaint

Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China

All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.

  • Removing and replacing objects in photos
  • Extending the frame to a required format
  • Inserting a product or detail from a text description
Sizes
based on SD 1.5
Hardware
from: Laptop
Commercial use allowedDetails
Photo editingNot maintained2024

IC-Light

Lvmin Zhang (author of ControlNet) · USA

Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.

  • Matching product lighting to a new background
  • Studio lighting for portraits without a reshoot
  • Consistent lighting style across a catalog
Sizes
based on SD 1.5
Hardware
from: Laptop
Commercial use allowedDetails
Image generationNot maintained2023–2024

PixArt

Huawei Noah's Ark Lab and partners · China

A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.

  • Illustrations for articles and social media
  • Backgrounds for product cards
  • Quick visual drafts
Sizes
0.6B
Hardware
from: Laptop
Commercial use allowedDetails
FacesNot maintained2023–2024

IP-Adapter-FaceID

Tencent AI Lab (h94) · China

One of the first adapters that transfer a face from a photo into a generated image. The SD 1.5 versions run on low-end cards, but the weights are non-commercial.

  • Portraits from a photo in different styles
  • Image series with one character
  • Avatar experiments
Sizes
adapters for SD 1.5 and SDXL
Hardware
from: Laptop
Non-commercial onlyDetails
Image generationNot maintained2023–2024

Playground

Playground AI · USA

An SDXL-based model focused on aesthetics: vivid colors, contrast, portraits. Compatible with SDXL ecosystem tools.

  • Aesthetic ad visuals
  • Portraits and lifestyle images
  • Post covers
Sizes
about 2.6B
Hardware
from: 1 GPU
Commercial use with conditionsDetails
FacesNot maintained2024

InstantID

InstantX (Xiaohongshu) · China

Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.

  • Portraits in different styles from one photo
  • Avatar and character sketches
  • Photoshoot prototypes
Sizes
adapter for SDXL
Hardware
from: 1 GPU
Non-commercial onlyDetails
Image generationNot maintained2023

DeepFloyd IF

DeepFloyd (Stability AI) · UK

An early model that was among the first to render text on images accurately. Today it is mostly of historical interest; development has stopped.

  • Research experiments
  • Prototype images with captions
Sizes
0.4B – 4.3B
Hardware
from: 1 GPU
Non-commercial onlyDetails

Collections

Need a model for your task?

An open model can run on your own server: data stays in-house, there is no per-request fee, and the model can be fine-tuned on your documents.

  1. SelectThe model and size for your task and hardware budget
  2. DeployOn your server or in a closed network, with an API
  3. Fine-tuneOn your data, or connect a knowledge base
  4. IntegrateInto your CRM, ERP, bot, website or team chat
Discuss deployment