Image generationGGUF2025–2026
Alibaba · China
Image generation and editing, including text in images. Earlier versions allow commercial use; the latest 2.1 is non-commercial only.
- Infographics for product cards
- Photo editing by text command
- Ad creatives
- Sizes
- 7B – 20B
- Hardware
- from: 1 GPU
Image generationRU2022–2026
Sber (Kandinsky Lab) · Russia
Sber's Russian family of image and video generation models. Understands Russian-language prompts and Russian cultural context well; released under MIT.
- Images from Russian-language descriptions
- Short promo videos from text or a photo
- Instruction-based image editing
- Sizes
- 2B – 19B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
lodestones (independent developer) · not disclosed
A community model retrained from FLUX.1-schnell with a simplified architecture. No style censorship; popular as a base for fine-tuning.
- Base for fine-tuning your own styles
- Artistic illustrations
- Images for games
- Sizes
- 4B – 8.9B
- Hardware
- from: 1 GPU
Image generationGGUF2026
Krea · USA
A 12B image model focused on realism without the glossy "AI look". The Turbo version produces a 2K image in a couple of seconds.
- Realistic photos for advertising
- High-resolution images
- Style fine-tuning
- Sizes
- 12B
- Hardware
- from: 1 GPU
Image generationGGUF2026
Ideogram · Canada
Open weights of the Ideogram model, known for precise typography. Under a non-commercial license: for business, suitable only for testing.
- Testing text-in-image generation
- Research and prototypes
- Sizes
- about 9B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
HiDream.ai · China
Open MIT-licensed image models: generation (I1), instruction-based editing (E1) and the unified O1-Image model that does both.
- Image generation from descriptions
- Editing images with words
- Variations of product photos
- Sizes
- about 9B – 17B
- Hardware
- from: 1 GPU
TextGGUF2025–2026
Baidu · China
Baidu's first open line: from a tiny 0.3B to MoE with 424 billion parameters, including versions that understand images. The mid-size 21B-A3B fits on one GPU; ERNIE-Image 8B draws images with text.
- Corporate assistant
- Analysis of documents and images
- Customer request classification
- Sizes
- 0.3B – 424B-A47B
- Hardware
- from: Laptop
Image generationGGUF2025–2026
Meituan · China
Meituan's 6B image generation and editing model. Renders Chinese text well; has a fast version for edits.
- Image generation from descriptions
- Instruction-based photo editing
- Visuals for product cards
- Sizes
- 6B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2026
Black Forest Labs · Germany
Image generation from the creators of Stable Diffusion. Renders text in images well and keeps the composition.
- Images for product cards
- Banners and covers
- Photo editing by description (Kontext)
- Sizes
- 4B – 32B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2026
Tencent · China
Tencent's image models. HunyuanImage 3.0 is the largest open MoE generation model at 80B; it can reason about the prompt and edit by instruction.
- Complex scenes from long descriptions
- Images with Chinese and English text
- Instruction-based image editing
- Sizes
- 1.5B – 80B-A13B
- Hardware
- from: 1 GPU
Image generationGGUF2025–2026
Alibaba (Tongyi-MAI) · China
A compact 6B model with photorealism on par with large models. The Turbo version produces an image in a few steps on a regular gaming GPU.
- Photorealistic ad images
- Images with English and Chinese text
- Bulk visual generation
- Sizes
- 6B
- Hardware
- from: 1 GPU
Image generation2026
Zhipu AI (Z.ai) · China
A hybrid of a 9B language model and a 7B decoder. Strong at text-heavy images: posters, infographics, slides.
- Posters and banners with text
- Infographics
- Illustrations for presentations
- Sizes
- 9B + 7B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2025
BAAI (Beijing Academy of Artificial Intelligence) · China
An all-in-one model: generates, edits and moves an object or person from a photo into a new scene without separate plugins.
- Placing a product or person into a new scene
- Instruction-based photo editing
- Generation from multiple references
- Sizes
- about 4B
- Hardware
- from: 1 GPU
Image generationGGUF2025
ByteDance Seed · China
A unified model that understands images, generates them and edits them in a conversation. Similar to how images work in ChatGPT.
- Photo editing in a conversation
- Answering questions about an image
- Image generation with explanations
- Sizes
- 14B-A7B
- Hardware
- from: 1 GPU
Image generationGGUF2024–2025
NVIDIA · USA
NVIDIA's fast image model: 4K images in seconds, runs even on a laptop GPU. The Sprint version generates in 1–2 steps.
- Bulk image generation
- High-resolution visuals
- Real-time generation inside apps
- Sizes
- 0.6B – 4.8B
- Hardware
- from: Laptop
Faces2025
ByteDance · China
FLUX-based image generation that preserves a face: follows the prompt better and less often pastes the face like a sticker. Research-only license.
- Portraits from one photo with a precise scene description
- Testing characters for advertising
- Comparing face-preservation methods
- Sizes
- adapter for FLUX.1-dev
- Hardware
- from: 1 GPU
Image + textGGUF2024–2025
DeepSeek · China
A single model that both understands images and draws them from a description. Janus-Pro-7B drew attention in early 2025, but its image quality is below specialised models.
- Answering questions about images
- Draft illustrations from a description
- Experiments with a unified vision and generation model
- Sizes
- 1B – 7B
- Hardware
- from: Laptop
Image generationGGUF2022–2024
Stability AI · UK
The model that started open image generation. A huge ecosystem of fine-tunes, styles and plugins; runs even on a home PC. The popular SDXL-Lightning and Hyper-SD accelerators were made by ByteDance.
- Illustrations and banners for advertising
- Backgrounds and scenes for product cards
- Fine-tuning to a brand style
- Sizes
- 0.9B – 8B
- Hardware
- from: Laptop
Faces2024
ByteDance · China
Preserves a person's face when generating images from one photo, with less damage to style and background. Versions exist for SDXL and FLUX; the latter runs on a 16 GB card.
- Portraits and avatars from one photo
- Ad characters with a recognizable face
- Photoshoot prototypes
- Sizes
- adapters for SDXL and FLUX.1-dev
- Hardware
- from: 1 GPU
Image generationNot maintained2023–2024
Lvmin Zhang (Stanford) and the community · USA
An add-on for image models: sets pose, outlines, depth or floor plan so the result follows the required composition exactly.
- Image from a sketch or outline
- Keeping pose and composition
- Interior visualization from a floor plan
- Sizes
- 0.4B – 1.3B
- Hardware
- from: Laptop
Photo editingNot maintained2023–2024
Shanghai AI Laboratory (OpenMMLab) and Tsinghua University · China
All-round photo inpainting: remove an object, insert a new one from a description, change a shape or extend the frame beyond its edges.
- Removing and replacing objects in photos
- Extending the frame to a required format
- Inserting a product or detail from a text description
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Photo editingNot maintained2024
Lvmin Zhang (author of ControlNet) · USA
Changes lighting in a photo: relights an object or person from a description or to match a given background, so a cut-out looks natural.
- Matching product lighting to a new background
- Studio lighting for portraits without a reshoot
- Consistent lighting style across a catalog
- Sizes
- based on SD 1.5
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Huawei Noah's Ark Lab and partners · China
A compact 0.6B image model with quality on par with much larger ones. The Sigma version does 4K; suits modest hardware.
- Illustrations for articles and social media
- Backgrounds for product cards
- Quick visual drafts
- Sizes
- 0.6B
- Hardware
- from: Laptop
FacesNot maintained2023–2024
Tencent AI Lab (h94) · China
One of the first adapters that transfer a face from a photo into a generated image. The SD 1.5 versions run on low-end cards, but the weights are non-commercial.
- Portraits from a photo in different styles
- Image series with one character
- Avatar experiments
- Sizes
- adapters for SD 1.5 and SDXL
- Hardware
- from: Laptop
Image generationNot maintained2023–2024
Playground AI · USA
An SDXL-based model focused on aesthetics: vivid colors, contrast, portraits. Compatible with SDXL ecosystem tools.
- Aesthetic ad visuals
- Portraits and lifestyle images
- Post covers
- Sizes
- about 2.6B
- Hardware
- from: 1 GPU
FacesNot maintained2024
InstantX (Xiaohongshu) · China
Generates images with a specific person's face from a single photo, without fine-tuning. Popular in ComfyUI, but the weights are for research only.
- Portraits in different styles from one photo
- Avatar and character sketches
- Photoshoot prototypes
- Sizes
- adapter for SDXL
- Hardware
- from: 1 GPU
Image generationNot maintained2023
DeepFloyd (Stability AI) · UK
An early model that was among the first to render text on images accurately. Today it is mostly of historical interest; development has stopped.
- Research experiments
- Prototype images with captions
- Sizes
- 0.4B – 4.3B
- Hardware
- from: 1 GPU