Z-Image: The New Powerhouse for Speed and Creation

A powerful 6B parameter image model offering sub-second inference, bilingual text rendering, and robust editing capabilities across three specialized checkpoints.

High-Speed Architecture

Efficient 6B Parameter Model
Delivers sub-second inference latency on enterprise GPUs and fits comfortably on 16GB consumer cards using only 8 steps (NFEs).

Advanced Capabilities

Bilingual Text & Photorealism
Generates stunning photorealistic images with distinct capabilities in rendering accurate English and Chinese text directly within the visual.

Specialized Variants

Turbo, Base, and Edit Modes
Three tailored solutions in one family: a distilled speed model, a foundation for fine-tuning, and a dedicated variant for natural language image editing.

Choose Your Checkpoint: Turbo, Base, vs. Edit

A side-by-side comparison of the Z-Image family. Explore the technical differences between our distilled high-speed inference model, the open foundation for fine-tuning, and the specialized variant for instruction-based editing.

Z-Image-Turbo

A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers sub-second inference latency on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices. 

Z-Image-Base

The non-distilled foundation model. By releasing this checkpoint, we aim to unlock the full potential for community-driven fine-tuning and custom development.

Z-Image-Edit

A variant fine-tuned on Z-Image specifically for image editing tasks. It supports creative image-to-image generation with impressive instruction-following capabilities

The New KING of LOCAL Image Models

See how this lightning-fast, locally-run AI image generator stacks up. Bijan Bowen tests Z-Image Turbo, a 6-billion parameter model, using diverse prompts. Discover if this open-source model truly reigns supreme.

Real-World Applications of Z-Image

Discover how the Z-Image suite transforms diverse industries. From sub-second generation in live apps to precision editing and custom model training, explore the four key scenarios where our architecture excels.

Powering Chatbots & Live Generators

Leverage Z-Image-Turbo to integrate image generation into consumer-facing apps or chatbots. With sub-second inference latency and low memory requirements (16GB VRAM).

You can deliver instant visual responses to user queries without the need for massive server clusters.

Creating Bilingual Visuals with Perfect Text

Use Z-Image-Turbo for designing posters, social media banners, and e-commerce visuals that require embedded text. Its ability to accurately render both English and Chinese characters ensures your typography is legible and perfectly integrated into photorealistic scenes.

Streamlining Edits with Natural Language

Deploy Z-Image-Edit for tasks that require modifying existing assets rather than starting from scratch. whether changing a background, swapping an object, or altering the lighting, designers can execute precise edits using simple text instructions, significantly reducing manual photoshop time.

Training Custom Models on Proprietary Data

Utilize Z-Image-Base when you need a model tailored to a specific style, brand identity, or medical/scientific dataset. As a non-distilled foundation, it serves as the perfect starting point for community fine-tuning, allowing developers to build niche tools on a powerful 6B parameter architecture.

Frequently Asked Questions (FAQ)

What are the hardware requirements to run Z-Image locally?

Z-Image is designed to be accessible. It fits comfortably within 16GB VRAM, making it compatible with many high-end consumer graphics cards (e.g., NVIDIA RTX 3090/4080/4090).

How fast is the model?

Z-Image-Turbo is capable of sub-second inference latency when running on enterprise-grade hardware like H800 GPUs. On consumer hardware, it remains exceptionally fast due to its optimized 8-step (NFE) generation process.

Does Z-Image support text rendering?

Yes. Z-Image excels at bilingual text rendering, capable of accurately generating legible text in both English and Chinese within images, making it ideal for poster design and marketing materials.

What is “Instruction Adherence”?

This refers to how well the model listens to your prompt. Z-Image-Turbo and Edit are tuned for robust instruction adherence, meaning they faithfully include the specific details, styles, and text strings you ask for, rather than ignoring parts of the prompt.

What makes Z-Image different from other models?

Z-Image stands out due to its balance of size and speed. With 6B parameters, it offers high fidelity, yet the Turbo variant requires only 8 Function Evaluations (NFEs), allowing for sub-second inference on high-end hardware and efficient performance on consumer GPUs.

Stay Updated on Z-Image Development

Be the first to know about new model releases, optimized checkpoints, and fine-tuning guides. Subscribe now to get the latest open-source AI news delivered straight to your inbox.