z-image-txt2img

Installation
SKILL.md

Z-Image Text-to-Image Workflows

Launch flag. Z-Image does not sample correctly under --use-sage-attention (black / garbled output). Launch ComfyUI with --use-pytorch-cross-attention for Z-Image. See comfyui-launch-flags.

Overview

Z-Image is a 6B-parameter image generation model from Alibaba's Tongyi Lab using a Scalable Single-Stream DiT (S3-DiT) architecture. It uses a Qwen text encoder (not CLIP-L/T5). Its VAE shares the Flux VAE architecture (same tensor shapes, so the file is the same 320MB size) but ships different weights. It is NOT byte-identical to Flux's ae.safetensors and must be kept as a separate file (z-image-ae.safetensors) to avoid clobbering the Flux VAE. Two variants:

  1. Z-Image Base (and RedCraft finetune). Full model, supports negative prompts, LoRA training, ControlNet. 10-30 steps.
  2. Z-Image Turbo. DMD-distilled, 8-10 steps, no effective negative prompts (CFG baked in).

Models

RedCraft Redzimage DX1 (Installed — Combined Checkpoint)

Installs
25
GitHub Stars
729
First Seen
Apr 8, 2026
z-image-txt2img — artokun/comfyui-mcp