πŸ€— Model Requests

Qwen-Image-2.1 Turbo Quantized

GGUF, Safetensors, and MLX quantized formats of Qwen/Qwen-Image-2.1-Turbo for local image generation using the official turbo base weights.

Benchmark

Qwen-Image-2.1 benchmark

Turbo Quantized Files

Q4_K_M is recommended for the best balance of size and quality for GGUF. INT8 ConvRot and NVFP4 offer high performance on supported GPUs, while MLX formats are optimized for Apple Silicon.

Text Encoders & VAE

Companion model files packaged for ComfyUI:

Type File Precision Size
Text Encoder text_encoders/qwen3vl_8b_bf16.safetensors BF16 17.53 GB
Text Encoder text_encoders/qwen3vl_8b_int8_convrot.safetensors Int8 9.35 GB
VAE vae/qwen_image_2.1_vae_bf16.safetensors BF16 676 MB

Usage

Use the model with ComfyUI and ComfyUI-GGUF.

All required companion files (GGUF transformer, text encoder, and VAE) are hosted directly in this repository.

1. Download & File Placement

Download the files and place them in their respective ComfyUI directories:

ComfyUI/
└── models/
    β”œβ”€β”€ diffusion_models/
    β”‚   └── qwen-image-2.1-turbo-Q4_K_M.gguf      # Or fp8 / int8_convrot / NVFP4 .safetensors
    β”œβ”€β”€ text_encoders/
    β”‚   └── qwen3vl_8b_bf16.safetensors        # Or qwen3vl_8b_int8_convrot.safetensors (recommended for lower memory)
    └── vae/
        └── qwen_image_2.1_vae_bf16.safetensors

2. ComfyUI Setup

  1. Install ComfyUI-GGUF: Use the maintained fork with native Qwen-Image 2.1 support by cloning leejet/ComfyUI-GGUF into your custom nodes:
    cd ComfyUI/custom_nodes
    git clone https://github.com/leejet/ComfyUI-GGUF
    
  2. Node Configuration:
    • Diffusion Model: Add the Unet Loader (GGUF) node (for .gguf files) or standard UNETLoader (for .safetensors files) and select your downloaded model.
    • Text Encoder: Add the standard CLIPLoader node, select qwen3vl_8b_bf16.safetensors (or int8), and set type to qwen_image.
    • VAE: Add the standard VAELoader node and select qwen_image_2.1_vae_bf16.safetensors.
  3. Official Workflows:

Memory & Performance Notes

  • Optimal Setup (GPU + RAM): Keep the GGUF / Quantized diffusion model in GPU VRAM and let the text encoder run in / offload to System RAM (CPU).
  • Low VRAM Mode: If you experience VRAM out-of-memory errors, start ComfyUI with the --lowvram argument.

Source and build

Downloads last month
2,040
MLX
Hardware compatibility
Log In to add your hardware

Quantized

GGUF
Model size
7B params
Architecture
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for abenzerps/Qwen-Image-2.1-Turbo-Quantized

Quantized
(33)
this model