AI News Feed
Market watch
Computer Vision

Alibaba's Qwen Releases Qwen-Image-2.1-Turbo, Cutting 2K Image Generation to 8 Steps

Alibaba's Qwen team has released Qwen-Image-2.1-Turbo, a 7B open-weight checkpoint that generates and edits 2K images in eight denoising steps instead of 40, backed by a hosted API priced at CNY 0.1 per image.

The checkpoint loads directly with QwenImage21Pipeline in Diffusers and ships with its recommended eight-step sampling schedule saved inside. Generation uses CFG=1 by default, and prefix KV caching reuses the text and reference-image context across denoising steps, which is what allows the step count to fall without recomputing conditioning at every stage.

The underlying architecture is a single-stream diffusion transformer with 32 layers and 7B parameters, paired with a Qwen3-VL 8B text encoder that processes both instructions and condition images. It uses block-causal attention, applying a token-level causal mask to text tokens and a chunk-level bidirectional mask to images. The VAE is a 64-channel RGBA autoencoder with 16x spatial compression, which enables native transparency, and the scheduler is Flow Matching with Euler discrete scheduling and dynamic shifting.

Qwen's model card showcases eight categories, including portraits, human poses, transparent images, typography and posters, and UI layouts. Editing examples cover single-image transformation, multi-reference composition and a four-image interior composition. The model retains the same 2K output and editing feature set as Qwen-Image-2.1; the base model supports up to ten reference images and local edits made through circles, painted annotations or masks. Supported resolution presets run from 2048x2048 square to 2752x1536 at 16:9.

The Turbo checkpoint runs on CUDA GPUs in BF16 via Diffusers. Qwen publishes no Turbo-specific VRAM minimum. Unsloth estimates the base model runs on 11 GB of VRAM with GGUF and 24 GB with INT8 or FP8 precision.

No Turbo-specific benchmark has been published. Qwen reports a score of 60.28 on Qwen-Image-Bench for the base Qwen-Image-2.1, which it describes as the top open-weight score it has measured.

Running the model requires Diffusers installed from source along with transformers 5.17.0 or later. The checkpoint depends on Diffusers pull request 14950, which adds pipeline-configured sampling sigmas. Qwen notes that setting num_inference_steps alone does not override the saved schedule; only an explicit sigmas argument does, and other schedules are untested.

Alibaba Cloud Model Studio now hosts both Turbo and Pro. The qwen-image-2.1-turbo endpoint costs CNY 0.1 per image in most regions with a 120 requests-per-minute limit, while qwen-image-2.1-pro costs CNY 0.25 per image with a 20 requests-per-minute limit, making Turbo 2.5 times cheaper per image with six times the request rate.

Among comparable fast image models, Z-Image-Turbo from Alibaba's Tongyi-MAI group uses a 6B generator with eight NFEs and lists Apache 2.0 licensing, but it does not perform editing in the same checkpoint, relying on a separate Edit model. FLUX.2 klein 9B from Black Forest Labs uses a 9B generator, four steps and multi-reference editing, and is shown at 1024x1024 in examples under a non-commercial license.

Qwen-Image-2.1-Turbo is released under the Qwen Research License, the same terms as the base model, so commercial self-hosting requires separate permission. Weights are available through ModelScope and Hugging Face.