
Turn natural language into polished images
Describe portraits, products, illustrations, marketing graphics, posters, or other visual assets in plain language and refine the result through prompt iteration.

Qwen Image 2.1 is Alibaba's compact image model for text-to-image generation and image editing. It supports native transparent assets, local edits, up to 10 reference images, 1K or 2K output, and stronger fidelity for people, products, and text.
1K: 2 credits / image × 1 image, rounded up for the estimate.






Image model overview
Qwen Image 2.1 is Alibaba's compact image model for text-to-image generation and image editing. It supports native transparent assets, local edits, up to 10 reference images, 1K or 2K output, and stronger fidelity for people, products, and text.
Choose text-to-image for a new visual, image-to-image for a reference-led result, transparent output for reusable assets, or local editing when only one area needs to change.

Product Features
Qwen Image 2.1 is Alibaba's compact image model for text-to-image generation and image editing. It supports native transparent assets, local edits, up to 10 reference images, 1K or 2K output, and stronger fidelity for people, products, and text.

Describe portraits, products, illustrations, marketing graphics, posters, or other visual assets in plain language and refine the result through prompt iteration.

Change objects, clothing, expressions, backgrounds, styles, and scene details without rebuilding the whole composition from scratch.

Generate transparent visual assets directly for stickers, product cutouts, design elements, graphics, and reusable creative libraries.
Image creation workflow
Choose a mode, organize prompts and references, configure the output, then review typography, layout, identity, and visual detail.




The unified generation and editing workflow makes Qwen Image 2.1 useful for both everyday visual creation and production-oriented image tools.



Qwen Image 2.1 uses a compact 7B architecture with 32 single-stream DiT layers, mixed-granularity attention, and KV-cache reuse to improve inference efficiency and reduce memory use.

A smaller model footprint makes the workflow practical for cost-conscious image applications and high-volume creative tools.
Try now
Keep creation, reference-led changes, transparent output, and local edits within one consistent image workflow.
Try now
Use the model for people, products, materials, typography, and other visuals where fidelity matters.
Try now
Start a new image project
Return to the image workspace, refine the prompt and references, choose the output settings, and start generating.