Alibaba has introduced Qwen-Image-Edit, the editing version of Qwen-Image – a novel image generation foundation model launched earlier this month. Built upon the 20-billion-parameter Qwen-Image model, the editing model excels in precise text editing, as well as sophisticated visual appearance and semantic editing.
Supporting both English and Chinese text editing, the model enables users to seamlessly add, delete, or modify text within images while preserving the original font size and style.
Beyond text, Qwen-Image-Edit supports detailed visual appearance editing—such as adding, removing, or altering visual elements—while keeping specific regions of the image intact. It also enables high-level semantic editing, including character generation, object rotation, and style transfer, all while maintaining semantic coherence and visual consistency.


The model achieves state-of-the-art (SOTA) performance across multiple benchmarks, making itself a powerful foundation model for image editing. It is now open sourced on Hugging Face, Github and Alibaba’s open-source community ModelScope. Users can also experience the model on Qwen Chat under “Image Editing”.
The model’s remarkable editing capabilities is made possible by Qwen-Image’s powerful text rendering capabilities. With a deep understanding of complex linguistic structures, Qwen-Image is able to produce visually compelling and semantically accurate outputs, establishing itself as a leading model in the field.

Through innovative approaches such as comprehensive data engineering, progressive learning strategies, enhanced multi-task training paradigms, and scalable infrastructure optimization, Qwen-image delivers exceptional precision in rendering intricate text within generated images. It excels in challenging scenarios involving multi-line layouts, paragraph-level semantics, and fine-grained visual details.
