Introducing Qwen-Image-Edit: Novel Model in Image Editing

Main Content

Introducing Qwen-Image-Edit: Novel Model in Image Editing

  • The model can render intricate texts with high precision in generated images.
  • The model achieves SOTA performance across multiple benchmarks.


Alibaba has introduced Qwen-Image-Edit, the editing version of Qwen-Image – a novel image generation foundation model launched earlier this month. Built upon the 20-billion-parameter Qwen-Image model, the editing model excels in precise text editing, as well as sophisticated visual appearance and semantic editing.

Supporting both English and Chinese text editing, the model enables users to seamlessly add, delete, or modify text within images while preserving the original font size and style.

Beyond text, Qwen-Image-Edit supports detailed visual appearance editing—such as adding, removing, or altering visual elements—while keeping specific regions of the image intact. It also enables high-level semantic editing, including character generation, object rotation, and style transfer, all while maintaining semantic coherence and visual consistency.

Image Editing In Adding Objects
image editing in adding objects
Image Editing In Adding Objects 2
image editing in photo restoration

The model achieves state-of-the-art (SOTA) performance across multiple benchmarks, making itself a powerful foundation model for image editing. It is now open sourced on Hugging Face, Github and Alibaba’s open-source community ModelScope. Users can also experience the model on Qwen Chat under “Image Editing”.

The model’s remarkable editing capabilities is made possible by Qwen-Image’s powerful text rendering capabilities. With a deep understanding of complex linguistic structures, Qwen-Image is able to produce visually compelling and semantically accurate outputs, establishing itself as a leading model in the field.

Qwen Image Benchmarks
Figure1: Qwen-Image exhibits strong general capabilities in both image generation and editing, while demonstrating exceptional capability in text rendering, especially Chinese.

Through innovative approaches such as comprehensive data engineering, progressive learning strategies, enhanced multi-task training paradigms, and scalable infrastructure optimization, Qwen-image delivers exceptional precision in rendering intricate text within generated images. It excels in challenging scenarios involving multi-line layouts, paragraph-level semantics, and fine-grained visual details.

Book
[Prompt: Bookstore window display. A sign displays “New Arrivals This Week”. Below, a shelf tag with the text “Best-Selling Novels Here”. To the side, a colorful poster advertises “Author Meet And Greet on Saturday” with a central portrait of the author. There are four books on the bookshelf, namely “The light between worlds” “When stars are scattered” “The silent patient” “The night circus”]
Reuse this content

Sign Up For Our Newsletter

Stay updated on the digital economy with our free weekly newsletter