Alibaba has launched Wan2.1-VACE (Video All-in-one Creation and Editing), an advanced open-source model poised to redefine the video production landscape through powerful AI-driven capabilities.
This innovative tool represents a significant advancement in video technology, integrating multiple sophisticated functions into a single, easy-to-use model that streamlines the video production process.
As part of Alibaba’s acclaimed Wan2.1 video generation AI model series, Wan2.1-VACE stands out as the industry’s first open-source model, offering a comprehensive solution for both video generation and editing tasks.
The company has released two versions of the model: a robust 14-billion-parameter edition and a compact 1.3-billion-parameter alternative, both of which are accessible through Hugging Face, GitHub, and Alibaba Cloud’s ModelScope community.
This initiative underscores Alibaba’s ongoing commitment to democratizing advanced AI technologies, empowering developers and businesses worldwide with state-of-the-art video production tools.

All-in-one AI Model Supports Versatile Applications
Wan2.1-VACE introduces powerful features, equipping content creators with unprecedented flexibility. Users can effortlessly generate videos using diverse inputs such as text, images, and existing video footage, enhancing creative possibilities.
Leveraging advanced editing technologies, the model empowers users to create dynamic videos from static images, animate specific interacting subjects based on provided samples, and employ sophisticated video repainting capabilities, including pose transfer, motion control, depth manipulation, and recolorization.
One notable highlight is Wan2.1-VACE’s ability to modify or remove specific video segments seamlessly without affecting surrounding content. Additionally, it intelligently extends video boundaries, automatically filling in gaps to create richer visual narratives.
The model enables users to effortlessly combine multiple editing functions, such as converting static images into lively videos, precisely controlling object movements, swapping characters or objects, animating referenced subjects, and transforming vertical images into horizontal video formats.

Cutting-Edge Technology
Central to Wan2.1-VACE is the innovative Video Condition Unit (VCU), a unified interface that is adept at processing various inputs, including text, images, videos, and masks. Another distinguishing feature, the Context Adapter structure, incorporates both temporal and spatial dimensions, enhancing the model’s flexibility to meet diverse editing needs.
The model’s versatile architecture positions it for adoption across multiple sectors, including social media content production and advertising, film and television post-production, special effects, and educational content creation.
Since August 2023, Alibaba has actively contributed to the open-source AI community, releasing multiple advanced AI models, including four Wan2.1 variants and an AI capable of generating videos from start and end frames. These open-source efforts have resonated widely, as evidenced by the Wan2.1 series surpassing 3.3 million downloads on platforms like Hugging Face and ModelScope.
First unveiled earlier this year, the Wan2.1 series notably pioneered video generation with text effects in both Chinese and English, reflecting Alibaba’s leadership and innovation in AI-driven multimedia technologies.
Watch the video below to discover more features about Wan2.1-VACE.