The world of digital content is constantly evolving, and video remains king. From marketing and entertainment to education and personal storytelling, high-quality video content is more in demand than ever. However, creating and editing compelling videos often requires specialized skills, expensive software, and significant time. What if there was a powerful, accessible, and versatile tool that could change that? Enter Wan2.1 VACE, an all-in-one AI model poised to revolutionize video creation and editing.
This comprehensive guide will delve into the exciting capabilities of Wan2.1 VACE, exploring its features, how to get started, and why it’s a game-changer for content creators and developers alike.
Key Takeaways From This Article:
- Wan2.1 VACE is a state-of-the-art, all-in-one open-source AI model designed to revolutionize video creation and editing.
- It offers broad accessibility by supporting consumer-grade GPUs, making advanced AI video tools available to more users.
- Wan2.1 VACE excels in diverse tasks including Text-to-Video, Image-to-Video, AI video editing, and unique visual text generation.
- The model is built on innovative technologies like Wan-VAE and Video Diffusion DiT, with strong community backing and a clear development roadmap.
Table of contents
- What is Wan2.1? An Open and Advanced Suite for Video Generation
- Unpacking the Power of Wan2.1 VACE: Key Capabilities
- Getting Started with Wan2.1 VACE: Installation and Model Downloads
- Exploring Wan2.1 VACE in Action: Generation Tasks
- The Engine Behind Wan2.1: Technical Innovations
- Wan2.1 VACE Performance: Benchmarks and Efficiency
- The Future is Open: Community and Development
- Why Wan2.1 VACE is a Game-Changer for Content Creators
- Conclusion: Embrace the Future of Video with Wan2.1 VACE
What is Wan2.1? An Open and Advanced Suite for Video Generation
Wan2.1 VACE is part of the larger Wan project, which aims to deliver open and advanced large-scale video generative models. This initiative has produced a comprehensive suite of video foundation models that are pushing the boundaries of what’s possible in AI-driven video generation. Wan2.1 stands out for its commitment to open-source principles, empowering a wider community to leverage and contribute to its development.
At its core, Wan2.1 is designed to be a versatile and powerful tool. It’s not just about one specific task; it’s an entire ecosystem for generating and manipulating video content through artificial intelligence.
Unpacking the Power of Wan2.1 VACE: Key Capabilities
Wan2.1 VACE brings a host of impressive features to the table, setting a new standard in the realm of AI video tools.

State-of-the-Art (SOTA) Performance
One of the most significant claims of Wan2.1 is its SOTA performance. Internal benchmarks and comparisons suggest that Wan2.1 consistently outperforms existing open-source models and even rivals some state-of-the-art commercial solutions. This level of performance is a testament to the advanced architecture and training methodologies behind the model.
Accessibility: Consumer-Grade GPU Support
Perhaps one of the most exciting aspects of Wan2.1 is its accessibility. The T2V-1.3B model, for instance, requires only 8.19 GB of VRAM. This makes it compatible with a wide range of consumer-grade GPUs, meaning you don’t necessarily need a high-end, expensive hardware setup to start creating. For example, it can generate a 5-second, 480P video on an RTX 4090 in about 4 minutes, even without advanced optimization techniques like quantization. This opens up powerful video generation capabilities to a much broader audience.
Versatility: A Multitude of Supported Tasks
Wan2.1 VACE is not a one-trick pony. It excels in a variety of video and image-related tasks, making it an incredibly flexible tool:
- Text-to-Video (T2V): Generate videos from simple text prompts.
- Image-to-Video (I2V): Bring still images to life by transforming them into dynamic video sequences.
- Video Editing: Modify existing videos with AI-powered tools.
- Text-to-Image (T2I): Create still images from text descriptions.
- Video-to-Audio: Generate accompanying audio for video content.
This multi-task proficiency means users can handle various stages of the creative workflow within the Wan2.1 ecosystem.
Groundbreaking Visual Text Generation
A unique and highly practical feature of Wan2.1 is its ability to generate visual text within videos. It is reportedly the first video model capable of generating both Chinese and English text robustly. This capability significantly enhances its practical applications, particularly for creating informational content, social media videos, or any project requiring embedded text.
The Backbone: Powerful Wan-VAE
Underpinning many of Wan2.1’s capabilities is the Wan-VAE (Variational Autoencoder). This component delivers exceptional efficiency and performance in encoding and decoding video data. It can handle 1080P videos of any length while crucially preserving temporal information. This makes Wan-VAE an ideal foundation not just for video generation but also for image generation tasks.

Getting Started with Wan2.1 VACE: Installation and Model Downloads
Ready to explore the potential of Wan2.1 VACE? Here’s how you can get started.
Quick Installation Guide
The first step is to set up the Wan2.1 environment on your system.
- Clone the Repository:
Open your terminal and clone the official GitHub repository:git clone https://github.com/Wan-Video/Wan2.1.gitcd Wan2.1
- Install Dependencies:
Ensure you have PyTorch version 2.4.0 or higher. Then, install the necessary dependencies listed in the requirements.txt file:pip install -r requirements.txt
Downloading Wan2.1 Models
Wan2.1 offers several models tailored for different tasks and resolutions. You can download these models from Hugging Face or ModelScope.
Here’s a rundown of the available models:
- T2V-14B: Supports 480P and 720P Text-to-Video.
- I2V-14B-720P: Supports 720P Image-to-Video.
- I2V-14B-480P: Supports 480P Image-to-Video.
- T2V-1.3B: Supports 480P Text-to-Video (and can generate 720P, though less stable).
- FLF2V-14B: Supports 720P First-Last-Frame-to-Video.
- VACE-1.3B: Supports 480P for various VACE tasks.
- VACE-14B: Supports both 480P and 720P for VACE tasks.
Important Notes:
- While the 1.3B model can technically generate videos at 720P, it has had limited training at this resolution. For optimal and more stable results with the 1.3B model, 480P resolution is recommended.
- For First-Last-Frame-to-Video generation, the model was primarily trained on Chinese text-video pairs. Therefore, using Chinese prompts is recommended for achieving better results in this specific task.
You can download models using the huggingface-cli:
pip install "huggingface_hub[cli]"
huggingface-cli download Wan-AI/Wan2.1-T2V-14B --local-dir ./Wan2.1-T2V-14B
Alternatively, use modelscope-cli:
pip install modelscope
modelscope download Wan-AI/Wan2.1-T2V-14B --local_dir ./Wan2.1-T2V-14B
Replace Wan-AI/Wan2.1-T2V-14B and ./Wan2.1-T2V-14B with the specific model and local directory you require.
Exploring Wan2.1 VACE in Action: Generation Tasks
With the setup complete, let’s explore how to use Wan2.1 VACE for various video and image generation tasks.
Text-to-Video (T2V) Generation with Wan2.1 VACE
Wan2.1 supports two primary Text-to-Video models (1.3B and 14B) at 480P and 720P resolutions.
Running T2V without Prompt Extension:
For a basic implementation, you can generate videos directly from your prompt.
- Single-GPU inference example:
python generate.py --task t2v-14B --size 1280*720 --ckpt_dir ./Wan2.1-T2V-14B --prompt "Two anthropomorphic cats in comfy boxing gear and bright gloves fight intensely on a spotlighted stage."If you encounter Out-of-Memory (OOM) issues, especially on GPUs like the RTX 4090, you can use options like –offload_model True and –t5_cpu to reduce VRAM usage:python generate.py --task t2v-1.3B --size 832*480 --ckpt_dir ./Wan2.1-T2V-1.3B --offload_model True --t5_cpu --sample_shift 8 --sample_guide_scale 6 --prompt "Two anthropomorphic cats in comfy boxing gear and bright gloves fight intensely on a spotlighted stage."For the T2V-1.3B model, a –sample_guide_scale of 6 is recommended, and –sample_shift can be adjusted (8-12) based on performance. - Multi-GPU inference using FSDP + xDiT USP:
For faster inference on multi-GPU setups, Wan2.1 utilizes FSDP (Fully Sharded Data Parallel) and xDiT USP. You’ll need to install xfuser:pip install "xfuser>=0.4.1" torchrun --nproc_per_node=8 generate.py --task t2v-14B --size 1280*720 --ckpt_dir ./Wan2.1-T2V-14B --dit_fsdp --t5_fsdp --ulysses_size 8 --prompt "Two anthropomorphic cats in comfy boxing gear and bright gloves fight intensely on a spotlighted stage."Strategies like Ulysses and Ring can be employed depending on your GPU configuration and model.
Running T2V with Prompt Extension:
Extending prompts can significantly enrich the details and quality of generated videos. Wan2.1 offers two methods:
- Dashscope API: Requires an API key and uses models like qwen-plus.
DASH_API_KEY=your_key python generate.py --task t2v-14B --size 1280*720 --ckpt_dir ./Wan2.1-T2V-14B --prompt "Your prompt" --use_prompt_extend --prompt_extend_method 'dashscope' --prompt_extend_target_lang 'zh' - Local Model: Uses Hugging Face Qwen models by default (e.g., Qwen/Qwen2.5-14B-Instruct). Larger models yield better results but need more GPU memory.
python generate.py --task t2v-14B --size 1280*720 --ckpt_dir ./Wan2.1-T2V-14B --prompt "Your prompt" --use_prompt_extend --prompt_extend_method 'local_qwen' --prompt_extend_target_lang 'zh
Running T2V with Diffusers:
Wan2.1 integrates with Diffusers for a streamlined experience (note: prompt extension and distributed inference for Diffusers are upcoming features).
import torch
from diffusers.utils import export_to_video
from diffusers import AutoencoderKLWan, WanPipeline
from diffusers.schedulers.scheduling_unipc_multistep import UniPCMultistepScheduler
model_id = "Wan-AI/Wan2.1-T2V-14B-Diffusers" # Or Wan-AI/Wan2.1-T2V-1.3B-Diffusers
vae = AutoencoderKLWan.from_pretrained(model_id, subfolder="vae", torch_dtype=torch.float32)
flow_shift = 5.0 # 5.0 for 720P, 3.0 for 480P
scheduler = UniPCMultistepScheduler(prediction_type='flow_prediction', use_flow_sigmas=True, num_train_timesteps=1000, flow_shift=flow_shift)
pipe = WanPipeline.from_pretrained(model_id, vae=vae, torch_dtype=torch.bfloat16)
pipe.scheduler = scheduler
pipe.to("cuda")
prompt = "A cat and a dog baking a cake together in a kitchen."
negative_prompt = "Bright tones, overexposed, static, blurred details..."
output = pipe(prompt=prompt, negative_prompt=negative_prompt, height=720, width=1280, num_frames=81, guidance_scale=5.0).frames[0]
export_to_video(output, "output.mp4", fps=16)
Running local Gradio demo for T2V:
Navigate to the gradio directory and run the appropriate Python script, specifying your checkpoint directory and prompt extension method if used.
Image-to-Video (I2V) Generation with Wan2.1 VACE
Transform static images into videos. Wan2.1 offers I2V models for 480P and 720P.
Running I2V without Prompt Extension:
- Single-GPU inference:
python generate.py --task i2v-14B --size 1280*720 --ckpt_dir ./Wan2.1-I2V-14B-720P --image examples/i2v_input.JPG --prompt "Summer beach vacation style, a white cat wearing sunglasses..."The –size parameter determines the generated video’s area, maintaining the input image’s aspect ratio. - Multi-GPU inference:
Similar to T2V, use torchrun with FSDP options.
Running I2V with Prompt Extension:
The process is similar to T2V prompt extension, using either Dashscope or a local model (e.g., Qwen/Qwen2.5-VL-7B-Instruct for I2V).
Running I2V with Diffusers:
A Diffusers pipeline is also available for I2V.
import torch
import numpy as np
from diffusers import AutoencoderKLWan, WanImageToVideoPipeline
from diffusers.utils import export_to_video, load_image
from transformers import CLIPVisionModel
model_id = "Wan-AI/Wan2.1-I2V-14B-720P-Diffusers" # Or Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
image_encoder = CLIPVisionModel.from_pretrained(model_id, subfolder="image_encoder", torch_dtype=torch.float32)
vae = AutoencoderKLWan.from_pretrained(model_id, subfolder="vae", torch_dtype=torch.float32)
pipe = WanImageToVideoPipeline.from_pretrained(model_id, vae=vae, image_encoder=image_encoder, torch_dtype=torch.bfloat16)
pipe.to("cuda")
image = load_image("your_image_path.jpg")
# Resize image appropriately
prompt = "An astronaut hatching from an egg..."
negative_prompt = "Bright tones, overexposed, static..."
output = pipe(image=image, prompt=prompt, negative_prompt=negative_prompt, height=height, width=width, num_frames=81, guidance_scale=5.0).frames[0]
export_to_video(output, "output.mp4", fps=16)
Running local Gradio demo for I2V:
Check the gradio directory for scripts like i2v_14B_singleGPU.py, allowing you to specify 480P, 720P, or both model checkpoint directories.
First-Last-Frame-to-Video (FLF2V) Generation
Generate a video sequence between a given first and last frame. Currently, this supports 720P.
Running FLF2V without Prompt Extension:
- Single-GPU inference:
python generate.py --task flf2v-14B --size 1280*720 --ckpt_dir ./Wan2.1-FLF2V-14B-720P --first_frame examples/flf2v_input_first_frame.png --last_frame examples/flf2v_input_last_frame.png --prompt "CG animation style, a small blue bird takes off..."As with I2V, –size refers to the video area, matching the input frames’ aspect ratio. Remember, Chinese prompts are recommended for optimal results. - Multi-GPU inference:
Use torchrun as with other tasks.
Running FLF2V with Prompt Extension:
Follow the same prompt extension methods (Dashscope or local model) as for T2V and I2V.
Running local Gradio demo for FLF2V:
The gradio directory contains scripts like flf2v_14B_singleGPU.py.
Leveraging VACE for Advanced Video Editing and Generation
The Wan2.1 VACE (Video Advanced Composition and Editing) component is where the all-in-one capabilities truly shine. It supports 1.3B and 14B models for 480P and 720P resolutions, respectively. VACE allows users to input text prompts along with optional videos, masks, and reference images for sophisticated video generation or editing.
Preprocessing for VACE:
For tasks like Reference-to-Video (R2V), preprocessing might be minimal. However, for Video-to-Video (V2V) editing and Masked Video-to-Video (MV2V) editing, additional preprocessing is needed to obtain video with conditions like depth maps, pose information, or masked regions. Refer to the vace_preproccess documentation for details.
CLI Inference with VACE:
- Single-GPU inference:
python generate.py --task vace-1.3B --size 832*480 --ckpt_dir ./Wan2.1-VACE-1.3B --src_ref_images examples/girl.png,examples/snake.png --prompt "A festive scene with a girl and her cartoon snake..." - Multi-GPU inference:
torchrun --nproc_per_node=8 generate.py --task vace-14B --size 1280*720 --ckpt_dir ./Wan2.1-VACE-14B --dit_fsdp --t5_fsdp --ulysses_size 8 --src_ref_images examples/girl.png,examples/snake.png --prompt "A festive scene..."
Running local Gradio demo for VACE:
Scripts like gradio/vace.py are available, supporting single-GPU and multi-GPU (FSDP + xDiT USP) inference.
Text-to-Image (T2I) Generation with Wan2.1 VACE
Because Wan2.1 is trained on both image and video data, it’s also proficient at generating still images. The command structure is similar to video generation.
Running T2I without Prompt Extension:
- Single-GPU inference:
python generate.py --task t2i-14B --size 1024*1024 --ckpt_dir ./Wan2.1-T2V-14B --prompt 'A simple and elegant beauty' - Multi-GPU inference:
torchrun --nproc_per_node=8 generate.py --dit_fsdp --t5_fsdp --ulysses_size 8 --base_seed 0 --frame_num 1 --task t2i-14B --size 1024*1024 --prompt 'A simple and elegant beauty' --ckpt_dir ./Wan2.1-T2V-14B
Running T2I with Prompt Extension:
Enable –use_prompt_extend with your chosen method (Dashscope or local Qwen model) for richer image details.
The Engine Behind Wan2.1: Technical Innovations
The impressive capabilities of Wan2.1 VACE are built upon a foundation of significant technical advancements.
Advanced 3D Variational Autoencoders (Wan-VAE)
A core innovation is the novel 3D causal VAE architecture, termed Wan-VAE, specifically designed for video generation. By combining multiple strategies, Wan-VAE improves spatio-temporal compression, reduces memory usage, and critically ensures temporal causality. This results in significant advantages in performance and efficiency compared to other open-source VAEs. Wan-VAE’s ability to encode and decode unlimited-length 1080P videos without losing historical temporal information makes it exceptionally well-suited for video generation tasks.
Cutting-Edge Video Diffusion DiT
Wan2.1 employs the Flow Matching framework within the mainstream Diffusion Transformers (DiT) paradigm. Its architecture uses a T5 Encoder to process multilingual text input. Cross-attention mechanisms in each transformer block embed this text information into the model structure. Furthermore, an MLP (Multi-Layer Perceptron) with a Linear layer and a SiLU activation layer processes input time embeddings to predict six modulation parameters individually. This MLP is shared across all transformer blocks, with each block learning a distinct set of biases, a design choice that experimental findings show leads to significant performance improvements at the same parameter scale.
The Foundation: High-Quality Data Curation

The quality and scale of training data are paramount for any AI model. Wan2.1 benefits from a meticulously curated and deduplicated candidate dataset comprising vast amounts of image and video data. The data curation process involves a rigorous four-step data cleaning pipeline, focusing on fundamental dimensions, visual quality, and motion quality. This robust data processing pipeline ensures the availability of high-quality, diverse, and large-scale training sets.
Wan2.1 VACE Performance: Benchmarks and Efficiency
Wan2.1 doesn’t just promise features; it aims to deliver top-tier performance.

Manual Evaluation: Outperforming the Competition
Through extensive manual evaluations using a carefully designed set of 1,035 internal prompts (covering 14 major dimensions and 26 sub-dimensions), Wan2.1 has been compared against leading open-source and closed-source models. The results reportedly demonstrate Wan2.1’s superior performance, especially when prompt extension is utilized. These evaluations cover both Text-to-Video and Image-to-Video tasks, indicating that Wan2.1 often surpasses other models in quality and adherence to prompts.
Computational Efficiency Across GPUs
The developers have tested the computational efficiency of different Wan2.1 models on various GPUs. The results, typically presented as “Total time (s) / peak GPU memory (GB),” show how the models perform under different hardware constraints. For example, specific settings are provided for running the 1.3B model on 8 GPUs (using Ring Strategy) or the 14B model on a single GPU (with model offloading). These tests are generally conducted without prompt extension to provide baseline performance figures. It’s noted, for instance, that T2V-14B is slower than I2V-14B because the former samples 50 steps while the latter uses 40 steps.
The Future is Open: Community and Development
Wan2.1 is a living project with an active community and a clear roadmap for future enhancements.
Community Contributions:
The open nature of Wan2.1 has already fostered exciting community works. Several projects have built upon or integrated Wan2.1:
- Phantom: Developed a unified video generation framework for single and multi-subject references based on Wan2.1-T2V-1.3B.
- UniAnimate-DiT: Trained a human image animation model based on Wan2.1-14B-I2V, with open-sourced code.
- CFG-Zero: Enhanced Wan2.1 (T2V and I2V models) from the perspective of Classifier-Free Guidance (CFG).
- TeaCache: Now supports Wan2.1 acceleration, potentially doubling speeds.
- DiffSynth-Studio: Provides extended support for Wan2.1, including video-to-video, FP8 quantization, VRAM optimization, and LoRA training.
Roadmap: The Todo List for Wan2.1:
The developers have an extensive to-do list, signaling ongoing improvements and feature additions across all components of Wan2.1, including Text-to-Video, Image-to-Video, First-Last-Frame-to-Video, and VACE. This includes plans for multi-GPU inference code, more checkpoints, Gradio demos, and deeper integration with Diffusers, including multi-GPU support.
Why Wan2.1 VACE is a Game-Changer for Content Creators
Wan2.1 VACE stands out for several key reasons:
- High-Quality Output: It aims for state-of-the-art results, rivaling even commercial offerings.
- Accessibility: Support for consumer-grade GPUs democratizes access to advanced AI video tools.
- Versatility: Its ability to handle text-to-video, image-to-video, video editing, and even visual text generation makes it incredibly flexible.
- Open Source: This fosters community involvement, innovation, and transparency.
- Continuous Development: A clear roadmap and active community ensure the platform will continue to evolve.
For independent creators, small businesses, researchers, and AI enthusiasts, Wan2.1 VACE offers an unprecedented opportunity to explore and create sophisticated video content without prohibitive costs or hardware requirements.
Conclusion: Embrace the Future of Video with Wan2.1 VACE
Wan2.1 VACE represents a significant leap forward in open-source AI video generation and editing. Its combination of SOTA performance, broad feature set, accessibility, and strong community support makes it an incredibly exciting tool. Whether you’re looking to generate captivating videos from text prompts, animate still images, edit existing footage with AI precision, or incorporate dynamic text into your visuals, Wan2.1 VACE provides a powerful and evolving platform to bring your creative visions to life.
To dive deeper and join the community, explore the Wan2.1 VACE resources on GitHub, Hugging Face, and ModelScope. The future of AI-powered video creation is here, and it’s more open and accessible than ever thanks to innovations like Wan2.1 VACE.
| Latest From Us
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure


