WAN 3.0 is for creators working from real source materials, bringing references, duration and pacing, consistency, and editing into one video workflow.
WAN 3.0 accepts multiple input types, letting visuals, sound, documents, and web page context shape the video output without first converting everything into a single format.
With native 30-second generation, WAN 3.0 gives scenes, pacing, and information structure more space to develop. Smart duration support helps the runtime match the creative direction.
Pixel-level consistency helps preserve important reference details, while video edits can follow your instructions or use reference assets to guide changes.
Use images, clips, audio, documents, or web pages as references to shape a 30-second video around existing source materials.
Use the longer native duration to plan a beginning, development, and ending with clearer pacing.
Edit with instructions or reference assets to refine video and move through multiple iterations more efficiently.
It can use images, videos, audio, documents, and web pages, with up to 20 reference assets supported in one generation.
WAN 3.0 supports native 30-second video generation and also includes smart duration support.
The model supports pixel-level consistency, helping the generated video stay closer to details in the reference assets.
Yes. Video editing can follow your instructions or be guided by reference assets.
Contains AI-generated content