Wan3.0 Goes Live
Alibaba Cloud released Wan3.0 on August 24. This is the third generation of the Wanxiang video generation series. The previous Wan2.7 held the top spot on multiple video generation benchmarks. Wan3.0 brings meaningful upgrades in duration, input formats, visual quality, and editing capabilities.
Key Changes
Native 30-second video. Previous models typically generated 5-10 seconds per run. Wan3.0 outputs 30 seconds of continuous video in a single pass, with automatic scene splitting. Creators no longer need to stitch clips together manually.
Document format input. The most distinctive update: Wan3.0 is the first model to accept doc, xls, ppt, pdf, and md files as input. Upload a PPT, and the model parses the content to generate a matching video. This has real utility for corporate presentations and educational content.
Up to 20 reference assets. In image-to-video, character-driven, and scene reconstruction modes, you can feed in up to 20 reference images or video clips. The model extracts character appearance, scene style, and other attributes to compose into the new video.
Visual quality. Alibaba describes it as "cinematic realism." Actual results depend on prompts and source material. Based on examples from the GitHub repo, physical details like motion blur, light reflections, and fabric wrinkles show improvement over the previous generation.
Text rendering. Wan3.0 supports text rendering in 12 languages, including Chinese, English, Japanese, Korean, and Arabic. You can generate video segments with text overlays, such as title cards and subtitles. Previous video generation models handled this poorly.
Generation Modes
| Mode | Description |
|---|---|
| Text to Video | Text prompt to 30s video |
| Image to Video | Static image to animated video |
| Reference to Video | Multiple reference assets to new scene video |
| Video Editing | Modify existing video based on prompts |
| Text to Image / Image Edit | High-fidelity image generation and conversational editing |
Pricing
| Resolution | API Price (CNY/sec) |
|---|---|
| 480P | ¥0.3 |
| 720P | ¥0.6 |
| 1080P | ¥1.2 |
From August 24 to September 23, Alibaba Cloud Bailian and Qwen AI platforms offer a limited-time 30% discount on API calls. That brings a 10-second 1080P clip down to roughly ¥8.4 (vs. ¥12 at full price).
Third-party platforms including TiaoYue Vision, Deevid, JuHuo, JD Lingjing, Meitu, and JingMeng have already integrated Wan3.0.
Where to Try It
- Alibaba Cloud Bailian platform
- wan.video (official site)
- WanJing YiKe
- Qwen AI platform / Qwen App
- DuiYou, IF STUDIO
Open-source code and model weights are published on GitHub: AlibabaCloud-Official/Wan3.0.
Competitive Landscape
The video generation space has several major players:
- ByteDance Jimeng Seedream 4.5: Doubao's video model, API via Volcano Engine, strong at Chinese scenes and character consistency
- Kuaishou Kling: One of the earliest commercial video generation models in China, realistic style
- Runway Gen-4: Leading Western player, known for stylized and cinematic output
- Pika 2.0: Lightweight video generation for rapid prototyping
Wan3.0's differentiators are document input and multi-reference support. For teams producing corporate videos or educational content at scale, these features address real pain points.
Final Notes
Competition in video generation models has intensified in the second half of 2026. Wan3.0 pushing single-pass generation to 30 seconds and adding document input represents a step toward practical utility. Actual results still depend heavily on input quality. If interested, try it out at wan.video.



