Wan3.0 Goes Live

Alibaba Cloud released Wan3.0 on August 24. This is the third generation of the Wanxiang video generation series. The previous Wan2.7 held the top spot on multiple video generation benchmarks. Wan3.0 brings meaningful upgrades in duration, input formats, visual quality, and editing capabilities.

Key Changes

Native 30-second video. Previous models typically generated 5-10 seconds per run. Wan3.0 outputs 30 seconds of continuous video in a single pass, with automatic scene splitting. Creators no longer need to stitch clips together manually.

Document format input. The most distinctive update: Wan3.0 is the first model to accept doc, xls, ppt, pdf, and md files as input. Upload a PPT, and the model parses the content to generate a matching video. This has real utility for corporate presentations and educational content.

Up to 20 reference assets. In image-to-video, character-driven, and scene reconstruction modes, you can feed in up to 20 reference images or video clips. The model extracts character appearance, scene style, and other attributes to compose into the new video.

Visual quality. Alibaba describes it as "cinematic realism." Actual results depend on prompts and source material. Based on examples from the GitHub repo, physical details like motion blur, light reflections, and fabric wrinkles show improvement over the previous generation.

Text rendering. Wan3.0 supports text rendering in 12 languages, including Chinese, English, Japanese, Korean, and Arabic. You can generate video segments with text overlays, such as title cards and subtitles. Previous video generation models handled this poorly.

Generation Modes

ModeDescription
Text to VideoText prompt to 30s video
Image to VideoStatic image to animated video
Reference to VideoMultiple reference assets to new scene video
Video EditingModify existing video based on prompts
Text to Image / Image EditHigh-fidelity image generation and conversational editing

Pricing

ResolutionAPI Price (CNY/sec)
480P¥0.3
720P¥0.6
1080P¥1.2

From August 24 to September 23, Alibaba Cloud Bailian and Qwen AI platforms offer a limited-time 30% discount on API calls. That brings a 10-second 1080P clip down to roughly ¥8.4 (vs. ¥12 at full price).

Third-party platforms including TiaoYue Vision, Deevid, JuHuo, JD Lingjing, Meitu, and JingMeng have already integrated Wan3.0.

Where to Try It

  • Alibaba Cloud Bailian platform
  • wan.video (official site)
  • WanJing YiKe
  • Qwen AI platform / Qwen App
  • DuiYou, IF STUDIO

Open-source code and model weights are published on GitHub: AlibabaCloud-Official/Wan3.0.

Competitive Landscape

The video generation space has several major players:

  • ByteDance Jimeng Seedream 4.5: Doubao's video model, API via Volcano Engine, strong at Chinese scenes and character consistency
  • Kuaishou Kling: One of the earliest commercial video generation models in China, realistic style
  • Runway Gen-4: Leading Western player, known for stylized and cinematic output
  • Pika 2.0: Lightweight video generation for rapid prototyping

Wan3.0's differentiators are document input and multi-reference support. For teams producing corporate videos or educational content at scale, these features address real pain points.

Final Notes

Competition in video generation models has intensified in the second half of 2026. Wan3.0 pushing single-pass generation to 30 seconds and adding document input represents a step toward practical utility. Actual results still depend heavily on input quality. If interested, try it out at wan.video.