Alibaba has pushed its Wan3.0 video generation model into public beta, delivering clips up to 30 seconds long with a radical new capability: ingest entire documents and turn them into moving video.
The model accepts text prompts, images, video clips, audio files, PDFs, PowerPoint presentations, and even live web pages as input sources. A single prompt can blend up to ten images, five video clips, and five audio tracks simultaneously — a dramatic shift from the text-to-video pipelines that dominated 2024 and early 2025.
This is not incremental. The 30-second runtime doubles what Wan2.5 could produce, and the document ingestion feature transforms static business assets — pitch decks, annual reports, technical manuals — into dynamic visual content without a single frame of manual animation.
How It Works
Wan3.0 processes text, images, video, and audio in parallel, fusing them into coherent outputs. Alibaba claims the model reduces "visual drift" — the tendency for AI-generated videos to distort faces, warp objects, or lose spatial consistency over time — by anchoring generation to reference material.
Upload a character design, a product shot, or a brand guideline PDF, and Wan3.0 keeps those elements visually stable across the entire clip. For enterprise use cases — marketing videos, training modules, product demos — this consistency is the difference between usable output and unusable slop.
The model also includes an extension tool that can lengthen existing videos, though Alibaba does not yet disclose whether this uses interpolation, generative fill, or a hybrid approach.
Pricing and Availability
Wan3.0 is accessible through three channels:
- wan.video — direct web interface
- Alibaba Cloud Model Studio — enterprise API access
- Qwen Cloud — developer API tier
Two tiers exist: Standard (currently discounted 30%) and Prime (faster inference). Pricing per second of output:
| Resolution | Standard | Prime |
|---|---|---|
| 480p | $0.05 | $0.068 |
| 720p | $0.10 | $0.14 |
| 1080p | $0.20 | $0.28 |
A full 30-second 1080p clip runs $6.00 on Standard or $8.40 on Prime. For comparison, Runway Gen-4 and Luma Dream Machine sit in similar price ranges, but neither accepts documents as input.
Where Alibaba Wants This to Go
Alibaba is positioning Wan3.0 across three verticals:
Film and entertainment — short dramas, social media clips, and pre-visualization for production teams. The 30-second window covers most TikTok, Reels, and YouTube Shorts formats natively.
Enterprise marketing — transform product brochures and brand guidelines into video ads without hiring animators. A retail brand could upload seasonal lookbooks and generate localized video content in hours instead of weeks.
Robotics and autonomous systems — this is the sleeper use case. Realistic simulation footage trains vision models for robots and self-driving cars. If Wan3.0 can generate consistent, photorealistic environments at scale, it becomes infrastructure for the embodied AI race China is already winning.
The Bigger Picture: Alibaba's AI Bet
Wan3.0 lands inside a broader financial story. Last week, Alibaba announced the largest share sale by a Hong Kong-listed company ever — HK$80 billion ($10.3 billion) — to fund its AI infrastructure buildout. Quarterly profits dropped 75% year-over-year, explicitly attributed to "sharply higher AI investments."
This is not a company testing the water. This is a capital-intensive commitment to win the Chinese AI ecosystem, from foundation models (Qwen) to video generation (Wan) to cloud infrastructure (Alibaba Cloud).
The timing is strategic. Western models like OpenAI's Sora and Runway's Gen-4 face export restrictions and data sovereignty concerns in China. Wan3.0 offers a domestic alternative with document ingestion — a feature Western models don't match — and pricing competitive enough for enterprise adoption.
🔥 Hot Takes
1. Document-to-video is the enterprise wedge Western models can't touch. Runway and Pika build pretty videos. Wan3.0 turns your Q3 earnings PDF into a boardroom presentation clip. That's not a creative tool — that's a productivity multiplier with a pricing model that undercuts Western alternatives by 40%.
2. The 75% profit drop is a feature, not a bug. Alibaba is sacrificing short-term earnings to build the AI stack China can't be sanctioned out of. When the US restricts H100 exports again, Alibaba Cloud will already have the compute, the models, and the talent pipeline. This is AI nationalism priced into the P&L.
3. 30 seconds changes the format math. Most AI video today tops out at 10-15 seconds. At 30 seconds, you can tell complete narrative arcs — setup, conflict, resolution — within a single clip. That pushes AI video from "cool demo" to "usable content format" for platforms that reward longer watch time.
Bottom Line
Wan3.0 is Alibaba's most capable video model yet, and it arrives at a moment when Chinese AI companies are turning capital into capability faster than Western competitors can respond. The document ingestion feature is a genuine differentiator. The pricing is aggressive. The timing — coinciding with an $80 billion funding round — signals this is just the opening move.
While Western models debate safety and alignment, Alibaba is shipping products that Chinese enterprises can actually use.