Wan 3.0 document to video: what the model reads and what Sume passes
Alibaba's Wan 3.0 can read an uploaded document to make a video. On Sume's video API today wan-3.0 takes prompts, frames and media references, not a file input.

Wan 3.0 can take a document as input on the vendor's own API, but Sume's wan-3.0 does not expose that field today. Alibaba Cloud describes a document-to-video feature that parses a file and generates a video from it; Sume's catalog lists wan-3.0 for text, first and last frames, and image, video and audio references, and notes that file_url, web_url and enable_thinking are not exposed in v1.
The Alibaba details are from its Wan3.0 guide and API reference; fal's Wan 3 page says documents need thinking mode. The Sume side is from Video generation and the current video catalog. All read 2026-09-29.
What can Wan 3.0 do with a document?
Per Alibaba, you provide one document or one public link and the model parses the content to build the video. The limits it publishes:
| Item | Alibaba's stated limit |
|---|---|
| File formats | docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md |
| File size | No more than 100 MB |
| File length | No more than 50 pages |
| Per request | One file or one link, not both |
Can I send a document to wan-3.0 on Sume?
No. The /v1/videos request has prompt, frame_images and input_references; a reference is an image_url, video_url or audio_url. There is no field for a document, so a PDF or a slide deck has nowhere to go on this route.
How do I get a video from a document on Sume, then?
Read the document yourself and turn it into a prompt. Extract the text, pick the three or four scenes worth showing, and write one prompt per scene; export the charts or pages you want on screen as images and pass them as input_references or a first_frame. PDF to video AI shows the agent route, where the extracted text goes in input.
Which route when?
- You want a clip per scene, with control of each prompt:
wan-3.0onPOST /v1/videos, 2 to 30 seconds each. - You want one call from a brief and material to a finished video: a Format or an Agent Completion, with the text in
input. - You need the vendor's own document parsing: that is Alibaba's API, which Sume does not front today.
Sources
Related posts
More in Models
- Wan 3.0 first and last frame: animate between two images
Wan 3.0 can make a clip that starts on one image and ends on another. On Sume, send frame_images with first_frame and last_frame to wan-3.0; last needs first.
- Wan 3.0 omni reference video: images, clips and audio in one prompt
Omni reference means one request carries mixed references. Wan 3.0 takes up to 10 images, 5 videos and 5 audio tracks; on Sume, send them as input_references.
- Wan 3.0 reference limits: how many images, videos and audio clips
Wan 3.0 reference-to-video takes up to 10 images, 5 video clips and 5 audio tracks. See the per-type caps, what Sume rejects with unsupported_capability, why.
- Wan 3.0 thinking mode: why document and web page input needs it
fal says Wan 3.0 needs thinking mode to read a document or web page; Alibaba says prompt_extend must be true. What each means, and why Sume's wan-3.0 skips it.
Written by Sume