Word document to video AI: turn a .docx into a narrated video

To turn a Word document into a video, extract its text and images first. Sume's agent takes text in input and images as attachments; a .docx isn't one.

4 min readSume
All posts

A Word document becomes a video in three steps: get the text and pictures out of the .docx, write a short brief for the video, and send both to a video agent. On Sume, a run takes the text in input (up to 2 MiB) and up to 30 images as attachments; a .docx is not an attachment type, so extract it first.

Limits are from the Format API and Create a run docs, read 2026-09-29. Alibaba's Wan 3.0 API lists docx among the files it can parse (API reference); Sume's does not pass a document to a model.

Where does each part of the document go?

From Create a run and the Format API, read 2026-09-29.
Part of the .docxWhere it goesLimit
Body text and headingsinput, as JSON2 MiB, at most 64 top-level keys
Pictures and chartsattachments, type input_image30 images; JPEG, PNG, WebP, GIF, AVIF; 30 MB each
What to makeinstruction8,000 accepted; about 4,000 reach the prompt

How do I extract the text?

Use any docx reader in your own code, or save the file as plain text. Keep headings, drop headers, footers and page numbers, and put the result in a key you choose, such as input.text. Sume publishes no fixed field list for input; the agent reads it as caller data.

How do I send it?

One Agent Completion with a required generation_spend_cap_usd, then poll the run. PDF to video AI has the request body; a Word file goes in the same way once it is text.

What can go wrong with a Word file?

  • Tracked changes and comments can end up in the extracted text; accept or strip them first.
  • Tables flatten into lines; restate the numbers that matter in your brief.
  • Embedded images are not extracted for you; export the ones you want and host them at public HTTPS URLs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume