AI product exploded view generator: two frames, real parts
An AI exploded view animation needs two stills: the product assembled and exploded. A video model makes the motion; the parts must come from you.

An AI product exploded view generator works from two stills: the assembled product as the first frame and the exploded view as the last, with an image-to-video model generating the motion between them. Swap the two for the assemble shot. The model cannot know what is inside your product, so if the parts must be real, render the exploded still from your CAD file and let AI make only the motion.
Below is how to do it with Sume's video API. Facts come from the Video generation, Image generation, Video frames and Timeline 1.0 docs, read on 2026-09-29, with prices read from Sume's pricing code.
How do I make an exploded view animation with AI?
- Make the assembled still: a clean product shot on a plain background.
- Make the exploded still in the same framing, camera angle and light. A CAD render is the only way to get the real parts in the real order.
- Host both stills at public HTTPS URLs. Sume's media-input docs say localhost, private-network and signed URLs are rejected before generation.
- Send both to
POST /v1/videosinframe_images: the assembled still asfirst_frame, the exploded still aslast_frame. Pick a model whosesupported_frame_imagesonGET /v1/videos/modelslistslast_frame; the docs'seedance-2entry lists both. - Keep the prompt to one motion with the camera locked, because every frame between your two stills is generated.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: explode-001" \
-d '{
"model": "seedance-2",
"prompt": "Locked-off studio shot. The parts separate slowly along one axis and stop, evenly spaced. No new parts appear.",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/assembled.png" }, "frame_type": "first_frame" },
{ "type": "image_url", "image_url": { "url": "https://example.com/exploded.png" }, "frame_type": "last_frame" }
],
"duration": 5,
"resolution": "720p",
"aspect_ratio": "16:9"
}'Can AI draw the exploded view for me?
It can draw one, but the parts are invented. An image model given your product photo as a reference in input_references on POST /v1/images sees only the outside; the screws, boards and cells it draws come from the prompt, not from your product. That is fine for a concept or mood piece. It is not fine for a spec sheet, a manual, or anything a buyer reads as how the product is built.
The video model has the same limit in the frames between your stills: a part can change shape mid-move or a new one can appear. Two stills from your own renders pin both ends, which is why the method above starts there. Product logo warping in image-to-video covers the same frames-versus-references choice for labels.
How do I make it come back together?
Make a second clip with the frames reversed: exploded still as first_frame, assembled still as last_frame. To play explode then assemble as one file, join the two clips with POST /v1/timeline-1.0/render. A render needs an audio spine, or audio.mode: "silence" for a declared length with no sound, and takes only this workspace's media.sume.com files, such as the clips Sume returned. In current code each clip's own sound is dropped, so any sound comes from the spine or a soundtrack. The default output is 1080×1920 (vertical), so set output.width and output.height (for example 1280 and 720) to keep a 16:9 shot.
How do I check the animation before I use it?
Pull stills from the clip with POST /v1/video-frames and a list of times in at[], then compare each part against your product or CAD model. The extract is unbilled and returns durable media.sume.com images. If a part changes shape or appears from nowhere, generate again or shorten the move.
How much does it cost, and what are the limits?
| Step | Price | Limit |
|---|---|---|
One clip, POST /v1/videos | Provider list × 1.25; $1.89 for 5 s of seedance-2 at 720p 16:9 | Most catalog models top out at 15 s |
Check stills, POST /v1/video-frames | Unbilled | One media.sume.com clip per extract |
Join, POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes | 1–200 slots, 1–1800 s |
Sources
Related posts
More in Use cases
- AI fitness video maker: what to generate, what to film
An AI fitness video maker can make the coach, intro, B-roll and on-screen text, but not trustworthy exercise demos. What to film, what to generate.
- AI game music generator: one track per level, cut to WAV
An AI game music generator makes each level, menu or boss track from a text brief. How to brief contrasting scenes, get WAV files, and what it costs.
- AI image for Instagram 4:5 (1080×1350): what the API returns
Instagram portrait is 4:5, 1080×1350, not 4:3. How to ask Sume's image API for 4:5 and why some models return a native size close to it, not exact pixels.
- Can AI make a lyric video? Yes, if you time the lines
AI can make a lyric video: pictures under the song, plus each lyric line burned in at the time it is sung. You supply the line timings.
Written by Sume