AI architecture video generator: renders to moving shots
An AI architecture video generator animates a still render: the render is the first frame, the prompt names one camera move. Limits, walkthroughs, cost.

An AI architecture video generator turns a still render or photo of a building into a short moving shot: the render becomes the clip's first frame, and the prompt names one camera move, such as a slow dolly in, a partial orbit or a crane up. Clips are short, so a walkthrough is several shots joined in an edit. The result is generated video, not a render of your model: everything beyond the frames you pin is invented, so it can't stand in for a walkthrough that has to match the plans.
The Sume steps below come from the Video generation and Timeline 1.0 docs, read on 2026-09-29. Limits marked as current behavior are read from Sume's code.
How do I turn an architectural render into a video?
Put the render at a public HTTPS URL and send it in frame_images with frame_type first_frame on POST /v1/videos. That makes the job image-to-video, so the shot opens on your exact image. The request has no camera setting, so the move goes in the prompt, one move per shot; AI video camera movement prompts has the vocabulary.
- To end the move on a second render, say a closer view of the entrance, add it as a
last_frameentry. In current codelast_framewithoutfirst_frameis refused. - Send
resolutionexplicitly, and read each model'ssupported_frame_imagesbefore relying on a last frame. - Wait for
completedon the polling URL, then download fromunsigned_urlswith your API key.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: facade-dolly-001" \
-d '{
"model": "seedance-2",
"prompt": "Slow dolly in toward the entrance of the building at dusk. The building does not change; only the camera moves. No people, no cars.",
"duration": 6,
"resolution": "1080p",
"aspect_ratio": "16:9",
"generate_audio": false,
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/renders/facade.png" }, "frame_type": "first_frame" }
]
}'How long can an AI architecture video be?
Each clip is capped by its model, so check GET /v1/videos/models before planning shots:
| Setting | What the docs say |
|---|---|
| Length per clip | seedance-2.5 4–30 s, wan-3.0 2–30 s; the others at most 15 s (supported_durations per model) |
| Aspect ratios | Include 16:9, 9:16, 1:1 and ultra-wide 21:9; each model lists its subset in supported_aspect_ratios |
| Resolutions | From 480p to 4K; each model lists its subset in supported_resolutions |
| Sound | generate_audio defaults to the model's audio capability; send false for a silent shot |
Can AI make a full architectural walkthrough?
Only as a sequence of short generated shots, joined in order: approach, entrance, lobby, a room, the view. The joining step is the same as for a listing tour from photos: the generated clips go in order into one Timeline 1.0 render (POST /v1/timeline-1.0/render) of up to 200 slots and 1,800 seconds. What changes for design renders:
- Each shot starts from its own render, so consecutive shots don't share a continuous camera path. To carry one move into the next shot, see chaining clips from a last frame.
- In current code the render keeps only the audio spine and the soundtrack, so any sound a clip generated is dropped. Put narration or music on the spine instead.
- The default output is 1080×1920, vertical. For a 16:9 presentation, set
output.width1920 andoutput.height1080.
Is an AI architecture video accurate enough to show clients?
Treat it as mood, not documentation. Only the frames you pin are your images; the model fills in everything between them.
- Anything the render doesn't show, such as a side facade or the room behind a door, is invented. A long orbit shows more of it.
- Geometry, materials, window counts and signage can drift between frames. Watch every clip before it reaches a client.
- Don't rely on it for dimensions, finishes or anything a buyer or planner might treat as a commitment. For off-plan marketing, label the footage as illustrative.
- A walkthrough that must match the drawings still comes from your 3D tool; an AI clip can be the moving teaser around it.
How much does an AI architecture video cost?
You pay per step from your workspace USD balance: one video job per shot, an optional narration, and one render per cut.
| Step | Call | Price |
|---|---|---|
| Animate each render | POST /v1/videos | By model, reserved at provider list × 1.25 per clip; see pricing_skus on GET /v1/videos/models |
| Narration (optional) | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join the shots into one MP4 | POST /v1/timeline-1.0/render | $0.10 per output minute |
Sources
Related posts
More in Use cases
- AI car video generator: from a car photo or a prompt
An AI car video generator animates a photo of your car or invents one from a prompt. How to get a driving shot, what drifts, and what it costs.
- AI Christmas video generator: a greeting that moves
An AI Christmas video is a few animated clips from a photo or a winter scene, an original instrumental, and your greeting burned on as typed text.
- AI product exploded view generator: two frames, real parts
An AI exploded view animation needs two stills: the product assembled and exploded. A video model makes the motion; the parts must come from you.
- AI fitness video maker: what to generate, what to film
An AI fitness video maker can make the coach, intro, B-roll and on-screen text, but not trustworthy exercise demos. What to film, what to generate.
Written by Sume