Generate an image, then animate it: image to video in two API calls

Generate a still with POST /v1/images, then pass its URL as first_frame to POST /v1/videos. The two calls, the URL rule between them, and what each reserves.

4 min readSume
All posts

Make the still with POST /v1/images, take the URL from data[0].url, and send it as a first_frame in frame_images to POST /v1/videos. That is the whole handoff: two calls, one Sume-hosted URL passed from the first to the second.

Google describes the same pattern for its own models, a fast image model feeding its Omni Flash video model. Sume's request shapes come from the Image API and Video generation docs, read 2026-09-29.

What do the two calls look like?

The image call returns 200 with the URL inside 30 seconds, or 202 with a job envelope if it is still running; in that case poll the job result for the image URL first.

curl -s -X POST https://api.sume.com/v1/images \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "google/nano-banana-2", "prompt": "A red sports car on a coastal road at dusk", "aspect_ratio": "16:9"}'

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: handoff-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "Slow push in as the car pulls away",
    "frame_images": [{"type": "image_url", "image_url": {"url": "IMAGE_URL_FROM_STEP_1"}, "frame_type": "first_frame"}],
    "duration": 5,
    "resolution": "720p",
    "aspect_ratio": "16:9"
  }'

Do the ratios have to match?

Pick one ratio for both. The video model lists its own supported_aspect_ratios, so generate the still at a ratio the video model accepts, such as 16:9 or 9:16 for Seedance 2, rather than resizing later.

What does each step reserve?

Each call reserves its own estimate when you submit. If the wallet cannot cover it, that submit returns 402 insufficient_credits.

Sume estimates from pricing code, 2026-09-29. Provider list × 1.25, rounded up to a cent.
StepModelEstimate
Stillgoogle/nano-banana-2 at 1KSee the Nano Banana 2 price table
Clipgemini-omni-flash-1.1, 5 s at 720p$0.63

Which field is the still: frame_images or input_references?

frame_images makes the still the exact first (or last) frame, which is image-to-video. input_references treats it as guidance for reference-to-video. If you send both, frame_images wins and the request is image-to-video.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume