frame_images plus input_references in one request: which wins?
When a Sume video request carries both frame_images and input_references, frame_images wins and the job runs as image-to-video. What that means for references.

If a POST /v1/videos request carries both frame_images and input_references, frame_images takes precedence and the job is treated as image-to-video. The references are then not used as the generation mode, so send only one of the two.
What is the difference between the two fields?
frame_images gives first or last frame images for image-to-video; each entry needs a frame_type of first_frame or last_frame. input_references gives style or content references for reference-to-video, and the model uses them as visual guidance rather than exact frames.
Which mode does each field trigger?
The docs tie each field to one generation mode.
| You send | Mode |
|---|---|
| frame_images only | image-to-video |
| input_references only | reference-to-video |
| both | image-to-video (frame_images takes precedence) |
How do I send a first frame?
Use a public HTTPS image URL and mark it first_frame. The example follows the docs shape.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: frame-001" \
-d '{"model":"seedance-2","prompt":"A character walking through a forest","frame_images":[{"type":"image_url","image_url":{"url":"https://example.com/first-frame.png"},"frame_type":"first_frame"}],"resolution":"1080p"}'Does every model accept every reference type?
No. Only models whose supported_input_references lists a type accept it, and supported_frame_images says which frame_type values a model takes. Check GET /v1/videos/models first; details are in the video docs.
What happens after the job is accepted?
The submit response is 202 Accepted with an id, polling_url, status of pending and the model. Video generation is asynchronous: submit, receive a job id and polling URL, poll until completed, then download from the content URL. The docs suggest a polling interval of about 30 seconds and say generation typically takes 30 seconds to several minutes depending on the model and parameters.
If a request fails, check that any image is reachable over public HTTPS and in a supported format, as the troubleshooting notes advise.
Sources
Related posts
More in Developers
- Gemini video understanding API vs Sume video inspect stills
Gemini's API now has agentic video understanding. Sume's video inspect is narrower: probe facts, 8 stills and optional STT you pass to your own model.
- Gemini CLI MCP timeout default vs Sume's 55-second jobs_wait
Gemini CLI's MCP timeout defaults to 600,000 ms. Sume's jobs_wait holds at most 55 seconds per call: repeat it on wait_slice_expired, never resubmit.
- Gemini CLI keeps the OAuth refresh token: what Sume's MCP does
Gemini CLI 0.62 retains the OAuth refresh token on refresh. Sume's MCP token endpoint accepts only authorization_code, so expect a fresh sign-in, not a refresh.
- Gemini CLI trust: true on a Sume MCP server: what it skips
Gemini CLI's trust setting bypasses every tool confirmation dialog. What that means for Sume's paid tools, and which Sume gates still apply.
Written by Sume