ChatGPT image storyboard: make frames with GPT Image 2.5

Make a storyboard with GPT Image 2.5 by API: one call per frame, the same reference images each time, and a prompt that names what stays fixed.

4 min readSume
All posts

To make a storyboard with GPT Image 2.5, generate one frame per API call and send the same reference images with every call, so the character and look are anchored to the same source. Then write each frame's prompt as "same as the references, but" plus the new action. A single call does not produce a sequence.

The request fields come from the Image API docs, read 2026-09-29. OpenAI's launch page, read the same day, is the source for what it says about consistency. Neither promises that frames will match, so review each one.

How do I keep the character the same across frames?

Send a character reference in input_references on every call; both GPT Image 2.5 ids take up to 16. OpenAI's page says Images 2.5 preserves the subjects in reference photos, and that in longer ChatGPT conversations earlier changes are more likely to stay consistent. The second point is about the ChatGPT app. Over the API each call is separate, so the references and the prompt carry the continuity.

  • Reference 1: the character. Say "keep this person's face, hair and outfit" in the prompt.
  • Reference 2, if needed: the location or the style.
  • Prompt: state the camera and the action for this frame only.
  • Set aspect_ratio the same on every frame, for example 16:9.

What does a three-frame loop look like?

The loop below sends the same reference each time. Each response carries the image in data[].url; a 202 means a job to poll instead.

REF="https://example.com/character.png"
i=1
for ACTION in "walks into the shop" "picks up a jacket" "looks at the price tag"; do
  curl -s -X POST "https://api.sume.com/v1/images" \
    -H "Authorization: Bearer $SUME_API_KEY" \
    -H "Content-Type: application/json" \
    -d "{
      \"model\": \"openai/gpt-image-2.5\",
      \"prompt\": \"Storyboard frame $i. Same character as the reference. She $ACTION. Wide shot.\",
      \"input_references\": [{\"type\": \"image_url\", \"image_url\": {\"url\": \"$REF\"}}],
      \"aspect_ratio\": \"16:9\"
    }" > "frame-$i.json"
  i=$((i+1))
done

Should I use n to get all the frames at once?

No. n asks for several images of one prompt, which suits alternatives of the same frame, not a sequence. The docs allow n up to 10 per call but say per-model ceilings are lower, so read the n range descriptor in the catalog for the model you pin. Each image is billed on its own.

From Image API, read 2026-09-29.
NeedFieldLimit in the docs
Anchor a character or lookinput_referencesUp to 16 on both GPT Image 2.5 ids
Several takes of one framen1 to 10 per call; per-model ceilings are lower
Same shape on every frameaspect_ratio16:9, 9:16, 1:1, 4:3 and more
Slow rendermodeasync returns a 202 job to poll

What happens when I submit many frames at once?

Sume's admission docs say concurrency is a dispatch limit, not a submit limit: extra valid jobs can wait as queued while queue capacity remains, and only a full queue fails with 429 queue_full. Do not treat queued as a failure. See Batch image generation: how many can run at once.

How do I turn the frames into a video?

Each finished frame is an image URL you can use as the start of a shot. Storyboard to video AI covers that step. Result URLs from /v1/images are signed, so copy the files you want to keep.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume