Nano Banana consistent characters and 14 objects: the Sume cap
Google says Nano Banana 2 keeps up to 5 characters and 14 objects consistent. Sume takes 10 references per request, so 19 subjects need a plan.

You cannot send 5 characters and 14 objects in one Sume request, because Sume's Nano Banana models take at most 10 input_references per call. Google's Nano Banana 2 announcement says the model keeps up to 5 characters and 14 objects consistent, which is 19 subjects. To use that many on Sume, split the scene across two requests, or fold several objects into one reference image.
Google's numbers are from its Nano Banana 2 post of 2026-02-26. Sume's are from the Image API docs and catalog code, read 2026-09-29.
How many references does Sume accept?
The catalog code sets one reference ceiling for every edit-capable model: 10. Only the two GPT Image 2.5 ids get 16. Both google/nano-banana-2 and google/nano-banana-pro therefore publish a maximum of 10 in their input_references descriptor.
| Item | Count |
|---|---|
| Characters Google says are kept consistent | up to 5 |
| Objects Google says are kept consistent | up to 14 |
| Subjects if you use both maximums | 19 |
input_references on Sume Nano Banana models | 0 to 10 |
Which subjects should get their own reference?
Spend the ten slots where drift hurts most. Faces and branded products give the most visible failures, so give each character and each logo-bearing object its own image. Plain props such as a mug or a chair can share a slot or move to the second pass.
How do I split the scene across two requests?
Generate the scene with the characters first. Then edit the result: pass the first output URL plus the remaining objects as references and set aspect_ratio to auto so the shape follows the reference. Reference URLs must be public HTTPS, and generated images come back as Sume-hosted URLs in data[].url.
IMG=$(curl -s -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"google/nano-banana-pro","prompt":"Both people from the references at a cafe table","input_references":[
{"type":"image_url","image_url":{"url":"https://example.com/anna.png"}},
{"type":"image_url","image_url":{"url":"https://example.com/ben.png"}}]}' \
| jq -r '.data[0].url')
curl -X POST https://api.sume.com/v1/images \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\":\"google/nano-banana-pro\",\"prompt\":\"Keep the people, add the laptop and the tote bag from the references\",\"aspect_ratio\":\"auto\",\"input_references\":[
{\"type\":\"image_url\",\"image_url\":{\"url\":\"$IMG\"}},
{\"type\":\"image_url\",\"image_url\":{\"url\":\"https://example.com/laptop.png\"}},
{\"type\":\"image_url\",\"image_url\":{\"url\":\"https://example.com/tote.png\"}}]}"Can I fold several objects into one reference image?
It is the other way to stay under ten. Lay out several small objects on one clean sheet, upload it as a single reference, and name each object in the prompt by its position. This is a prompting technique, not a documented Sume feature, and how well the model separates the objects is something to test on your own products before you rely on it. Keep one image per character either way.
Whichever route you take, write down which reference carries which subject. When a result drifts, that list tells you which slot to replace instead of rerunning the whole scene blindly.
Does the second pass keep the first pass intact?
Sume's docs make no consistency guarantee, so check it. Compare the second result against the first for faces and logos, and regenerate if something moved. Google's figures describe Google's model in its own products, not a promise about what a Sume request returns. Read the input_references range from GET /v1/images/models before you plan a scene, since the catalog is the source of truth.
Sources
Related posts
More in Use cases
- Fundraising video ideas for nonprofits, made with AI
Fundraising video ideas built on real material: one true story, a program update, the ask and the link. Where AI helps, where it must stop, and what it costs.
- On hold music for business: make the music and messages
On hold music for a business is a calm track, often with spoken messages, played to waiting callers. How to generate both, in phone-ready formats.
- Online course promo video with AI: what to put in it
An online course promo video says who it's for, what they'll learn, what's inside, who teaches it, and where to enroll. How to cut one from AI lessons.
- Perfume ad AI video: from one bottle photo to a short spot
Make a perfume ad with AI video: animate a real bottle photo with one move per shot, add a music bed, and check the glass and label in every frame.
Written by Sume