Google Vids caption font, position, animation: Sume design overrides
Google Vids lets you change caption font, position and animation. Sume's caption overrides cover colour, size, placement, phrasing and motion, not a Latin font.

Google Vids captions can now be customized for font and typography, position on screen, and animation, according to a Google Workspace Updates post dated 30 September 2026. Sume's standalone caption job maps to two of those three directly: placement and motion are design overrides. Typography is partly covered, and a free choice of Latin font is not available, because the font field only takes Hangul faces.
How do Google's three controls map to Sume fields?
| Google Vids control | Sume field | Notes |
|---|---|---|
| Font and typography | design.typography | Weights, active scale, size ratio, safe width, stroke width. The font field is Hangul faces only. |
| Where captions sit | design.placement | anchor_ratio and landscape_anchor_ratio: the line centre as a fraction of frame height. |
| How they animate | design.motion | Enter, exit and emphasis timings in seconds. |
| Brand colours | design.colors | Base, active, stroke, accent and card; hex or rgba. |
Can I pick a brand font in Sume captions?
Not for Latin text. The docs list font as an optional Hangul face, and naming one beside a Latin style returns a 400 (caption_font_requires_hangul_style). The Latin styles slam, punch and tiktok-green draw in their own display faces. Google's post says the point of the update is matching captions to brand fonts and colors; in Sume you can match colours and weights, not the family.
Also note that design is not supported on punch or tiktok-green, so use slam (or a Hangul style for Korean speech) when you want overrides.
What does a request with overrides look like?
Each field is optional and merges over the style's own value, so one key changes one thing. Numbers outside their documented range come back as a 400 instead of rendering wrong, so check the response before batching. The video_url must be a fetchable public HTTPS URL, and a media.sume.com clip qualifies.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: caption-design-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/promo.mp4",
"style": "slam",
"language": "en",
"design": {
"colors": { "active": "#22D3EE" },
"placement": { "anchor_ratio": 0.75 },
"phrasing": { "max_words": 3 },
"motion": { "enter_seconds": 0.15 }
}
}'How do I try a second look without paying for a second transcript?
Pass source_caption_id instead of video_url. Sume reuses the first caption's source video and word timings, so no second speech-to-text runs. Billing is unchanged, because a restyle is still a render: the docs list $0.20 per accepted standalone caption job for videos up to 60 seconds, under the current fixed estimate. Confirm the live figure in GET /v1/catalog.
Sources
Related posts
More in Developers
- GPT-5.5 retires Oct 14 in Codex: what to change for Sume
Codex retires GPT-5.5 on October 14, 2026. A Sume server entry names no model, and Agent Completions only accepts sume-agent, so check your own settings.
- gpt-image-2.5-flare-2026-09-08 on Sume: use openai/gpt-image-2.5
OpenAI lists the snapshot gpt-image-2.5-flare-2026-09-08. Sume documents catalog slugs, so send openai/gpt-image-2.5 and check /v1/images/models for ids.
- GPT Image 2.5 Flare: 5 images a minute at Tier 1 vs Sume limits
OpenAI limits Flare to 5 images per minute at Tier 1 and 250 at Tier 5. Sume applies plan concurrency, queue capacity and 429 codes instead. Both tables here.
- GPT Image 2.5 largest square: 2880x2880, not 3000x3000
GPT Image 2.5 caps pixels at 8,294,400. A square tops out at 2880x2880; 3000x3000 (9,000,000) is rejected. OpenAI calls sizes above 2560x1440 experimental.
Written by Sume