Polly speech marks need a separate request; Sume returns timings
Polly returns speech marks instead of audio when you ask for them: sentence, word and viseme metadata. Sume TTS returns words[] timings on the same job result.

Amazon Polly's docs say that when you request speech marks, Polly "returns this metadata instead of synthesized speech": where a sentence or word starts and ends in the audio, plus visemes for lip-sync. So you make one call for the audio and another for the marks. On Sume TTS 1.0, one request with timestamps.words: true returns the audio and words[] timings together on the completed job.
Polly's description is from its speech marks page, read 2026-10-01; Sume's from the OpenAPI schema behind the API reference.
Which engines support speech marks on Polly?
The page says speech marks are available with the neural, long-form or standard engines. It does not list generative voices there, so check the engine before you plan on marks.
What does Sume return, and what does it not?
| Item | Amazon Polly | Sume TTS 1.0 |
|---|---|---|
| Calls needed | One for audio, one for speech marks | One job |
| Word timings | Word speech marks | words[] with start and end seconds |
| Sentences | Sentence speech marks | segmentation.mode: sentence, gapless |
| Visemes | Viseme speech marks for lip-sync | Not in the TTS contract |
How do I do lip-sync without visemes?
Sume's lip-sync is a separate stage that takes audio, not mouth shapes. Use sentence segments to cut clips and feed each audio slice to the avatar or lip-sync step; see sentence segments to drive lip-sync clips.
What does the Sume request look like?
One call, wav output, word timings on.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Highlight each word as it is spoken.",
"avatar_handle": "@narrator",
"output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
"timestamps": { "words": true }
}'Sources
Related posts
More in Developers
- Polly takes 3,000 billed characters; Sume TTS takes 20,000
Polly SynthesizeSpeech accepts 3,000 billed characters and cuts audio at 10 minutes. Sume TTS 1.0 accepts 20,000 characters and 1,200 seconds of audio per job.
- Postgres 17.11 pgcrypto change: verify a Sume webhook with hmac()
The Postgres minor release changed legacy pgcrypto ciphers. Its listed changes never mention hmac(), so a Sume sume-v1 signature check in SQL still works.
- Punch-in zoom on video by API: crop or zoompan, no keyframes
Sume has no auto zoom switch. Use a crop op for a fixed punch-in or the allowlisted zoompan filter in a video-filter graph; there are no keyframes.
- remove.bg API rate limit: 500 per minute, weighted by megapixels
remove.bg allows 500 images per minute at about 1 MP, less for larger inputs. Sume limits requests per minute per key and sends retry-after on 429.
Written by Sume