AssemblyAI Sync API file limit vs Sume STT duration_seconds
AssemblyAI's launch post frames its Sync API for short clips. Sume STT 1.0 takes duration_seconds from 1 to 600 and reserves one minute if you omit it.

AssemblyAI's launch post describes its Sync API as built for short clips and says the clip is decoded to a single 16 kHz audio array with no chunking; the post text I read states no byte or minute cap, so check AssemblyAI's own docs for the exact file limit. Sume's STT 1.0 numbers are explicit: duration_seconds takes 1 to 600, and omitting it reserves 1 minute.
AssemblyAI facts are from its launch post; Sume facts from the OpenAPI schema and docs, read 2026-10-01.
What does AssemblyAI say about clip length?
The post lists the Sync API for "short clips: dictation, voice agents with turn detection" and sends "long audio, batch, latency-tolerant" work to the Async API. It does not give a number in the part I read, so I do not quote one.
What are Sume's length numbers?
| Setting | Value | Source |
|---|---|---|
duration_seconds | 1 to 600; improves the usage reservation | STT schema |
Omitted duration_seconds | Reserves 1 minute | STT schema |
| Maximum hint | 10 minutes | STT schema |
| Audio Detach output | At most 900 s | Audio Detach docs |
| Rate | $0.01 per audio minute | Rate card |
What happens if I leave duration_seconds out?
The reservation is one minute. For a 5-minute recording, send duration_seconds: 300 so the reservation matches the audio. The schema says the field improves the usage reservation; it is a hint for billing, not a cut-off you set on the audio.
What about audio longer than 10 minutes?
Split it into parts under 600 seconds and submit each as its own job, then join the word lists by offset. If the source is a video, Audio Detach can extract a 16 kHz mono track first; see audio detach for speech-to-text. Its own output limit is 900 s, so long files need a range there too.
Sources
Related posts
More in Developers
- Avatar video captions error over 60 seconds: split the script
Inline captions on an avatar video are rejected when the estimated duration is over 60 seconds, the same cap as the job. Split long scripts into jobs.
- Azure batch takes 10,000 inputs per job; Sume sends one text per job
Azure batch synthesis accepts up to 10,000 text inputs in a 2 MB JSON body. Sume TTS 1.0 takes one transcript of up to 20,000 characters per job.
- Azure batch: 95% of outputs within 120 seconds; Sume TTS waits 30
Azure says half of batch outputs finish in 10 to 20 seconds and 95% within 120. Sume TTS sync mode waits at most 30 seconds, then hands you a status URL.
- Azure word boundaries in ms vs Sume TTS timestamps in seconds
Azure writes word timings as AudioOffset and Duration in milliseconds, in a separate file. Sume returns words[] with start and end seconds on the job result.
Written by Sume