ElevenLabs STT 3 GB / 10 hour limit vs Sume STT 600 s jobs

ElevenLabs speech to text accepts files up to 3 GB and 10 hours in standard mode. Sume STT 1.0 reserves up to 10 minutes per job, so cut long recordings.

4 min readSume
All posts

ElevenLabs' speech-to-text FAQ says files up to 3 GB are supported, and standard mode takes up to 10 hours. Sume STT 1.0 plans around much shorter jobs: the duration_seconds hint is 1 to 600, and the OpenAPI text says to omit it to reserve for 1 minute, with a maximum of 10 minutes. The schema describes that hint as a usage reservation rather than a stated audio-length cap, so cutting a long recording into pieces of 600 seconds or less is the safe plan, not a documented hard limit.

Vendor numbers are from the ElevenLabs page; Sume numbers from the OpenAPI reference and the Audio detach docs, read 2026-10-01.

What are the ElevenLabs limits?

The page lists a 3 GB maximum file size, and a maximum duration of 10 hours in standard mode (use_multi_channel=false). In multi-channel mode the combined duration of all channels must be under 10 hours, and the key facts list says 1 hour. Both audio and video files are accepted.

What are Sume's numbers?

Limits compared, read 2026-10-01.
LimitElevenLabs (page)Sume
Max file size3 GBNot stated for STT; send a public HTTPS audio_url
Duration per request10 hours standardduration_seconds 1 to 600
Default reservationBilled on audio duration1 minute if the hint is omitted
RateSee ElevenLabs pricing$0.01 per audio minute

How do I transcribe a long recording on Sume?

Split it into chunks of 600 seconds or less, send one STT job per chunk with its duration_seconds, and add each chunk's start offset to its word timings when you join the results. The audio-detach route can produce mono 16 kHz audio for STT and accepts a range, but its source cap is 1800 seconds, so a multi-hour file has to be cut before it reaches that route. See audio detach for speech-to-text.

Word timings are always returned, so each chunk gives you per-word starts and ends to shift.

Where can this go wrong?

Cut points mid-word can split a word across two jobs, so cut on silence if you can. Keep the chunk list and each chunk's offset with your job ids, and set language_code when you know the language; omit it for auto-detect.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume