AssemblyAI Sync API: one call, and Sume STT sync mode
AssemblyAI's Sync API returns a short-clip transcript in one POST. Sume STT mode sync waits up to 30 seconds, then returns the job id to poll if not done.

AssemblyAI's Sync API is a single HTTP POST that returns a finished transcript for a short clip in the same response. Sume's closest match is mode: "sync" on STT 1.0: the request blocks for at most wait_timeout_seconds (max 30) and, if the transcript is not ready, returns the job so you can read it by id. It is a bounded wait on a job, not a separate endpoint.
AssemblyAI facts are from its launch post; Sume facts from the OpenAPI schema and docs, read 2026-10-01.
What does AssemblyAI say the Sync API does?
The post says to send a short audio clip in one HTTP request and get a finished Universal-3.5 Pro transcript back in the same response, "in ~134 ms" at p50 for a 2-second clip, with no polling, no WebSocket and no job to manage. That figure is AssemblyAI's own and says nothing about Sume.
What does Sume sync mode do?
| Question | AssemblyAI Sync API | Sume STT 1.0 sync |
|---|---|---|
| Transport | One HTTP POST | One HTTP POST with mode: "sync" |
| Result in the response | Yes, for short clips | If terminal within the wait; otherwise the current job state |
| Wait cap | Not stated in the post | wait_timeout_seconds, max 30 |
| Job id | "No job to manage" | Returned in the first response in every mode |
| Word timings | Not covered in the post | Always returned |
What if the transcript is not ready in 30 seconds?
The response is still 2xx and carries the current queued or running state. Keep the job id and read it, as in the 30-second wait guide. The docs also say sync and subscribe are the same bounded wait and recommend async or webhook for new integrations.
How do I make the one call?
Send audio_url, set duration_seconds for the real clip length, and reuse the same Idempotency-Key on a retry. Billing is $0.01 per audio minute on the rate card.
const res = await fetch("https://api.sume.com/v1/stt-1.0/transcribe", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": "dictation-0001",
},
body: JSON.stringify({
audio_url: "https://media.sume.com/artifacts/example/clip.m4a",
duration_seconds: 5,
mode: "sync",
wait_timeout_seconds: 30,
}),
});
console.log(res.status, await res.json());Sources
Related posts
More in Developers
- Avatar video captions error over 60 seconds: split the script
Inline captions on an avatar video are rejected when the estimated duration is over 60 seconds, the same cap as the job. Split long scripts into jobs.
- Azure batch takes 10,000 inputs per job; Sume sends one text per job
Azure batch synthesis accepts up to 10,000 text inputs in a 2 MB JSON body. Sume TTS 1.0 takes one transcript of up to 20,000 characters per job.
- Azure batch: 95% of outputs within 120 seconds; Sume TTS waits 30
Azure says half of batch outputs finish in 10 to 20 seconds and 95% within 120. Sume TTS sync mode waits at most 30 seconds, then hands you a status URL.
- Azure word boundaries in ms vs Sume TTS timestamps in seconds
Azure writes word timings as AudioOffset and Duration in milliseconds, in a separate file. Sume returns words[] with start and end seconds on the job result.
Written by Sume