Gemini 3.5 Transcribe Live pricing and Sume's async STT job
Google lists Gemini 3.5 Transcribe Live at about $0.009 per minute for streaming audio. Sume's STT is an upload-and-poll job priced per audio minute.

Google's pricing page lists Gemini 3.5 Transcribe Live at an estimated blended rate of about $0.009 per minute, for bidirectional streaming over WebSockets. Sume's speech transcription is a different shape: you send a finished audio file as a job and it is priced per audio minute, at $0.01 per audio minute on the rate card.
Both are per-minute numbers, but they buy different things, so compare the workflow before the rate. Google figures are from its page and Sume figures from the rate card and Video inspect docs, read 2026-09-30.
How does Google price the Live model?
The page shows per-token prices: $3.50 per 1M input tokens (about $0.005 per minute of audio) and $21.00 per 1M output tokens (about $0.004 per minute of text). It states the estimate uses 25 audio tokens per second for input and 175 text tokens per minute for output, giving about $0.009 per minute.
How is Sume's STT job billed?
Sume STT 1.0 is priced per audio minute after Sume margin. You can send a duration_seconds hint; if you omit it, Sume reserves 1 minute, and the maximum hint is 600 seconds. Paid usage is reserved at submit and captured on completion, as the generation admission docs describe.
| Gemini 3.5 Transcribe Live | Sume STT 1.0 | |
|---|---|---|
| Input | Streaming audio over WebSockets | Uploaded audio, submitted as a job |
| Unit | Audio and text tokens, about $0.009 per min blended | Per audio minute, $0.01 per audio minute |
| Duration known up front | No, the stream runs as long as you hold it | Optional hint, 1 minute reserved if omitted, max 600 s |
Which one fits my use case?
If you need words while a person is still speaking, streaming is the shape Google's Live model is built for. Sume does not document a streaming transcription endpoint, so for live captions use a streaming service. If you have recordings, a job with a duration hint gives you a known reservation before work starts. For long files see transcribe long audio files.
What should I check before comparing prices?
Confirm the unit (tokens versus audio minutes), whether output tokens are billed, and the file length limits. Sume's own limits are covered in Gemini 3.5 Transcribe limits versus Sume STT. Read the live rate with GET /v1/catalog rather than copying a number from a blog.
Sources
Related posts
More in Pricing
- Gemini 3.8 Flash TTS price: audio-time vs per-character
Gemini 3.8 Flash TTS Standard costs $0.00225 per 10 s of audio until Dec 31, 2026, then $0.0045. Sume bills TTS per character. How to compare.
- Gemini tiers unlock by spend and days. Sume limits follow the plan
Gemini upgrades a project after spend and elapsed days. Sume sets request and concurrency limits by plan, and a prepaid top-up does not raise concurrency.
- Genjutsu 1080p: Sume serves it at 480p or 720p only
Higgsfield lists Genjutsu output up to 1080p. On Sume, Genjutsu Motion Transfer accepts 480p (default) or 720p, and any other value is rejected.
- GPT Image 2.5 reference image cost: Runway credit vs Sume tokens
Runway charges 1 credit per GPT Image 2.5 reference image, once per request. Sume prices reference images as input image tokens inside the endpoint price.
Written by Sume