Remove filler words from a talking video with an API
Sume has no one-call filler remover. Transcribe with video inspect for word timings, then cut the clean ranges with video trim at $0.02 per job.

Sume has no single call that strips filler words. You can build it from documented parts: POST /v1/video-inspect with transcribe: true returns word timings, and POST /v1/video-trim cuts each range you keep into a new MP4. You decide which words are filler, and you join the pieces yourself.
What does HeyGen's Speech Cleanup do?
HeyGen's June 2026 notes say it "removes every filler word, awkward pause, and false start, then stitches the remaining footage into a single seamless take with no visible jump cuts." That automated stitching is HeyGen's feature; Sume's docs do not describe an equivalent.
How do I get word timings?
Send the clip, hosted on media.sume.com, to video inspect with transcribe: true. The result carries transcript with text, words[] and optional sentence segments[]. Transcription is billed at $0.01 per audio minute; probe and stills are unbilled. A clip with no audio track fails as inspect_source_has_no_audio. Details: Video inspect.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: filler-inspect-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"duration_seconds": 120
}'How do I cut the clean ranges?
Scan words[] for the words you count as filler, and take the gaps between them as keep ranges. Send each keep range to video trim with start and exactly one of end or duration. The default precision is exact, a frame-accurate re-encode, which suits speech cuts. Then place the returned MP4s in order on a Timeline.
| Step | Route | Documented price |
|---|---|---|
| Word timings | POST /v1/video-inspect | $0.01 per audio minute |
| Cut one range | POST /v1/video-trim | $0.02 per job |
| Join the pieces | Timeline 1.0 | See the Timeline docs |
What are the limits?
Source video must be at most 1800 seconds for both routes, a trim output is 0.2-900 seconds, and each trim is one job, so a talk with 40 cuts is 40 jobs. This is a build-it-yourself recipe, and the docs make no promise of seamless joins.
Sources
Related posts
More in Developers
- Remove filler words from a video by API: cut at word timestamps
Descript's API lists Remove Filler Words as an Underlord edit. the Sume docs list no such op; here is how to cut um and uh yourself from words[] and a Timeline.
- OpenAI response_format json_schema on a Sume scheduled run
Sume accepts an OpenAI-shaped response_format as an alias for output_schema on schedule runs. Sending both returns 400, and the schema must be strict.
- Responses steering after a video job started: what Sume does
response.steer does not cancel started tools. A Sume video job past its start runs to completion; jobs_cancel returns 409 job_generation_already_started.
- Change caption style without transcribing again: source_caption_id
Pass source_caption_id instead of video_url to re-burn a clip under a new style. Sume reuses the stored word timings, so speech-to-text runs only once.
Written by Sume