Automate video editing in Python with an editing API
Automate video editing in Python by calling an editing API with Requests: submit a caption, cut, or crop job, poll until it ends, chain the output.

To automate video editing in Python, either drive a local library and FFmpeg on your own machine, or call a video editing API from a script. With an API, the script sends each edit (a caption, a cut, a crop) as a job, waits for it to finish, and feeds the output file into the next edit, with no FFmpeg install and no rendering on your hardware.
This post shows the API route with Sume and the Requests library. The Sume facts come from the Video captions, Video trim, Video filter, and Jobs and results docs and the Sume API reference, read on 2026-09-29. Anything described as current behavior is read from Sume's code.
Should I edit locally or call an API from Python?
Edit locally when your files live on your machine, you need full FFmpeg control, and you have the CPU to spare. Call an API when you'd rather not install or scale an encoder, when the edits are standard (cut, crop, caption, overlay, join), and when the source videos are already online. The trade-off with Sume: its cut, crop, and join tools read only files already on Sume's media host, so a pipeline starts from a public URL that the caption job can take, or from an earlier Sume output.
What does an editing script look like?
One helper submits a job and polls it; each edit is one call. This one captions a public clip, then cuts the captioned file to its first 15 seconds; the caption job's output is a media.sume.com file, which in current code the trim tool admits. Keep the key in a server-side environment variable, never in browser or mobile code.
import os, time, requests
API = "https://api.sume.com/v1"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def call(method, url, key=None, **kw):
headers = {**AUTH, **({"Idempotency-Key": key} if key else {})}
r = requests.request(method, url, headers=headers, timeout=30, **kw)
r.raise_for_status()
return r.json()["data"]
def run(path, body, key):
job = call("POST", f"{API}/{path}", key, json=body)
status = call("GET", job["status_url"])
while not status["terminal"]:
time.sleep(status["next_poll_after_seconds"] or 5)
status = call("GET", job["status_url"])
if status["sume_status"] != "completed":
raise RuntimeError(f"{path}: {status['sume_status']}")
return job
cap = run("video-captions", {"video_url": "https://example.com/clip.mp4"}, "clip-cap-1")
captioned = call("GET", f"{API}/video-captions/{cap['video_caption_id']}")["video_url"]
cut = run("video-trim", {"video_url": captioned, "start": 0, "duration": 15}, "clip-cut-1")
print(call("GET", cut["result_url"])["result"]["video_url"])Why poll the status before reading the result?
Because only a completed job has a result. terminal is true once the job is completed, failed, or canceled, and sume_status says which. In current code GET /v1/jobs/{id}/result answers job_not_completed, marked retryable, even for a failed or canceled job, so a script that retries the result call alone can loop forever. Store each job id as soon as the submit returns, so a restarted script resumes polling instead of paying for the edit again.
Send an Idempotency-Key on every submit and reuse it, with the same body, only when retrying after a timeout or network failure. To skip polling on long pipelines, pass a webhook_url and receive a signed callback when each job ends.
Which edits can the script call, and what do they cost?
Each call is a separate job, billed on its own, plus a 5.5% agent fee by default. The media tools reserve their listed rate and capture their own compute, never above the reservation.
- In current code a caption job refuses a source over 60 seconds or one with no audio stream.
- Trim and filter return a new MP4 and leave the source untouched; filter programs can be checked free at
POST /v1/video-filter/checkfirst. - There is no public route to upload a file from your computer, so local footage has to be reachable at a public HTTPS URL for the caption step.
| Edit | Endpoint | Input it takes | Price |
|---|---|---|---|
| Burn captions | POST /v1/video-captions | Public HTTPS URL, ≤ 60 s, with sound | $0.20 per job |
| Cut a range, resize | POST /v1/video-trim | Sume-hosted file | up to $0.02 per job |
| Crop, dim, other allowlisted filters | POST /v1/video-filter | Sume-hosted file, ≤ 300 s | up to $0.02 per job |
| Logo over video | POST /v1/timeline-1.0/compose | Sume-hosted files | up to $0.02 per job |
| Join clips over audio | POST /v1/timeline-1.0/render | Sume-hosted files | up to $0.10 per output minute |
Sources
Related posts
More in Developers
- Bash for loop with curl: one API request per line
Loop over a file with while IFS= read -r, build each JSON body with jq --arg, send it with curl --fail-with-body, and pace it under the API's rate limit.
- Batch image generation API: how many images can run at once?
Sume queues extra generation jobs instead of rejecting them. The per-plan concurrency and queue table, the 429 queue_full case, and how to size an image batch.
- Caption API script_alignment_mismatch: what it means and how to fix it
If script_text mismatches the speech, POST /v1/video-captions fails with script_alignment_mismatch or script_alignment_failed. Simplify or omit it.
- claude -p with --mcp-config: run Sume video tools from a script
Use claude --bare -p with --mcp-config and --allowedTools to call Sume's hosted MCP tools in CI; read mcp_server_errors so an unloaded server fails the job.
Written by Sume