OpenRouter Batch API vs Sume async jobs: which for video?

OpenRouter's Batch API is for text and embeddings with a 24 hour window. For video files, Sume returns a job id per request: async, sync, or webhook mode.

4 min readSume
All posts

For video, use per-job async submission, not a batch file. OpenRouter's Batch API, announced for more than 70 models, bundles requests and lets the provider finish them anywhere in a 24 hour window; its announcement lists chat completions, responses, messages and embeddings as the supported shapes. Sume has no batch file: each video request is its own job with a job id returned in the first response.

OpenRouter facts are from its announcement; Sume facts from Jobs and results and Generation admission, read 2026-10-01.

What does the OpenRouter Batch API do?

You call api/v1/batches with a list of requests and an endpoint shape, then poll GET /api/v1/batches/:id until the status is completed, failed, expired or cancelled. Completed batches return results inline. In exchange for the open completion window, the announcement says providers generally charge 50% (and sometimes less) of their normal per-token price. It also reports that, over its two-week beta, the median batch finished in 7 minutes and 90% within an hour.

How do Sume's submit modes differ?

Every Sume submit endpoint takes a mode. The docs say the mode decides how you learn the outcome; it never changes whether a job is created, what it costs, or how long it runs.

Sume communication modes, from Jobs and results, read 2026-10-01.
ModeServer blocksWhat you do next
async (default)No, returns 202 with the envelopePoll status_url until terminal is true, then read result_url
syncUp to 30 secondsNot terminal yet? Poll the job id; do not resubmit
subscribeSame as syncSame as sync
webhookNo, returns 202Wait for the signed callback and keep polling as a backup

Will many video jobs be refused at the door?

No. The admission docs state that concurrency is a dispatch limit, not a submit limit: with the workspace at its concurrency limit, Sume can still accept jobs as queued while queue capacity remains, and workers move them to processing later. When the queue is full, new paid submissions fail with 429 queue_full. So the equivalent of a big batch is a loop of submits, each keeping its own job id and its own Idempotency-Key; see idempotency keys for AI video APIs.

Which should I pick?

If your work is text or embeddings that can wait hours, a batch API fits that shape. If each result is a video file, submit per job in async or webhook mode and track job ids. A sync wait is capped at 30 seconds, so treat it as a convenience for short jobs, not a way to hold a connection open. For a side-by-side of the two modes, see sync vs async video generation API.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume