OpenAI async tool calling for long-running render jobs
OpenAI's async tool calling lets the model keep working while your tool runs. For a slow Sume render, return the job id fast, then wait in slices.

Async tool calling is a Responses API control: the model keeps working while your application runs a function or custom tool, and you return the result when it is ready. A video render fits that shape if your tool does two small things: submit the job and hand back its id, then check the job later instead of holding one call open for the whole render.
The OpenAI wording is from its changelog (Sep 3 entry, read 2026-10-01); the Sume behavior is from Jobs and results and MCP tools and gates.
What does the changelog say async tool calling does?
The Sep 3 entry lists it among new controls for long-running work with GPT-6 Astra in the Responses API: "Let the model continue working while your application runs function or custom tools, then return results as they become available." It sits next to mid-turn steering over WebSockets; that is a separate control, covered in mid-turn steering and a started Sume job.
How should a render tool behave?
Sume generation endpoints create durable jobs, and the docs tell you to store the job id from the submit response so work can be recovered after a process restart. Every submit mode returns that id in its first response, and a 2xx means the job exists, not that it finished. So the tool you expose to the model can return the id and status_url immediately.
| Step | Tool result | Docs rule |
|---|---|---|
| Submit | Job id plus status URL | async is the default mode and returns 202 with the envelope |
| Check | terminal and result_ready flags | Poll status_url until terminal is true |
| Not done yet | Same job id again | "Poll. Do not resubmit." |
| Done | Artifacts from the result route | GET result_url once result_ready is true |
How do I wait without one huge call?
Over hosted MCP, jobs_wait is bounded: timeout_seconds defaults to 50 and is capped at 55, and the docs say to wait for a ten-minute render by repeating the wait, not by asking for a longer one. On wait_slice_expired, call jobs_wait again with the same ids. Each slice can be its own tool result, so the model is never blocked on a single long call. See MCP jobs_wait for long video jobs.
What happens if my side times out?
Nothing happens to the job. The docs state that a client-side timeout does not cancel it; the job keeps running and still bills, and you have only stopped watching. Pick it back up from status_url, or cancel it explicitly. Never resubmit the create because your tool call expired.
Sources
Related posts
More in Developers
- OpenAI image-encoding fix: rerun workflows with the right key
OpenAI fixed an image-encoding bug and advised retrying affected workflows. On Sume, a replayed idempotency key returns the original; use a new key to rerun.
- mTLS instead of an API key? How Sume authenticates calls
OpenAI's docs list mutual TLS and workload identity federation. Sume authenticates API calls with one API key header and signs webhooks with HMAC.
- OpenAI service-account-only keys, and Sume Format run limits
OpenAI admins can allow only service-account keys. On Sume, service-account keys cannot create Format runs or bulk queues; use a user-issued key.
- Punch-in zoom on video by API: crop or zoompan, no keyframes
Sume has no auto zoom switch. Use a crop op for a fixed punch-in or the allowlisted zoompan filter in a video-filter graph; there are no keyframes.
Written by Sume