OpenAI Agents SDK imageGenerationTool action edit and Sume
imageGenerationTool() takes action: 'generate' | 'edit' | 'auto' in openai-agents-js 0.18.0. To edit with Sume instead, wrap its image route as a function tool.

In openai-agents-js v0.18.0, imageGenerationTool() accepts action: 'generate' | 'edit' | 'auto', and leaving it unset keeps the provider default. That option belongs to the hosted tool. If you want the edit done by Sume, give the agent a function tool that calls Sume's image route with prompt and image_urls.
Vendor facts are from the openai-agents-js release notes; Sume facts are from Image generation and edit, read 2026-09-30.
What changed in the Agents SDK?
The v0.18.0 notes (2026-09-10) say the option is forwarded through streaming and non-streaming Responses requests. The Python release v0.22.2 (2026-09-09) is listed as "support current image generation tool options". The notes say nothing more about how the hosted tool receives the images to edit, so check the SDK reference before relying on a particular input shape.
What does Sume's image edit take?
Sume uses the same image route for generation and edit. The shape of the request decides which one you get.
| Goal | Request |
|---|---|
| Text to image | prompt only |
| Edit or reference | prompt + image_urls (1–10 public HTTPS URLs) |
| Masked edit | add mask_image_url with image_urls |
How do I expose the edit as a function tool?
Wrap the route in a function tool. The docs example sends an Idempotency-Key on the create call and says to reuse it only for the same operation and payload when retrying, so let the caller pass a stable key.
import { tool } from "@openai/agents";
import { z } from "zod";
export const editImage = tool({
name: "edit_image",
description: "Edit images at public HTTPS URLs with Sume.",
parameters: z.object({
prompt: z.string(),
image_urls: z.array(z.string().url()).min(1).max(10),
key: z.string(),
}),
async execute({ prompt, image_urls, key }) {
const res = await fetch("https://api.sume.com/v1/image-1.0/generate", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify({ prompt, image_urls }),
});
return await res.text();
},
});What comes back, and what should I store?
The call returns a job, not a finished image. Poll the job and read its result, as in Jobs and results. Sume mirrors outputs into Sume-owned media URLs, and the docs say integrations should store the Sume URL, not raw provider URLs. Sources must be public HTTPS URLs; localhost, private-network and non-HTTPS URLs are rejected before submission.
Edits cost money, so consider an approval step; see human-in-the-loop for paid tools. For video from the same SDK, see MCP video generation.
Sources
Related posts
More in Developers
- OpenAI Agents SDK MCP require_approval for Sume write tools
Use require_approval with a tool_names list to gate Sume write tools by name, or connect with OAuth mcp:read so those tools are never visible.
- OpenAI Agents SDK MCP tool_input_guardrails for Sume spend
tool_input_guardrails on an Agents SDK MCP server can reject a call before it runs. For Sume paid tools, check idempotency_key and max_spend_usd.
- OpenAI Batch API limits vs Sume bulk runs: 50,000 vs 100
OpenAI Batch allows 50,000 requests per file with a 24h window. A Sume Format bulk run takes 1-100 items at concurrency 1-16, so chunk your file accordingly.
- OpenAI-compatible /v1/audio/speech: what Sume's TTS route is
Sume has no `/v1/audio/speech` clone. Its TTS is `POST /v1/tts-1.0/generate` or the Router, a job-based request authenticated with a Sume key.
Written by Sume