GPT-6 Astra tool calling needs the Responses API: meaning for Sume
OpenAI says GPT-6 Astra supports Chat Completions but its tool calling requires Responses. Use the Responses mcp tool for Sume, not a chat.completions loop.

If you want GPT-6 Astra to call tools, use the Responses API. OpenAI's latest-model guide says Astra supports Chat Completions, but its tool calling requires Responses, so a chat.completions loop with tools is the wrong shape. For Sume, that means attaching the hosted MCP server as an mcp tool on a Responses request.
Both statements about Astra come from OpenAI's pages, read 2026-09-29: the guide for the Responses requirement and the model page for the model id gpt-6-astra and its endpoint list (Responses, Chat Completions, Batch).
Which endpoint does what for Astra?
The model page lists three endpoints. Only one of them is the tool-calling path.
| Endpoint | Listed for Astra | Tool calling |
|---|---|---|
| Responses | Yes | Supported |
| Chat Completions | Yes | Not supported for tool calling; use Responses |
| Batch | Yes | See OpenAI's guide |
How does Sume connect to a Responses request?
Sume's hosted MCP server is at https://mcp.sume.com/mcp and accepts an API key as a Bearer token or x-api-key. Add it as a remote MCP tool in your Responses request, with the server URL, and let the model list and call tools.
Existing pages on this site show the full request; the point here is that the server does not care which model asks. Paid tools still need an idempotency_key, and dry_run=true previews the cost.
Is Sume's Agent Completions API a Chat Completions drop-in?
No. POST /v1/agent/completions borrows the messages[] shape, but it returns 202 with an agent.run receipt to poll, not a choices[] response. Streaming and a synchronous OpenAI-compatible response are listed as not available yet.
It also has a required generation_spend_cap_usd with no default, accepts system and user turns only, and rejects assistant turns.
What should I do with an existing chat.completions agent?
Move the tool loop to Responses, keep any plain text Chat Completions calls as they are, and connect Sume through the mcp tool. If you would rather hand the whole task to Sume's own agent, call Agent Completions from your backend and poll the run.
Sources
Related posts
More in Developers
- GPT Image 2.5 4K: how to request a 3840x2160 image by API
To get a 4K image from GPT Image 2.5 on Sume, send image_size 3840x2160 to POST /v1/images and be ready for a 202 job response. Request, cost, and polling.
- GPT Image 2.5 image editing API: edit a photo with a prompt
Edit a photo with GPT Image 2.5 on Sume: send the image in input_references, describe the change, and set aspect_ratio to auto. Up to 16 references per call.
- GPT Image 2.5 request returned 202: how to get the image from the job
When GPT Image 2.5 takes longer than 30 seconds on Sume, POST /v1/images returns 202 with a job. Poll the status URL, then read the result, or use a webhook.
- GPT Image mask edit API: how to use mask_url on GPT Image 2.5
GPT Image 2.5 on Sume takes an optional mask_url for edits. OpenAI says the mask needs an alpha channel and only guides the edit. The request and the limits.
Written by Sume