Can ChatGPT make music? Getting a song or beat as audio
To get a song or a beat as an audio file, give ChatGPT a music tool: in developer mode it can call one over MCP and hand back a link to the track.

To get music you can play, a song or a beat as an audio file, ChatGPT needs a music-generation tool. In developer mode it can call one on a remote MCP server: ChatGPT writes the music brief, the tool renders the track, and ChatGPT hands you a link to the file.
ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. The tool used here is music_create on Sume's hosted MCP server, described in Music 1.0, Music Router, and MCP tools and gates. All were read on 2026-09-29. Sume has no official ChatGPT connector: this is a remote MCP connection, and Sume's basics page says hosted MCP still works but is not the primary path today.
How do I give ChatGPT a music tool?
OpenAI says developer mode provides "full Model Context Protocol (MCP) client support for all tools, both read and write", and lists it as available to Pro, Plus, Business, Enterprise, and Education accounts on the web. Add https://mcp.sume.com/mcp as a developer-mode app; How to add an MCP server to ChatGPT walks through the screens. At Sume's sign-in, Write is off by default, and without mcp:write paid tools such as music_create are hidden, so turn it on. In a chat, choose Developer mode from the Plus menu, select the app, and ask for the track by tool name.
What should I ask ChatGPT for?
A specific brief, not a genre word. Sume's music docs suggest seven axes: emotion, genre, tempo as a number, key and mode, two to four instruments with texture, an arc with one named moment, and era or production. Put exclusions in the prompt itself, and steer length there too, because there is no duration setting. For example:
Use the Sume app's music_create tool. Run it with dry_run
first and show me the cost, then submit.
Prompt: Cocky, restless drill beat, 142 BPM half-time, F minor.
808 with long glide, crisp hi-hats, dark piano stabs. Drop to
bass and claps at 0:20, full return at 0:28. About 1 minute.
Instrumental, no vocals.Can it make a song with vocals, or a set length?
Not as settings. The docs describe the engine behind the tool, sume/music-auto (Lyria 3.5 today), as producing full-length structured songs up to a few minutes, steered by the prompt, and call every brief axis a creative direction, not a guaranteed output setting. Listen to the result before you use it. The full call is in Lyria MCP; here is what it means for a song or a beat:
| You want | What the docs say |
|---|---|
| Vocals or lyrics | No lyrics or vocals field is documented; the prompt guide closes with "Instrumental, no vocals." |
| A set length | No duration field; it is rejected. Ask for length in the prompt |
| A tempo or key | Write them in the prompt, such as "142 BPM half-time" |
| Leave something out | Say so in the prompt; a non-empty negative_prompt is rejected |
| The same track twice | No seed, temperature, or guidance parameter |
| A file | An audio artifact on media.sume.com, typically audio/mpeg |
How much does it cost, and will ChatGPT ask first?
A music generation is listed at $0.125 per audio on API pricing, plus a 5.5% agent fee by default, charged to your Sume wallet. The docs say the price does not vary by prompt length or image conditioning. Sending dry_run=true previews the cost without submitting the job, max_spend_usd caps a call when you send it, and every paid call needs an idempotency_key.
On ChatGPT's side, OpenAI says write actions require confirmation by default and that tools without the readOnlyHint annotation are treated as write actions. In current code Sume marks music_create readOnlyHint: false, so ChatGPT asks before the dry run and before the real submit, unless you choose to remember your approval for that tool in the conversation. Check the prompt text in the tool input each time.
Where is the track when it's done?
In the job result. music_create answers with a job id, not audio. ChatGPT then waits with jobs_wait, which holds at most 55 seconds per call, and reads jobs_result. The tool's current description says a track typically takes under a minute. The finished file is a link on media.sume.com that you can open or download outside the chat. To drop the track under a video, see Lyria MCP for the same tool from Claude, or the music generation API for a backend.
Sources
Related posts
More in Agents
- Can ChatGPT make music videos? Song, clips, one render
With tools, yes: ChatGPT can have a track generated, a short clip made per section, and the clips cut over the song in one MP4 render.
- Can ChatGPT make podcasts? Script, voices and one audio file
ChatGPT can write a podcast script. For a finished episode, give it speech tools over MCP: it voices each host's lines and joins them into one file.
- Can ChatGPT make product videos from a product photo?
Yes, with a video tool: in developer mode ChatGPT can send your product photo to a video model as the first frame and return a link to the clip.
- Can ChatGPT make slideshow videos? Stills, voice, one MP4
Yes, through tools: ChatGPT can make the slides, voice the script, and have a timeline tool hold each still for its line in one MP4.
Written by Sume