Can ChatGPT make podcasts? Script, voices and one audio file

ChatGPT can write a podcast script. For a finished episode, give it speech tools over MCP: it voices each host's lines and joins them into one file.

5 min readSume
All posts

ChatGPT can write a podcast script, but a finished episode is an audio file with each host's voice in it, and for that ChatGPT needs speech tools. In developer mode it can call tools on a remote MCP server: one text-to-speech call per line in each host's voice, then a join that puts the lines in order into one file, and it hands you a link to the episode.

ChatGPT's side comes from OpenAI's ChatGPT Developer mode guide. OpenAI's help center could not be read for this post, so it says nothing about ChatGPT's built-in voice features. The tools used here are tts_create and timeline_audio on Sume's hosted MCP server, from MCP overview, Timeline audio and the Sume API reference, all read on 2026-09-29. Sume has no official ChatGPT connector, and its basics page says hosted MCP still works but is not the primary path today.

How do I make a podcast with ChatGPT?

OpenAI lists developer mode for Pro, Plus, Business, Enterprise, and Education accounts on the web. How to add an MCP server to ChatGPT walks through turning it on and adding https://mcp.sume.com/mcp as an app. At Sume's sign-in, Write is off by default and paid tools stay hidden until it's on, so turn it on. Then, in a chat with the app selected, name the tools, as OpenAI advises ("Be explicit"):

Write a 3-minute two-host episode about our spring launch
from the notes below, as lines tagged HOST_A or HOST_B.
Then use the Sume app only:
1. tts_create for each line: HOST_A with avatar_handle
   "host_a", HOST_B with avatar_handle "host_b", wav output.
   Run the first call with dry_run and show me the cost.
2. timeline_audio with operation "concat", parts in script order.
3. jobs_wait, then jobs_result, and give me the audio_url.

How does each host get a different voice?

Each TTS request takes one voice selector: a ready avatar's avatar_handle or avatar_id, or a voice.id. So the script is voiced line by line, each with its host's voice. Set language on every line that isn't English; left out, it defaults to English. In the Sume app today, the Voices screen lets you clone a voice or invent one from a prompt; Sume's MCP docs list no tool for that step, so do it in the app before the chat. Text to speech with multiple voices covers the voice rules.

What are the limits?

The join is a sample-domain concat: no re-synthesis and no silence at the seams. It joins end to end, so a music bed under the voices is a different job, a Timeline 1.0 render with a soundtrack; AI podcast generator from text covers that and the full API recipe.

From the TTS schema in the Sume API reference, Timeline audio, MCP tools and gates and Jobs and results, read 2026-09-29.
StepToolLimit
Voice each linetts_createUp to 20,000 characters per request; spaces and punctuation count
Join the linestimeline_audio (concat)1–20 ordered parts of this workspace's media.sume.com audio; output up to 1,800 seconds; WAV by default
Wait for a jobjobs_waitHolds at most 55 seconds per call; retry with the same ids, never resubmit
Any paid calltts_create, timeline_audioidempotency_key required; dry_run=true previews cost without submitting

Can Claude make podcasts?

The same way. Anthropic says all current Claude models support text and image input and text output, so Claude writes the script, and a speech tool makes the audio. Custom connectors using remote MCP are available on Free, Pro, Max, Team, and Enterprise plans, with Free users limited to one custom connector. Add Sume to Claude as a custom connector shows the setup, and MCP OAuth and API keys covers the Write toggle.

How much does a podcast episode cost?

Speech is billed on transcript characters at $0.0475 per 1,000 characters, and each concat reserves $0.01 per job, each plus a 5.5% agent fee by default, charged to your Sume wallet. A longer episode needs more lines, and past 20 lines the join runs in stages, since a concat's output can be a part of the next one. On ChatGPT's side, write actions require confirmation by default, so check each tool input before you approve it.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume