C# text to speech: call a TTS API with HttpClient
Text to speech in C#: POST the text and a voice with HttpClient, poll the job until it finishes, then stream the MP3 from its audio_url to a file.

To do text to speech in C#, send the text and a voice to a TTS API with HttpClient, poll the job until it finishes, then download the audio file the result links to. With Sume TTS 1.0 that is POST https://api.sume.com/v1/tts-1.0/generate, then GET status_url until terminal is true, then GET result_url for the MP3's audio_url. Any .NET server app can run it, because it's plain HTTPS and JSON.
Sume's SDK is a TypeScript client, so from C# you call the HTTP API directly. The Sume facts come from the TTS 1.0 schema in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results; the .NET calls follow Microsoft's HttpClient and HTTP requests pages. All were read on 2026-09-29. The same flow in Java is Text to speech API in Java.
What goes in the request?
The text, plus a way to pick the voice: an avatar reference or voice.id. If you send both, they must match, or the request fails with 400. Everything else is optional.
| Body field | What it does |
|---|---|
transcript | The text, up to 20,000 characters; spaces and punctuation count toward usage |
avatar_handle or avatar_id | Uses the voice of an avatar from GET /v1/avatar-1.0/avatars whose voice.status is ready |
voice.id | A TTS voice UUID or a Voices library id (voi_ plus 32 hex) you already hold |
language | Set it for every non-English transcript; omitted means English |
output_format | Defaults to MP3 at 44,100 Hz and 128 kbps; wav and raw are the other containers |
How do I send the request from C#?
Create one HttpClient and reuse it: Microsoft's docs say it is “intended to be instantiated once and reused throughout the life of an application,” because a new client per request can exhaust sockets under load. DefaultRequestHeaders carries the key on every call; read it from a server-side environment variable, never from client code. JsonContent.Create serializes the dictionary as the JSON body, and the Idempotency-Key header makes a retried submit return the original job instead of billing a second one.
using System.Net.Http.Headers;
using System.Net.Http.Json;
using System.Text.Json;
var api = new HttpClient(); // create once, reuse
api.DefaultRequestHeaders.Authorization = new AuthenticationHeaderValue(
"Bearer", Environment.GetEnvironmentVariable("SUME_API_KEY"));
async Task<JsonElement> Data(HttpResponseMessage res)
{
res.EnsureSuccessStatusCode(); // throws outside 200-299
return (await res.Content.ReadFromJsonAsync<JsonElement>()).GetProperty("data");
}
var submit = new HttpRequestMessage(HttpMethod.Post, "https://api.sume.com/v1/tts-1.0/generate")
{
Content = JsonContent.Create(new Dictionary<string, object>
{
["transcript"] = "Your order has shipped and arrives on Thursday.",
["avatar_handle"] = "acme",
}),
};
submit.Headers.Add("Idempotency-Key", "order-shipped-001");
var job = await Data(await api.SendAsync(submit));How do I wait for the audio and save it?
Poll status_url at least once, waiting next_poll_after_seconds between reads, until terminal is true; the submit envelope has no sume_status, and a replayed submit can return a job that has already finished. Read result_url only when sume_status is completed: /result answers 409 job_not_completed until result_ready is true. In current code audio_url is the file's public Sume media URL, so the download uses a second client without the key, and GetStreamAsync returns the body as a stream you copy to disk.
var status = job; // poll at least once: a replayed submit may already be done
do
{
var hint = status.GetProperty("next_poll_after_seconds");
await Task.Delay(TimeSpan.FromSeconds(hint.ValueKind == JsonValueKind.Number ? hint.GetInt32() : 2));
status = await Data(await api.GetAsync(job.GetProperty("status_url").GetString()));
} while (!status.GetProperty("terminal").GetBoolean());
var state = status.GetProperty("sume_status").GetString();
if (state != "completed") throw new Exception($"TTS job ended as {state}");
var result = (await Data(await api.GetAsync(job.GetProperty("result_url").GetString())))
.GetProperty("result");
var media = new HttpClient(); // no API key on media downloads
await using var audio = await media.GetStreamAsync(result.GetProperty("audio_url").GetString());
await using var file = File.Create("speech.mp3");
await audio.CopyToAsync(file);Which voices can I use from C#?
Two kinds. Any avatar in your workspace whose voice.status is ready lends its voice through avatar_handle or avatar_id; GET /v1/avatar-1.0/avatars lists them. A voice you cloned or designed in the Sume app has a voi_ id for voice.id; creating that voice isn't an API call, as Voice cloning API explains. A voice.id of any other shape fails with 400 invalid_voice_id before a job is queued.
What does it cost, and what are the limits?
TTS 1.0 costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default. Synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and captures no credits. The job is async and non-streaming: you get a finished file, not audio chunks. A webhook_url replaces the poll loop; C# webhook receiver verifies the signed callback in ASP.NET Core.
Sources
Related posts
More in Developers
- Text to speech streaming API: what Sume returns instead
Sume's text to speech API doesn't stream audio chunks. It returns a finished file per job; split long scripts into sentence jobs to start playback sooner.
- Duck background music under a voiceover with the Timeline API
Set soundtrack.duck_db (0 to 20) on POST /v1/timeline-1.0/render so the music dips under your voiceover spine. It needs a real spine; silence mode is refused.
- Render a silent video from clips with the Timeline API
Set audio.mode to silence and a duration_seconds on POST /v1/timeline-1.0/render to join clips with no audio file. Which fields are illegal there, and pricing.
- Timeline error too_many_chained_transitions: how to fix it
Timeline 1.0 refuses more than 8 adjacent fades with too_many_chained_transitions. Insert a hard cut. Also transition_too_long, transition_not_frame_aligned.
Written by Sume