AI voice generator for games: voice every line as a file
An AI voice generator for games turns each dialogue line into an audio file in a character's voice. Voices, engine formats, batch runs and cost.

An AI voice generator for games turns each line of your dialogue script into an audio file in a character's voice, so NPC barks, quest lines and cutscene dialogue can be voiced in a batch and re-voiced when the script changes. It is an offline asset step: you generate the files ahead of time and ship them with the game.
On Sume each line is one text to speech request to POST /v1/tts-1.0/generate, billed at $0.0475 per 1,000 characters plus a 5.5% agent fee by default. The fields come from the TTS schema in the Sume API reference (see API reference), and batch behavior from Jobs and results and Generation admission, read on 2026-09-29.
How do I give each character its own voice?
Map every character to one voice and send it on each of that character's lines. A request selects the voice with voice.id (a TTS voice UUID or a Voices library id, voi_ plus 32 hex) or with a ready avatar's avatar_id / avatar_handle. New character voices are made in the Sume app under Assets → Voices, where you can clone from a recording or describe a person and have a voice generated; there is no API route for that step. AI voice generator with a prompt covers describing one.
If a voice is cloned from a real actor, the Sume Terms of Service say you represent that you have permission to use any person's voice you submit. That is what the terms say, not legal advice.
What audio format can I get for my game engine?
Set output_format to what your engine imports. The default is MP3 at 44,100 Hz and 128 kbps; for uncompressed game assets, ask for WAV:
| Field | Values |
|---|---|
container | mp3, wav, raw |
sample_rate (Hz) | 8000, 16000, 22050, 24000, 44100, 48000 |
encoding (wav/raw) | pcm_f32le, pcm_s16le, pcm_mulaw, pcm_alaw |
bit_rate (mp3) | 32000, 64000, 96000, 128000, 192000 |
How do I voice hundreds of lines without duplicates?
Loop over your dialogue table and send one request per line. Build the Idempotency-Key from the line id and its script revision: a retry after a timeout then returns the original job instead of billing a second one. Reuse a key only with the same body; a changed line gets a new key. metadata stores your line id on the job.
Jobs past your workspace's concurrency limit are accepted as queued while queue capacity remains; when the queue is full, submits fail with 429 queue_full. In current code a same-key retry replays that refusal, so after the queue drains, resubmit that line with a new key. Use mode: "webhook" with a webhook_url, or poll the job's status, then download each file and name it by line id.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: blacksmith-greet-012-r3" \
-d '{
"transcript": "Back again? That blade will not sharpen itself.",
"voice": { "mode": "id", "id": "9f3e2b1c-6a4d-4c8e-9f1a-2b3c4d5e6f70" },
"generation_config": { "speed": 0.95, "emotion": "gruff" },
"output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 48000 },
"metadata": { "line_id": "blacksmith-greet-012" },
"mode": "async"
}'Can I direct the performance of each line?
Partly. generation_config takes speed (0.6 to 1.5), volume (0.5 to 2.0) and emotion, a free-form guide of up to 64 characters. The schema documents no list of emotion values, so treat it as a hint and listen to the result. Set language for any non-English line; omitted, it defaults to English. For a scene with several characters joined into one file, see Text to speech with multiple voices.
How much does AI game voice acting cost, and what can't it do?
Speech is billed per character at $0.0475 per 1,000 characters; spaces and punctuation count, and each retake is billed again. The line counts and lengths below are example assumptions.
- No live voice: TTS 1.0 is an async job with polling or a webhook, not streaming, so it can't voice player-driven dialogue at runtime.
- One request takes up to 20,000 characters, and audio longer than 1,200 seconds fails with
tts_duration_exceeded.
| Example | Lines | Characters | Price |
|---|---|---|---|
| Short game: 3 characters | 200 | 12,000 | $0.57 |
| Indie RPG: 20 characters | 2,000 | 160,000 | $7.60 |
| Larger game: 60 characters | 10,000 | 900,000 | $42.75 |
Sources
Related posts
More in Use cases
- AI voiceover for ads: one brand voice across every cut
Make AI voiceover for ads by pinning one voice and voicing each 6, 15 or 30 second cut as its own text to speech request. Fields, fit and cost.
- AI voiceover for short videos: 9:16 with voice and music
Make an AI voiceover for a Reel, Short or TikTok: script to speech, clips on the beat of the voice, optional music bed, one vertical MP4.
- AI wedding invitation video maker: art, motion, music, text
An AI wedding invitation video is your invitation art animated into short clips, set to music, with names, date and venue burned on as typed text.
- AI whiteboard animation generator: blank board to sketch
AI can imitate whiteboard animation: a blank-board first frame, a finished-sketch last frame, the drawing in between. Words go on as captions.
Written by Sume