Fish Audio Drama 3 single-word fix vs Sume sentence segments

Fish Audio says Drama 3 preview can fix a single word. Sume has no word-level repair: regenerate a sentence and join takes with Timeline audio concat.

4 min readSume
All posts

Fish Audio's post says its Drama 3 preview can "fix only a single word", and its API docs call drama-3-preview a preview model whose behavior and availability may change. Sume has no word-level repair. The nearest workflow is to regenerate one sentence as a new take and splice it in with Timeline audio concat, which joins without re-TTS.

Fish Audio facts are from its API docs and a Fish Audio post; Sume facts from the Timeline audio docs and the OpenAPI schema, read 2026-10-01.

What does Fish Audio claim for Drama 3?

The post says you can describe tone, pacing and character in simple language, shift voice mid-sentence, generate a multi-character scene, or fix a single word. The API docs list drama-3-preview among allowed model values and warn that unrecognized values fall back to s2.1-pro. These are vendor claims; this post did not test them.

What can Sume do about one wrong word?

Sume TTS can return gapless sentence segments[], where each segment ends exactly where the next starts. That gives you sentence-sized units to replace. Regenerate only the sentence containing the error, then join old and new pieces.

Single-word fix versus Sume's sentence-level route, read 2026-10-01.
StepFish Audio Drama 3 (claimed)Sume
Unit you redoOne wordOne sentence
Split the takeNot describedSentence segments[], gapless
Join the piecesNot describedTimeline audio concat, 1 to 20 ordered parts
SeamsNot describedSample-domain join: no re-TTS, no added silence

How do I splice a new take in?

Import the audio so it lives on media.sume.com, then call POST /v1/timeline-1.0/audio with operation: "concat" and ordered parts[]. Each part takes a url and optional source_in and duration, so you can trim the old take around the bad sentence. Produced audio is capped at 1800 seconds.

const body = {
  operation: "concat",
  parts: [
    { url: "https://media.sume.com/artifacts/example/take1.wav", duration: 12.4 },
    { url: "https://media.sume.com/artifacts/example/fixed-sentence.wav" },
    { url: "https://media.sume.com/artifacts/example/take1.wav", source_in: 15.1 },
  ],
};
console.log(JSON.stringify(body));

Will the fixed sentence match the voice and mood?

Not guaranteed. A new take can differ in delivery, and emotion is only an optional guide. Listen at the seams, and keep speed and volume identical across takes. For longer scripts, see the TTS length limit and chapter splitting.

Sources

Related posts

More in Models

All Models posts

Written by Sume