TikTok AI voiceover label: generic TTS vs a real person's voice

TikTok's 2026-H2 guidelines: generic text-to-speech narration needs no label, but AI audio that mimics a real person's voice does. How that maps to Sume voices.

4 min readSume
All posts

TikTok's Community Guidelines (version 2026-H2) say disclosure is not needed when you use generic text-to-speech narration, as long as the TTS is not a recognizable voice of a known individual. Disclosure is needed when AI-generated audio mimics the voice of a real person. A stock voice id used for narration is the first case; a voice made to sound like a specific person is the second.

Read from TikTok's page on 2026-09-30; the guideline text is TikTok's, and its enforcement is its own.

Where does TikTok draw the line?

The page defines likeness as a recognizable image, video or audio representation of a person, including their voice. It asks creators to label AI-generated or significantly edited content that shows realistic-looking scenes or people, and lists the voice case among the situations where content that is not harmful could still be confusing.

Voice cases in TikTok's 2026-H2 guidelines, read 2026-09-30.
Voice caseDisclosure per the page
Generic TTS narration, not a recognizable voice of a known individualNot needed
AI-generated audio that mimics a real person's voiceNeeded

Which one is a Sume voice?

Sume's TTS takes a voice id that must be "a TTS voice UUID (8-4-4-4-12 hex) or a Voices library id", not a voice name from another TTS ecosystem; other shapes are rejected with 400 before any credits are reserved. Picking a library voice for narration is a choice of a listed voice. Whether a particular voice sounds like a known person is still a question you answer against TikTok's wording, so do not pick a voice to imitate someone.

Text input is capped: spaces and punctuation count toward usage, with a maximum of 20000 characters per request.

What about talking-face shots?

Sume's models overview says every on-camera speaking shot is Fabric with an accepted still plus TTS audio. The voice side follows the table above, and the face is a separate question: realistic-looking generated people fall under TikTok's label request for realistic scenes or people, and the voice exemption does not cover the picture.

Does joining a voiceover into a video change the answer?

Timeline 1.0 takes "one audio spine + ordered" video slots and returns one MP4, so the narration becomes the audio track of the render. That does not change what the voice is. For the comparable YouTube question, see does YouTube require a label for AI voiceover Shorts.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume