TikTok AI voiceover label: generic TTS vs a real person's voice
TikTok's 2026-H2 guidelines: generic text-to-speech narration needs no label, but AI audio that mimics a real person's voice does. How that maps to Sume voices.

TikTok's Community Guidelines (version 2026-H2) say disclosure is not needed when you use generic text-to-speech narration, as long as the TTS is not a recognizable voice of a known individual. Disclosure is needed when AI-generated audio mimics the voice of a real person. A stock voice id used for narration is the first case; a voice made to sound like a specific person is the second.
Read from TikTok's page on 2026-09-30; the guideline text is TikTok's, and its enforcement is its own.
Where does TikTok draw the line?
The page defines likeness as a recognizable image, video or audio representation of a person, including their voice. It asks creators to label AI-generated or significantly edited content that shows realistic-looking scenes or people, and lists the voice case among the situations where content that is not harmful could still be confusing.
| Voice case | Disclosure per the page |
|---|---|
| Generic TTS narration, not a recognizable voice of a known individual | Not needed |
| AI-generated audio that mimics a real person's voice | Needed |
Which one is a Sume voice?
Sume's TTS takes a voice id that must be "a TTS voice UUID (8-4-4-4-12 hex) or a Voices library id", not a voice name from another TTS ecosystem; other shapes are rejected with 400 before any credits are reserved. Picking a library voice for narration is a choice of a listed voice. Whether a particular voice sounds like a known person is still a question you answer against TikTok's wording, so do not pick a voice to imitate someone.
Text input is capped: spaces and punctuation count toward usage, with a maximum of 20000 characters per request.
What about talking-face shots?
Sume's models overview says every on-camera speaking shot is Fabric with an accepted still plus TTS audio. The voice side follows the table above, and the face is a separate question: realistic-looking generated people fall under TikTok's label request for realistic scenes or people, and the voice exemption does not cover the picture.
Does joining a voiceover into a video change the answer?
Timeline 1.0 takes "one audio spine + ordered" video slots and returns one MP4, so the narration becomes the audio track of the render. That does not change what the voice is. For the comparable YouTube question, see does YouTube require a label for AI voiceover Shorts.
Sources
Related posts
More in Use cases
- TikTok API 10-minute video limit: trim a longer video with Sume
TikTok's media guide says a developer can send at most 10 minutes through initialize upload. Sume video-trim cuts up to 900 s from a 1800 s source into an MP4.
- TikTok API accepts MP4, WebM or MOV: which format Sume returns
TikTok's media guide lists MP4 (recommended), WebM and MOV. Sume video trim and timeline write an MP4, so a Sume result already matches the recommended type.
- TikTok upload chunk size: 5-64 MB and 1000 chunks for AI video
TikTok's media guide sets 5 MB to 64 MB chunks, 1 to 1000 chunks and a 4GB maximum. Sume gives no file size control, so measure the MP4 before you split it.
- Do you have to label a face swap video on TikTok?
TikTok's disclosure list includes a face replaced with someone else's. Sume Avatar Face Swap output is covered: Beta endpoint inputs and disclosure.
Written by Sume