Kling 3 lip sync with your own audio: what each API does

Kling lists lip sync driven by text or audio. Sume's kling-3 takes no audio input, only a generate_audio switch; models with audio references are listed.

5 min readSume
All posts

Kling lists Lip Sync as a capability of all its model versions, combined with text or audio to drive a character's mouth, but Sume's kling-3 takes no audio file: it accepts no input_references, only a generate_audio on/off switch for sound the model makes itself. To drive a video from your own audio on Sume, use a model that lists audio_url references, or a tool built for talking video.

Kling's capability is from its capability map; Sume's limits from Video generation and code, read 2026-09-29.

What does Kling say about lip sync and voice?

The map's Global Capabilities table lists Lip Sync for all model versions, "combined with text or audio to drive the mouth shape of characters in the video", and Text to Audio and Video to Audio for all versions. For Kling 3.0 text-to-video it marks Native Audio supported and Voice Control (human voice) not supported.

What can I send on Sume?

The docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. "Honored" means the reference is used as guidance; the docs do not promise frame-accurate mouth sync, so test a short clip.

Audio inputs by Sume model from the video docs and catalog code, read 2026-09-29.
ModelAudio behavior
kling-3generate_audio on or off; no audio input
seedance-2.5, seedance-2, seedance-2-fast, seedance-2-miniAudio references honored; optional generate_audio
wan-3.0Audio references honored
minimax-h3, minimax-h3-maxAudio references honored
gemini-omni-flash-1.1Native audio; no audio references
grok-imagine-video-1.5No audio

How do I send an audio reference?

Add an audio_url entry to input_references, alongside an image reference.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-ref-001" \
  -d '{
    "model": "seedance-2.5",
    "prompt": "The woman in the photo speaks to camera, warm light.",
    "input_references": [
      {"type":"image_url","image_url":{"url":"https://example.com/face.png"}},
      {"type":"audio_url","audio_url":{"url":"https://example.com/voice.mp3"}}
    ],
    "duration": 8
  }'

Is there a better tool for talking video?

For a face that must speak your exact audio, the tools built for it are covered in talking video: avatar vs lip sync vs motion control and lip sync from a photo and audio.

Sources

Related posts

More in Models

All Models posts

Written by Sume