HeyGen lipsync API: speed vs precision mode

HeyGen's POST /v3/lipsyncs re-syncs an existing video to new audio in speed or precision mode. Body fields, statuses, and how Sume's still-plus-audio differs.

5 min readSume
All posts

HeyGen's lipsync API is one endpoint, POST /v3/lipsyncs, that takes an existing video plus new audio and redraws the mouth to match; mode chooses between speed and precision. HeyGen describes speed as a fast audio-only resync for drafts and precision as frame-accurate mouth movement for long-form, higher-fidelity work.

HeyGen's fields are from its Speed and Precision pages, read 2026-09-29. The Sume comparison is from the Models overview.

What is the difference between speed and precision?

HeyGen says both modes run on one lip sync engine and trade latency for fidelity. The request is identical apart from the mode value.

Modes, from HeyGen's Speed and Precision pages and Sume's Models docs, read 2026-09-29.
speedprecision
HeyGen's descriptionFast audio-only resync; rapid drafts and batch processingFrame-accurate mouth movements on long-form video
Suited toPreviewing edits before a final renderCinematic content where fidelity matters
EndpointPOST /v3/lipsyncsPOST /v3/lipsyncs
Bodyvideo, audio, modevideo, audio, mode

What does a request look like?

video and audio are objects, each either a url or an asset_id. The response returns a lipsync_id.

curl -X POST "https://api.heygen.com/v3/lipsyncs" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "video": { "type": "url", "url": "https://example.com/source.mp4" },
    "audio": { "type": "url", "url": "https://example.com/new-audio.mp3" },
    "mode": "precision"
  }'

How do I get the result back?

The job is asynchronous. Poll GET /v3/lipsyncs/{lipsync_id} until status is completed or failed (it moves from pending to running first), or pass a callback_url to get a POST when it finishes. Optional fields include enable_caption, enable_speech_enhancement, start_time and end_time for a partial lipsync, and enable_dynamic_duration, which defaults to true and lets the output length follow the new audio.

Which mode should I start with?

Follow HeyGen's own framing: speed for drafts and for previewing a dub before you commit, precision for the render you publish, especially on long-form or cinematic footage where mouth accuracy shows. Because the body is the same, switching is a one-word change to mode, so you can run a short segment in speed first using start_time and end_time, check the timing of your new audio, and then submit the full video in precision.

Keep in mind that this engine only replaces dialogue. HeyGen's page says the Lip Sync API does no translation and no audio generation: you supply the audio, and its separate Video Translation API can run the same engine after translating.

Can Sume re-sync an existing video the same way?

No. Sume's two lip-sync routes, VEED Fabric 1.0 and MiniMax H3 Max Lip Sync, take a still image plus audio hosted on Sume's media host, not a video. The docs also state that video models do not lip-sync to generated speech or to a later voice-over. To change the language of a talking clip on Sume you generate a new clip; the steps are in dubbing an AI avatar video, and the differences between the two jobs are in lip sync vs dubbing.

The snapshots of HeyGen's pages carry no prices, so this post does not quote one; check HeyGen's own pricing page.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume