Sync react-1 model_mode and emotion vs Sume TTS emotion and speed
Sync react-1 edits an existing video with model_mode lips, face or head plus an emotion prompt. On Sume, emotion and speed are set in the TTS audio request.

Sync's react-1 takes a video plus audio and edits the face: options.model_mode picks lips, face (the default) or head, and options.prompt takes a single-word emotion. Sume has no such switch on the video; emotion and speed are set in the text-to-speech request, and the face is then driven by that audio.
react-1 facts are from Sync's React models page; Sume facts from the OpenAPI document and the Models docs, read 2026-10-01.
What do react-1's modes do?
| model_mode | Lip sync | Facial expressions | Head movements |
|---|---|---|---|
| lips | Yes | No | No |
| face (default) | Yes | Yes | No |
| head | Yes | Yes | Yes |
What can the emotion prompt say?
Sync documents single-word emotions only: happy, angry, sad and neutral. Without a prompt the model follows the emotional context of the input video. The page says model_mode works only with react-1 and is ignored by other models, and Sync's free-trial page lists react-1 as paid plans only, with a 15 second input limit in its generation-times table.
Where do emotion and pace live on Sume?
In the TTS request. Its generation_config takes emotion ("Optional emotion guide for generation."), speed ("Speed multiplier in [0.6, 1.5].") and volume ("Volume multiplier in [0.5, 2.0]."). The Models page says every on-camera speaking shot is Fabric with an accepted still plus TTS audio, and video models do not lip-sync to generated TTS or a later voice-over. So you shape the performance in the audio, then pass that audio to the talking-video call. The emotion field is free text up to 64 characters, not a fixed list, and it is a guide, not a guarantee: listen to the result.
Is there a head-movement control?
Not in the docs I read. Sume's talking-video docs describe an avatar, a script or scenes, a quality tier and an aspect ratio, not a face or head mode. For related emotion and pacing advice see AI avatar emotion and speaking speed.
Sources
Related posts
More in Models
- Synthesia Interactive Avatar API vs Sume rendered avatar clips
Synthesia headlines a live Interactive Avatar API. Sume avatar video is script-driven and rendered as a job: submit, poll, then fetch the clip.
- Veo 3.1 only makes 16:9 and 9:16; which Sume video models add more
Google's Veo 3.1 supports 16:9 and 9:16. On Sume, Seedance, MiniMax and Wan add 4:3, 1:1 and 3:4, Kling adds 1:1, and Grok takes no ratio.
- Veo 3.1 prompts cap at 1,024 tokens; Sume's Omni at 20,000 characters
Google caps a Veo 3.1 text prompt at 1,024 tokens. On Sume, gemini-omni-flash-1.1 documents a 20,000-character cap. Tokens and characters differ.
- Vidu Q2 Pro Fast image-to-video 1080p vs Sume first frame
QwenCloud lists vidu/viduq2-pro-fast_img2video at 720P and 1080P. Sume has no Vidu id; send image_url as the first frame to a 1080p model.
Written by Sume