Kling motion control with multiple people in the video
Kling motion control animates one character. With two or more people in the reference, it uses the one with the largest share of the frame.

Kling motion control drives one character at a time. If the motion reference video has two or more people, Kling uses the motion of the character who takes up the largest portion of the frame and ignores the rest.
That rule is from Kling's own Motion Control User Guide, read 2026-09-29. What Sume's endpoint accepts comes from the Sume API reference and the API reference docs.
What does Kling do when there are several people?
The guide describes the same idea for the motion reference and for element binding, the facial-consistency feature of 3.0.
| Where | Rule |
|---|---|
| Motion reference video | Upload a single-character reference. With two or more characters, the motion of the one occupying the largest portion of the frame is used. |
| First frame with element binding | May contain several people, but only one element is supported. The person with the largest on-screen presence is selected as the element. |
| Similar sizes in the first frame | If the people occupy similar portions of the frame, no element is selected. |
Can I animate two characters at once?
The guide's own description is that you assign the motion to one character in the image. The guide describes no way to drive two characters from two different people in one motion reference, so a group dance reference will animate only the person who fills the most frame. If you need two moving characters, the guide's rules point to one clip per character, each with its own single-person reference.
How do I make sure the right person is used?
The guide states the selection rule, and the rest follows from it. Three habits keep the choice from being a surprise:
- Use a reference with one person in it, which is what Kling asks for.
- If the clip has a crowd, make sure the person whose motion you want fills more of the frame than anyone else.
- Avoid two people of similar size in the first frame when you bind an element: Kling then selects no element.
What does Sume's endpoint take?
POST /v1/kling/3.0/motion-control takes one visual source (image_url, or a ready avatar_id / avatar_handle) and one motion_video_url, a fetchable public HTTPS video of at most 30 seconds. The schema has no field to name which person in the reference to follow, so the reference itself has to make the choice clear.
The reference describes the result as the still animated with the driving video's motion, with output length following the driving video. For request fields and price, see Motion control API: animate an image with a driving video; for dance-style clips see AI avatar with hand gestures.
Sources
Related posts
More in Models
- Kling motion control not working: causes and fixes
A short, cut-off or wrong-person Kling motion control result usually traces to the reference video. Kling's causes, and the errors Sume's endpoint returns.
- Kling motion control prompt: what it can change
In Kling motion control the video supplies the movement and the prompt covers background and looks. Kling's examples, and Sume's 2,000-character prompt.
- Kling motion control reference video requirements
Kling asks for one character in one continuous shot, 3 to 30 seconds, short edge 340 px or more. Which limits Sume's endpoint also enforces.
- LTX 2 open source: what Lightricks released, and the license
LTX-2 is an audio-video model with open weights under the LTX-2 community license. What the Hugging Face cards list, and why Sume's catalog has no LTX.
Written by Sume