Vidu S2-Editing: live stream edits vs a Sume Omni clip edit
Vidu S2-Editing edits an incoming video stream in real time. Sume edits one finished clip with a prompt through Gemini Omni video_to_video, as an async job.

Vidu S2-Editing works on a live feed; Sume's edit works on a file. With Sume you send one finished clip and a prompt describing the change to gemini-omni-flash-1.1, and you get a new clip back from an async job, not a stream.
Vidu facts are from its update log; Sume facts are from Video Router and Webhooks, read 2026-09-30.
What does Vidu S2-Editing do?
The September 15, 2026 entry describes "Vidu S2-Editing: Real-Time Editing Model", which uses style, clothing, character and background reference images to edit an incoming video stream in real time. The page gives no request schema, so any integration detail has to come from Vidu.
How does a Sume edit work?
The video_to_video (edit) capability takes a video_url. The prompt describes the edit, resolution is optional (default 720p), and there is no aspect_ratio or duration. video_url is the edit source, not a reference, so it cannot be combined with image_url, end_image_url or reference_*_urls.
Native audio is always on; generate_audio: false is rejected. Send the request to POST /v1/video-router/generate with an Idempotency-Key.
Which parts of the Vidu feature have no Sume equivalent?
| Aspect | Vidu S2-Editing | Sume Omni edit |
|---|---|---|
| Input | Incoming video stream | One clip at video_url |
| Reference images | Style, clothing, character, background | Not combinable with video_url |
| Output | Edited stream in real time | New clip from a job |
| Completion signal | Not stated on the page | job.completed webhook or polling |
How do I know when the edit finished?
Poll the job or take a webhook. The docs list job.completed as "The job completed and a public result is available", plus job.failed and job.canceled. Nothing is sent while the edit is still running.
What should I check before editing a clip?
Uploaded-video edits have an availability limit covered in Gemini Omni edit of an uploaded video: EEA, UK and Switzerland; read it first. For prompt wording, see Gemini Omni video edit prompts.
Sources
Related posts
More in Models
- Wan 3.0 on Runway, or the wan-3.0 id as an API job on Sume
Runway's changelog lists Wan 3.0 in tool mode and workflows on paid plans. On Sume, wan-3.0 is a model id you call from code and poll as a job.
- Dictation API: AssemblyAI cleaned text vs Sume STT word timings
AssemblyAI's Dictation API returns send-ready text. Sume STT returns a transcript with words[] timings and no cleanup flag, so you do the filler removal.
- AssemblyAI text to speech: coming soon, and what Sume has today
AssemblyAI's product menu lists a Text-to-Speech API as coming soon. Sume TTS is available now: what a call takes, the 20000-character cap, and word timings.
- Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.
Written by Sume