PixVerse API and R2 world model: Sume has no live session
PixVerse R2 is a persistent real-time world. Sume has no live session or action controls: video jobs return a clip you poll for or get by webhook.

PixVerse R2 is a real-time world model, and Sume does not ship that kind of product. Sume's video API is request and response: you submit a job to POST /v1/videos, then poll it or receive one terminal webhook, and you get a fixed clip.
R2 facts are from PixVerse's announcement; Sume facts are from Video generation and Webhooks, read 2026-09-30.
What is R2, according to PixVerse?
The September 22, 2026 post says a real-time world model "generates a world that keeps running while a user is in it" instead of returning a clip. R2 takes text, references, audio and action controls into the same running world, and the post says input persists in the world's state. The example shows a player moving with WASD keys while prompts change the scene.
What does Sume's video API do instead?
Sume's docs say video generation is asynchronous: submit, receive a job id and polling URL, poll until completed, then download. You can pass callback_url and Sume POSTs to it. It sends terminal job events only, with no progress or partial deliveries, so nothing arrives while a clip renders.
Which R2 inputs map to a Sume request field?
| R2 input (vendor post) | Closest Sume field | Match |
|---|---|---|
| Text | prompt | Yes, one prompt per job |
| References | input_references (reference images) | Partly: stills for style guidance |
| Audio | generate_audio | Partly: audio generated with the clip, not steered live |
| Action controls | None | No request field exists |
| Persistent world state | None | Each job is independent |
How do I see what a model accepts?
The catalog lists each model's capabilities, including supported_input_references, which says which reference types the model accepts. Frame control exists as frame_images (first and last frames for image-to-video). See OpenRouter-compatible video generation for the request shape.
When is a clip job the wrong tool?
If the user steers the output while it plays, as in R2's game example, a request-response clip cannot do that. If you know the shot ahead of time and want a file for editing or publishing, a job fits.
Sources
Related posts
More in Models
- Recraft V4.1 Flash API: Sume lists V4, webp, text-only
Recraft V4.1 Flash is live in Recraft Studio and its API. Sume lists recraft/recraft-v4, not a Flash id: webp output, text-to-image only, no references.
- Medical transcription API: what Sume STT does and doesn't
Sume has one speech-to-text route, sume/stt-1.0, with no medical model option. What it takes and returns, so you can decide for clinical audio.
- Seedance 1.5 Pro retires Nov 11: which Seedance ids Sume lists
ByteDance retires Seedance 1.5 Pro on November 11, 2026. Sume's video docs list seedance-2.5 and seedance-2; use bare catalog ids and check the models endpoint.
- Seedance 2.5 reference limit: BytePlus says 50, Sume varies by model
BytePlus allows up to 50 multimodal references for one Seedance 2.5 clip. Sume documents limits per model, so read the catalog for the id you send. Examples.
Written by Sume