Grok Imagine Video 1.5 text-to-video at 1080p: xAI vs Sume
xAI lists native 1080p text-to-video for grok-imagine-video-1.5. Sume's row is image-to-video only at 480p and 720p. How the two differ and what to send.

On xAI's own API, grok-imagine-video-1.5 supports text-to-video with native 1080p. Sume's grok-imagine-video-1.5 row is image-to-video only and lists 480p and 720p, so a prompt-only request to that row is not what the Sume row is for. Send a start image, or pick a catalog model whose capabilities include text-to-video at 1080p.
The xAI claim is from its July 31 release note; the Sume claims are from the Video Router docs, all read 2026-10-01.
What did xAI release?
The July 31 note says grok-imagine-video-1.5 now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for text-to-video and image-to-video. It also says text-to-video on this model runs as text-to-image and then image-to-video under the hood.
What does the Sume row accept?
The Sume model table lists the row as grok-imagine-video-1.5, 480p and 720p, 4 to 15 seconds, noted "Image-to-video only". The reference-images side of this model has its own post: Grok Imagine Video 1.5 reference images, xAI vs Sume.
| Item | xAI release note | Sume row |
|---|---|---|
| Text-to-video | Supported | Not listed (image-to-video only) |
| Image-to-video | Supported | Supported |
| Reference-to-video | Supported | Not listed |
| Native 1080p | Text-to-video and image-to-video | Resolutions are 480p and 720p |
How do I get a text-prompt video on Sume?
Two documented routes. Generate a still from your text with an image model, then pass it as the image of an image-to-video request on the Grok row; this mirrors what xAI describes happening under the hood, split into two Sume jobs. Or choose a model whose catalog entry covers text-to-video: the Video Router example sends only a prompt to seedance-2.5, which accepts 480p, 720p, and 1080p.
Do not assume either envelope. The docs say to read capabilities from GET /v1/video-router/models rather than assuming one. See the image-to-video handoff for passing a still between jobs.
Does 1080p from xAI carry over to Sume?
No. The xAI 1080p statement describes xAI's endpoint. The Sume row lists 480p and 720p, and that list is what a Sume request is checked against.
Sources
Related posts
More in Models
- HappyHorse 1.0 API: what Runway lists and how to check Sume's models
Runway lists HappyHorse 1.0 at 3-15 seconds, 720p and 1080p. Sume's video docs name no such id, so read /v1/videos/models before you plan around it.
- HappyHorse 1.1 API ids: 480P to 1080P, 3 to 15 seconds
Alibaba lists happyhorse-1.1-t2v, -i2v and -r2v at 480P, 720P or 1080P for 3 to 15 seconds. Sume's catalog docs name HappyHorse 1.1 as not in the v1 catalog.
- Hedra Character 3 aspect ratios vs Sume Avatar Video ratios
Hedra Character 3 lists seven aspect ratios including 9:21 and 21:9. Sume Avatar Video supports five: 1:1, 3:4, 9:16, 4:3 and 16:9, default 9:16.
- HeyGen cinematic avatar API: 3 looks per shot vs Sume's one
HeyGen cinematic_avatar takes a prompt and 1-3 avatar looks with no script. Sume Avatar Video resolves one avatar per final video from a script.
Written by Sume