Grok Imagine video sound: xAI audio vs Sume's silent row
xAI's Grok Imagine video makes audio unless generate_audio is False. Sume's grok-imagine-video-1.5 row lists audio: false and rejects generate_audio.

Grok Imagine video sound is an xAI-side feature: its docs turn audio generation on and let you disable it with generate_audio=False. The Sume Grok row lists audio: false and rejects generate_audio, so it returns silent video.
xAI's facts are from its video docs and July 31, 2026 post, read 2026-09-30. Sume's are from its catalog and Video Router docs.
What does xAI offer for sound?
Its docs list audio generation, disabled with generate_audio=False, and up to 3 voices per request for reference audio. The July 31 post adds voice consistency through character images and voice references, with voice reference support in the API on request.
What does the Sume row do?
The catalog sets audio: false and reference_audios: false, and its constraints include "no aspect_ratio / bitrate_mode / generate_audio". Treat the output as silent and add sound separately.
| Where | Audio |
|---|---|
| xAI, Grok | Generated; off with generate_audio=False |
| Sume, Grok | None (audio: false) |
Sume, gemini-omni-flash-1.1 | Native synced audio always on |
Sume, wan-3.0 | audio: true in the catalog |
Which Sume ids carry sound?
The docs say gemini-omni-flash-1.1 has native synced audio always on, and generate_audio: false is rejected there. MiniMax H3 Max is documented with native stereo audio. For a general walkthrough see AI video with sound.
What should I do for a Grok clip that needs sound?
Call xAI directly, or keep the silent clip and add music or a voiceover afterwards with a separate tool. Read audio in the catalog before assuming a track exists.
Sources
Related posts
More in Models
- HeyGen ElevenLabs v3 model_id per request vs Sume TTS
HeyGen's speech endpoint takes settings.model_id such as eleven_v3 per request. Sume TTS 1.0 hides model ids; its separate TTS Router lists models explicitly.
- HeyGen photo avatar hand gestures vs Sume Motion Control
HeyGen's motion_prompt still needs an animation reference for Avatar V photo avatars. On Sume, body motion is a separate Motion Control route, not a prompt.
- HeyGen text to video API: 5-15 s, 768p vs Sume durations
heygen-video-1 makes 5-15 second clips at 480p or 768p from a 5,000-character prompt. Sume lists durations and resolutions per model in its catalog.
- HeyGen reference to video: 9 images, 3 videos, 3 audio
HeyGen's heygen-video-1 reference mode takes 9 images, 3 videos and 3 audio files, 12 in total. Here is how Sume's reference limits compare.
Written by Sume