AI sound effects for video: what Sume lists and what it does not
Sume does not list a dedicated sound-effects model. Sound for video comes from video models with generate_audio, the music route, or your files on the Timeline.

Sume's docs do not currently list a dedicated AI sound-effects generator. Sound for a video comes three ways: a video model that generates audio with the clip, a music track from the Music Router, or audio you already have, mixed in Timeline 1.0. If you need a one-shot effect like a door slam as a standalone file, Sume does not list an endpoint for it.
From the Video generation, Music Router and Timeline 1.0 docs, read 2026-09-29.
Which Sume routes make sound?
| Route | Sound it produces |
|---|---|
POST /v1/videos | Models with generate_audio in the catalog can generate audio alongside the video |
POST /v1/music-router/generate | A music track from a prompt; put a sound-design feel in the brief |
POST /v1/tts-1.0/generate | Speech |
POST /v1/timeline-1.0/render | Mixes a voice spine with a soundtrack bed |
How do I get sound with a generated clip?
Check each model's generate_audio flag in the catalog. The request field generate_audio defaults to the model's audio capability, so set it to true and describe the sounds in the prompt: the docs list seedance-2 with audio, and minimax-h3-max with native stereo audio.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sfx-clip-001" \
-d '{
"model": "seedance-2",
"prompt": "Rain on a tin roof, a kettle whistles in the kitchen, close camera.",
"duration": 5,
"resolution": "720p",
"aspect_ratio": "16:9",
"generate_audio": true
}'Can I mix my own effects?
Yes, with files you already have. Import them to media.sume.com, join or trim with Timeline audio, and set a bed as the render's soundtrack. That is mixing existing sound, not generating new effects.
Sources
Related posts
More in Media tools
- Amazon online video ad specs: OLV size, length, bitrate
Amazon online video (OLV) ads run 6–120 s in 16:9, at least 1920×1080 and 4 Mbps, with 192 kbps AAC on 2+ channels and up to 500 MB site-served.
- Audio ad specs: Spotify, Amazon, SiriusXM, and YouTube
Audio ad specs by seller: Spotify wants 192–320 kbps at -16 LUFS, Amazon a 10–30 s file up to 3 MB, SiriusXM a 44.1 kHz MP3, YouTube a video.
- Extract 16 kHz mono audio from a video for speech-to-text
Set sample_rate 16000 and channels mono on Sume's audio detach to get the speech-to-text shape from a video. Options, the 900 second cap and the price.
- Bulk add a watermark to videos: same logo, every file
To bulk add a watermark to videos, run one overlay job per video with the same logo and layout. How to keep it identical by API, and what it costs.
Written by Sume