Lyria 3.5 outputs MP3 or WAV: what you get from Sume music
Google lists MP3 by default or WAV for Lyria 3.5. Sume's music request has no format field and returns an audio file; timeline audio can make a wav.

Google's page lists MP3 by default or WAV for Lyria 3.5 at "44.1 kHz high-fidelity stereo". The Sume docs list no format field on the music request: the result's audio artifact is typically audio/mpeg on media.sume.com. A wav file may come from a one-part concat in timeline audio, whose default output is wav; the docs do not list accepted input formats, so test one job.
Sume details are from the Music 1.0 docs and Timeline audio docs, read 2026-09-30.
What does each side list?
Format choices differ at the request, not the sample rate claim.
| Item | Google Gemini API | Sume |
|---|---|---|
| Default | MP3 | Artifact typically audio/mpeg |
| WAV | Available for Lyria 3.5 | Via timeline audio, output.format wav (its default) |
| Format field on the music request | Not stated on the page | None in the docs |
How do I get a wav from a Sume track?
Timeline audio takes parts[] of one to twenty items, so a single part is a valid concat. Send the music artifact URL as the only part. The docs require this workspace's media.sume.com audio and do not say whether a generated music artifact is accepted directly or must be imported first (POST /v1/media-imports), so try one job before building on it. The result is kind: timeline_audio with an audio_url.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: music-to-wav-001" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/track.mp3" }
],
"output": { "format": "wav" }
}'Does a wav restore detail?
No. Re-encoding a lossy file into wav gives a sample-exact container, not the detail that was dropped. The reason to do it is joining or lip-sync work: the docs say to keep wav when the file will be joined again, and that mp3 output re-adds priming padding at every edge. The job costs $0.01 flat.
How do I submit the job and fetch the track?
A music request takes mode async, sync, subscribe or webhook. With sync or subscribe, wait_timeout_seconds is 0 to 30; with webhook, webhook_url must be a public HTTPS callback. Send an Idempotency-Key on the submit, then poll GET /v1/jobs/{job_id}/status and read GET /v1/jobs/{job_id}/result. The audio is the entry in result.artifacts[] where type is audio, hosted on media.sume.com; raw provider URLs are not public outputs. Typically the file is audio/mpeg, per the Music 1.0 docs.
The prompt is 1 to 5000 characters. Put exclusions in the positive prompt ("Instrumental, no vocals"), because a non-empty negative_prompt is refused. Full field list: Music Router docs.
Sources
Related posts
More in Media tools
- Start a video's audio 45 seconds into a song: audio.source_in
Timeline 1.0's audio.source_in sets the in-point into a single audio spine. Output length stays duration_seconds, and it is illegal with parts or silence mode.
- Timeline audio segments: re-base video start times after concat
After a concat, use segments[] (index, start, duration_seconds) as the on-spine start of each video slot in Timeline 1.0. Declared starts are authoritative.
- How to assemble a long-form video with the Timeline 1.0 API
Timeline 1.0 renders one audio spine plus 1 to 200 ordered video slots into one MP4. Every URL must be Sume-hosted; the plan preflight is unbilled.
- How to burn captions onto a video with the Sume API
Send a public HTTPS video URL to POST /v1/video-captions and get a job-backed captioned video, timed by speech-to-text or by text you supply.
Written by Sume