Check whether a reference video is silent before adding music

Reference ingest reports audio.silent at a -60 LUFS gate, speech presence and beats, so you know whether to keep, replace or add a soundtrack before a remix.

4 min readSume
All posts

To check whether a video is silent before you add music, run it through POST /v1/reference-ingest and read audio.silent. The docs define it as true when integrated loudness is at or below −60 LUFS, or the true peak is −inf, so a clip with an audio track that carries no signal counts as silent.

The endpoint is dev-first: the Reference ingest page, read 2026-09-29, says production is opt-in, so confirm it is listed for your workspace.

What else does the audio block report?

Audio facts, from Reference ingest, read 2026-09-29.
FieldMeaning
silentIntegrated loudness at or below −60 LUFS, or true peak −inf
Speech presenceSilero voice-activity detection
BeatsFound with librosa when the track is music
TranscriptOnly with speech.allow_billed_stt, and only if the track is not silent and speech is found

What should I do when it is silent?

The docs say a silent source means: plan new music and discard the source audio, and never request a transcript on it. If you ask for one anyway, the transcript is skipped with stt_skipped_silent and settles to zero. Warnings stt_skipped_no_speech and stt_skipped_no_audio_track cover the other skip cases.

How do I put music under the result?

Build the final video with Timeline 1.0. A render takes one audio spine and ordered video[] slots. A silent video with a music bed can use audio.mode: "silence" for the spine plus a soundtrack, though duck_db needs a real spine. The next posts cover ducking music under a voice-over and rendering a silent video.

Is the check billed?

No. The manifest is unbilled. Only the optional transcript reserves the sume/video-inspect-1.0#transcript per-minute rate, and it settles to what actually ran.

How does the optional transcript work?

Set speech.allow_billed_stt and Sume transcribes only when the track is not silent and voice-activity detection finds speech; otherwise it settles to zero. speech.language_code is a hint and duration_seconds (at most 300) sizes the reservation, but both need allow_billed_stt, or the call fails with reference_ingest_stt_required. The transcript comes back with words and sentence segments.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume