Open source AI video generator with audio: LTX-2 and MiniMax H3

Two open-weights video models make sound with the picture: LTX-2 and MiniMax H3. What each page says, and how a hosted API reports audio.

4 min readSume
All posts

Yes: two open-weights video models generate audio together with the picture. Lightricks describes LTX-2 as generating synchronized video and audio within a single model, and MiniMax says its open-sourced H3 generates video with native stereo audio.

Read 2026-09-29: the LTX-2 model card and MiniMax's H3 open-source announcement. Sume's side comes from Video generation.

What do the two pages say about sound?

From the vendors' own pages, read 2026-09-29.
ModelWhat the page says about audio
LTX-2 (Lightricks)A DiT-based audio-video foundation model that generates synchronized video and audio within a single model; audio without speech may be of lower quality
MiniMax H3Generates video with native stereo audio; output audio is 32 kHz stereo; stable dialogue support for 11 languages

Can I get generated audio from a hosted API instead?

Sume's video catalog reports audio per model through a generate_audio field, and the generate request accepts a generate_audio boolean that defaults to the model's audio capability. Two listed models are described with audio in the docs: minimax-h3-max with native stereo audio, and gemini-omni-flash-1.1 with native synced audio.

Read the field from the catalog rather than assuming it, because models differ. The docs do not say that the hosted MiniMax models run the open-sourced checkpoints, so do not treat the hosted ids and the downloadable weights as the same thing.

curl "https://api.sume.com/v1/catalog" \
  -H "Authorization: Bearer $SUME_API_KEY"

What input modes does the open MiniMax H3 release have?

MiniMax's page lists two base variants. H3-Base-FL2VA supports zero, one, or two input images (text-to-video, first-frame or last-frame, or first-and-last-frame). H3-Base-Ref2VA is an omni-reference mode with up to 9 images, up to 3 video clips and up to 3 audio clips, and audio must be accompanied by image or video input rather than used alone.

What audio can a hosted model take as input?

Audio in is a different feature from audio out. Sume's docs say audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3, and MiniMax H3 Max, while Gemini Omni Flash 1.1 accepts video references but not audio. MiniMax's open-source page lists H3 output at 4-15 seconds and 24 FPS, so check the hosted catalog for each hosted id's own limits.

How do I add sound to a model that has none?

Video models that return silent clips can be paired with separate sound. AI video with sound: generate audio covers the options on Sume, and MiniMax H3 open weights on Hugging Face covers what the H3 release contains.

What does this answer leave out?

  • Only cards read on 2026-09-29 count: other open models may also generate audio.
  • Audio quality is not compared here; neither page gives a like-for-like measure.
  • Both licenses are their own agreements (ltx-2-community-license-agreement and MiniMax's H3 community license); read them before commercial use.

Sources

Related posts

More in Models

All Models posts

Written by Sume