Creatify Boreal talking clips vs Sume's still-plus-audio route

Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.

4 min readSume
All posts

Creatify's own numbers show single-person talking clips are where its Boreal model gains least. Sume sidesteps the question: every on-camera speaking shot is Fabric, built from an accepted still plus TTS audio, and video models are used only for wordless beats.

Boreal claims are from Creatify's 2026-09-15 launch post; Sume's routing is from the Models docs, both read 2026-10-01.

What does Creatify say about talking clips?

In its blind review against its own untouched base model, evaluators preferred Boreal in 81% of decisive comparisons overall: 83% on product ads and 85% on creator and UGC scenes, but 60% on single-person talking clips. These are Creatify's own measurements across 40 production cases, not a general result.

How does Sume route a speaking shot?

The docs say it plainly: "Every on-camera speaking shot is Fabric with an accepted still + TTS", covering short UGC, presenter ads and testimonials. They add that video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is never a video-model clip with narration laid underneath.

Which route fits which shot?

Sume routing by shot type, read 2026-10-01.
ShotRoute in the docs
Person speaking to cameraFabric: accepted still + TTS audio
Speaking, still plus audio, other modelPOST /v1/minimax/h3-max/lip-sync
Wordless beat, B-roll, product motionAuto image, inspect, then Auto video

What should I do with a talking-head brief?

Generate and inspect the still, produce the audio, then send both to Fabric with the measured audio length. Keep product shots on the video route. See avatar vs lip sync vs motion control for how the endpoints differ.

Sources

Related posts

More in Models

All Models posts

Written by Sume