sync-3 image formats JPEG PNG WebP vs Sume image URL rules
sync-3 accepts JPEG, PNG and WebP stills. Sume's docs set URL rules instead: a fetchable public HTTPS image, with private and signed URLs rejected.

sync-3 lists JPEG, PNG and WebP as its supported image formats. Sume's docs do not publish a format list for the Fabric still; what they state is the URL rule: image_url must be a fetchable public HTTPS image URL, and private or signed URLs are rejected before generation.
sync-3 facts are from its model page; Sume facts from Media inputs and Avatar 1.0, read 2026-10-01.
What does sync-3 accept?
The page says a single face image works, with JPEG, PNG or WebP, in image plus audio or image plus text (TTS) combinations. For images with several faces it asks for manual speaker selection with coordinates; auto-detect is not supported for images.
What URL rules does Sume apply?
| URL type | Result |
|---|---|
| Public HTTPS image | Accepted |
| Localhost or private-network | Rejected before submission |
| Non-HTTPS | Rejected before submission |
| Signed or private URL | Rejected before submission |
| Non-image response or mismatched content type | Rejected before submission |
How does the still reach the talking-clip route?
The Fabric body is audio_url, measured duration_seconds, and exactly one visual source; the docs prefer image_url of a generated, inspected posed still. Check the live OpenAPI schema for the exact field shapes before you send.
Which format should I send?
The Sume docs only require that the URL serve an image with a matching content type. Since a format list is not published, test with a small request and read the rejection; see lip-sync API with photo and audio.
Sources
Related posts
More in Developers
- sync-3 lip sync takes 10-15 min: async job and webhook pattern
Sync lists sync-3 at about 10-15 minutes for a 30 s video. Do not hold a request open: submit async, take a terminal webhook, and keep polling as a backup.
- Sync Labs batch API (20 to 500) vs Sume bulk runs (1 to 100)
Sync Labs batch takes 20 to 500 lip sync generations in one JSONL file on Scale and Enterprise. Sume bulk runs queue 1 to 100 Format runs with a 1 to 16 window.
- Sync Labs API rate limit: 100 per minute, and Sume headers
Sync Labs allows 100 POST /v2/generate and 600 GETs a minute. Sume has a per-plan requests-per-minute budget; read ratelimit headers and retry-after.
- Sync lip sync output: H.264 crf 17, HDR to SDR, alpha removed
Sync re-encodes lip sync output to H.264 at crf 17, turns HDR into SDR and drops alpha. Here is that next to what Sume returns for avatar video.
Written by Sume