App Store screenshot has an alpha channel: request jpeg
App Store Connect says screenshots can't include alpha channels or transparencies. Ask the Sume image API for output_format jpeg instead of png.

Apple's screenshot specifications say "Images can't include alpha channels or transparencies" and accept .jpeg, .jpg and .png. The simplest way to avoid the error from a Sume image is to request output_format jpeg: JPEG has no alpha channel, so there is nothing to strip.
Apple's rule is from its screenshot specifications page; Sume's formats from the Images docs, read 2026-10-01.
Which formats can I ask for?
| Item | Value |
|---|---|
| Apple accepted extensions | .jpeg, .jpg, .png |
| Apple rule | No alpha channels or transparencies |
| Sume image output_format | png, jpeg, webp, or svg |
| Sume video_frames format | jpeg (default) or png |
Why not just keep png?
A PNG can carry an alpha channel, and Apple's page does not say that a fully opaque PNG is rejected or accepted either way beyond the quoted rule. Choosing jpeg takes the question off the table. The Images docs also list background as auto, transparent or opaque for ChatGPT Image 2.5; for a store screenshot, leave it off transparent.
What about frames pulled from a video?
video_frames returns jpeg by default and png only when you ask for it, described in the docs as "lossless inspection". If the stills are headed for App Store Connect, keep the default. The frame extraction post shows the request shape.
Does jpeg also fix the size?
No. Size is a separate check. With custom pixels on the GPT image models, both edges must be multiples of 16, with a maximum edge of 3840 and an aspect ratio of at most 3:1. Compare your target against the device list on Apple's page; the screenshot size post covers that side.
Sources
Related posts
More in Use cases
- Sync dubbing in 92 languages vs building a dub on Sume
Sync's built-in dubbing now lists 92 languages with speaker detection. Sume has no dubbing endpoint: chain transcription with a language hint, translation, TTS.
- Synthesia dub speaker attribution vs Sume speech-to-text and audio
Synthesia lets Enterprise users reassign and rename speakers in a dub. Sume gives you audio detach, speech-to-text and audio joins; speaker mapping is yours.
- Synthesia Avatar Builder credits: 14 per option, and Sume jobs
Synthesia Avatar Builder charges 14 credits per generated option. Sume creates an avatar with one job per request via POST /v1/avatar-1.0/generate.
- Synthesia brand kit fonts vs Sume caption fonts: Hangul only
Synthesia Motion Graphics now use brand kit fonts. Sume's caption font field takes a Hangul face only, so Latin brand fonts cannot be set there.
Written by Sume