Tenacity retry in Python: safe settings for a paid API POST
Bare @retry in tenacity retries forever with no wait. For a paid POST, cap attempts, add jittered backoff, retry only transient errors, and reuse one key.

Tenacity's bare @retry retries forever without waiting whenever the function raises, so a POST that costs money needs three settings: stop=stop_after_attempt(n), a jittered backoff such as wait_random_exponential, and a retry= predicate that matches only transient failures. Then create the idempotency key outside the decorated function, so every attempt sends the same one and a retry can't pay twice.
Tenacity facts come from its documentation and API reference. The retry rules for the example API come from Sume's Errors and spend, Create a run and Jobs and results pages. All were read on 2026-09-29. For urllib3's built-in retries instead, see Python requests retry.
Which tenacity settings does a paid POST need?
stop=stop_after_attempt(5)bounds the attempts. Combine it withstop_after_delay(...)using|to also bound the time.wait=wait_random_exponential(multiplier=1, max=60)waits a random time up to 2^x seconds, capped at 60. Tenacity says exponentially increasing jitter helps when several processes compete for a shared resource.retry=retry_if_exception_type((...))names the exceptions worth retrying. Anything else is raised at once.reraise=Trueraises your own exception after the last attempt instead of tenacity'sRetryError.- A custom
waitis any function ofretry_statethat returns the seconds to wait, which is how you honor a server'sretry-after.
Which errors should tenacity retry?
Only the ones where resending the same request can succeed. Not every 5xx qualifies: Sume's 502 attachment_fetch_failed means one of your URLs couldn't be fetched, and resending won't change that. A 4xx at create means nothing ran and nothing was charged. Retryable HTTP status codes covers the general rules; the table lists the create answers the code below retries.
| Response | Retry? | What Sume says to do |
|---|---|---|
409 idempotency_key_in_use | Yes | Wait about a second and resend |
429 rate_limited | Yes | Wait retry-after |
503 studio_agent_upstream_unavailable | Yes | Retry later with the same Idempotency-Key |
402 insufficient_credits | No | Top up; retrying returns the same answer |
409 idempotency_conflict | No | Fix your key derivation; do not retry as is |
401, 403, 404 | No | Fix the key or the address |
How do I write the retry for a paid POST?
Raise a dedicated exception for the retryable answers, retry that plus connection errors and timeouts, and pass the key in. The wait uses retry-after when the server sent one:
import os, requests
from tenacity import retry, retry_if_exception_type, stop_after_attempt
from tenacity import wait_random_exponential
URL = "https://api.sume.com/v1/formats/acme/product-promo/runs"
backoff = wait_random_exponential(multiplier=1, max=60)
class Transient(Exception):
def __init__(self, status, retry_after=None):
super().__init__(f"Sume answered {status}")
self.retry_after = retry_after
def wait_sume(retry_state):
after = getattr(retry_state.outcome.exception(), "retry_after", None)
return float(after) if after else backoff(retry_state)
@retry(stop=stop_after_attempt(5), wait=wait_sume, reraise=True,
retry=retry_if_exception_type((requests.ConnectionError, requests.Timeout, Transient)))
def start_run(body, idempotency_key):
r = requests.post(URL, json=body, timeout=30, headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": idempotency_key})
code = r.json().get("error", {}).get("code") if not r.ok else None
if code in ("idempotency_key_in_use", "rate_limited") or r.status_code == 503:
raise Transient(r.status_code, r.headers.get("retry-after"))
r.raise_for_status()
return r.json()["data"]
run = start_run({"instruction": "15-second vertical promo."}, "order-8823-promo-v1")Why must the idempotency key live outside @retry?
Tenacity calls the whole decorated function again on every attempt. A uuid.uuid4() inside it makes each attempt a new request as far as the server knows; Sume's docs say a per-request UUID makes the header decorative. Derive the key from the thing being made, such as an order id and a version, and pass it in.
With one key, a retried create is safe: the same key and the same body return 200 with the original run and idempotency_hit: true, with no second run and no second charge. That also covers the case the retry exists for, a timeout after the server already accepted the request. Without a key, don't resubmit a paid request because a local process timed out; poll the job instead.
What does tenacity not handle?
- It doesn't know the work finished. A
202means the run was accepted, not that it finished; learn the outcome fromstatus_urlor a webhook, not by calling the create again. - It doesn't re-run failed runs. A run that ends
failedafter202needs a newIdempotency-Key, because the old one is bound to the failed run. - It doesn't dedupe across processes. Two workers that start the same order at once get one run and one
409 idempotency_key_in_use, which the code above retries into the original run.
Sources
Related posts
More in Integrations
- Workato HTTP connector: call an API with a key and JSON
Workato's HTTP connector calls any HTTP API: a Header auth connection holds the key, and Send request via HTTP POSTs a JSON body and maps the reply.
- How to add an MCP server to ChatGPT with developer mode
Turn on ChatGPT developer mode, create an app for the server's URL, and sign in with OAuth. The steps, with Sume's hosted MCP server as the example.
- How to add subtitles to a video in Python
Add subtitles to a video in Python with Requests: POST the video URL to Sume's /v1/video-captions, poll the job, then read the captioned video_url.
- Add Sume to Claude as a custom connector (remote MCP)
Add Sume's hosted MCP server to Claude under Customize > Connectors, see what Sume's OAuth consent grants, and decide whether to allow paid tools.
Written by Sume