Engineering guide
Reliable webhook delivery: retries, timeouts, and idempotency
Reliable webhook delivery combines durable handoff, selective retries, stable request identity, receiver-side idempotency, signed requests, and attempt history. Retries alone cannot prevent lost work or duplicate side effects.
A webhook sender has two easy outcomes. It receives a success response, or it fails before opening a connection. The trouble lives between those outcomes. A connection can disappear after the receiver has read the body, updated its database, and started sending a response. The sender sees a timeout. The receiver has already acted.
That uncertainty changes the design. A reliable delivery system needs durable acceptance, a bounded retry policy, a stable identity for repeated attempts, an idempotent receiver, authenticated requests, and enough history to explain what happened. Leave out any one of those and “we retry webhooks” is doing too much work as a reliability claim.
Start with the delivery contract
Reliable delivery is a set of promises between the publisher, the delivery system, and the receiver. It helps to separate them before discussing retry counts.
- 01AcceptValidate and store the request
- 02AttemptOpen a connection and send it
- 03ObserveRecord the response or transport error
- 04DecideFinish or schedule another attempt
- 05ReconcileInspect, cancel, or replay later
Each stage answers a different question:
- Acceptance. Who owns the work after the publisher disconnects?
- Delivery. Did an HTTP request reach the destination?
- Processing. Did the receiver complete the intended business operation?
- Recovery. What happens after a temporary failure or unknown outcome?
- Evidence. Can an operator reconstruct every attempt?
A 2xx response can confirm delivery-system success, but only the receiver can
define what business success means. A payment handler might acknowledge the
request after committing a database transaction. Another receiver might
acknowledge after placing the event on its own durable queue.
A 202 response needs a product contract
RFC 9110 defines 202 Accepted
as an intentionally noncommittal response. The server accepted the request for
processing, but processing has not finished and might never happen. HTTP has no
built-in way to send a later status code on the original exchange.
This means 202 alone does not promise durability. An API must say what it has
done before returning the response.
A useful durable-acceptance contract should answer these questions:
- Has the complete request definition been stored?
- Will work survive the accepting process restarting?
- Does the response contain an ID for later inspection?
- Can the publisher distinguish rejection from acceptance?
- Where can an operator read the eventual outcome?
Stormedo’s contract is explicit. It returns 202 Accepted after storing the
request for durable delivery. The returned req_... value identifies the
request, and delivery proceeds independently of the publisher connection. The
destination has not necessarily received anything when Stormedo returns that
response. See the request lifecycle for
the exact statuses and final outcomes.
That distinction matters during recovery. If Stormedo returns 502 or 503
instead, it did not durably accept the submission. If the publisher loses the
API response altogether, it needs publisher-side idempotency before trying the
submission again.
Classify the failure before retrying
“Retry on failure” sounds reasonable until every instance retries the same struggling destination at once. A retry consumes more connections, CPU, and capacity on the system that just failed. It should happen only when another attempt might work and repeated processing is safe.
This table is a starting point for outbound HTTP. A destination’s documented contract can override it.
| Observed outcome | What the sender can conclude | Policy starting point |
|---|---|---|
| DNS resolution failed | No connection reached the named origin | Retry with backoff |
| TCP or TLS connection could not be opened | No HTTP response exists | Retry with backoff |
| Connection closed while sending | The receiver may have read part or all of the request | Retry only with idempotency |
| Timeout while waiting for a response | The receiver may have completed the operation | Retry only with idempotency |
HTTP 408 Request Timeout |
The server says the request did not complete in time | Retry with backoff |
HTTP 429 Too Many Requests |
The server applied a rate limit | Wait, honor valid guidance, then retry |
HTTP 5xx |
The server or an intermediary failed while handling the exchange | Retry with idempotency |
Other HTTP 4xx |
The request probably needs a caller or configuration change | Stop unless the destination says otherwise |
HTTP 2xx received |
The destination accepted the attempt under its contract | Finish delivery |
RFC 6585 defines 429
for rate limiting and allows the server to include Retry-After.
RFC 9110 defines Retry-After
as either a date or a delay in seconds. A sender should validate the value,
bound it, and fall back to its own retry schedule when the header is invalid.
HTTP status is not always enough. Some APIs use a 4xx response for a
temporary condition. Others use 200 while reporting a rejected operation in
the body. Integrations with such APIs need an explicit destination policy.
Guessing from status classes is safer than retrying everything, but it is not a
substitute for the destination contract.
The timeout case is the one to design around
Here is the failure that catches otherwise careful implementations. The times are illustrative.
12:00:00.000The receiver reads the request.It authenticates the sender and starts a database transaction.
12:00:00.018The receiver commits the side effect.The order is marked paid and the request ID is stored in the same transaction.
12:00:00.021The receiver starts its response.A proxy or network connection fails before the delivery system records it.
12:00:50.000The delivery attempt times out.The sender knows that it did not receive a response. It cannot infer that the receiver did nothing.
laterThe same logical request arrives again.The receiver recognizes its stable ID and returns success without charging the order twice.
This is an unknown outcome. Retrying is the right availability decision, but it can repeat the business effect. Refusing to retry avoids the duplicate and risks losing the operation instead.
There is no timeout value that removes the ambiguity. A shorter timeout reaches it sooner. A longer timeout occupies resources for longer. Neither lets the sender see across a broken connection.
AWS’s guidance on timeouts and retries describes the same trap: a timeout or failure does not prove that side effects did not occur. The practical answer is at-least-once delivery with idempotent processing.
Bound every retry policy
Retries can help a destination recover from a short outage. Unbounded or coordinated retries can keep it down.
A production policy needs four limits:
- A classification rule. Retry failures that may clear without changing the request.
- A delay rule. Increase the interval between attempts instead of retrying immediately.
- Randomization. Add jitter so many failed requests do not return on the same schedule.
- A stopping rule. Limit total attempts, elapsed time, or both.
AWS Well-Architected recommends exponential backoff, jitter, and a maximum retry count. It also warns against retrying errors that need a permission or configuration change.
Retries should live at one deliberate layer. If an application retries a Stormedo submission, Stormedo retries delivery, and the receiver retries its own downstream call, a single failure can expand into many calls. Count the complete path, not only the number configured in one component.
Do not make correctness depend on an exact retry timestamp. Backoff can change,
a service can be busy, and a valid Retry-After can move the next attempt. The
receiver should remain safe whenever the request returns.
Idempotency has two separate boundaries
Webhook systems often discuss idempotency as if it were one switch. There are two network calls and therefore two places where a response can disappear.
Publisher idempotency prevents duplicate submissions
The publisher calls the delivery API. If that response disappears, the publisher cannot tell whether the delivery system accepted the request.
Use a key that names the logical operation, such as
invoice-paid-in_8Ls7, when submitting the work. Repeating the same
submission with that key should return the original request rather than create
a second one.
Stormedo accepts an Idempotency-Key on request creation. The reservation is
scoped to the project and retained for 24 hours. The same key with the same
input returns the original request. Reusing it with different input returns
409 Conflict.
Receiver idempotency prevents duplicate side effects
The delivery system calls the destination. Every automatic retry of one
Stormedo request carries the same Stormedo-Request-Id header.
For a database-only operation, store that validated request ID and apply the business change in the same transaction:
BEGIN;
INSERT INTO processed_deliveries (request_id, processed_at)VALUES ('req_33uZSQsVaf8aZjDuzkuAq', now())ON CONFLICT (request_id) DO NOTHING;
-- Continue only if the insert created a row.UPDATE ordersSET paid = trueWHERE id = 'ord_123';
COMMIT;The application must check whether the insert created a row before running the update. If the ID already exists, it should skip the effect and return a successful response. The unique constraint settles concurrent duplicate attempts without a race.
Do not commit the deduplication row before the business change. A crash between those commits would mark unprocessed work as complete. One transaction avoids that gap when both changes use the same database.
An external side effect cannot join a local database transaction. If the receiver calls another API, pass a stable business idempotency key when that API supports one. Otherwise, commit an outbox record with the deduplication row and process the outbox separately. An arbitrary network call plus a local database write cannot be made exactly once by adding retries.
Automatic retry and manual replay are different work
A stable delivery ID solves duplicates among automatic attempts of one request. It does not necessarily identify the business event forever.
Stormedo replay creates a new request with a new req_... ID and records the
old ID in replayed_from_request_id. That preserves the history of both
executions. It also means a receiver that deduplicates only on
Stormedo-Request-Id will process the replay as new work.
That behavior is useful when replay means “perform this operation again.” If replay means “recover the same payment event without repeating its effect,” put a stable business identifier in the payload and deduplicate on that value as well. The right key depends on what replay means for the application.
Read cancel and replay before building an operational replay button.
Authenticate before recording the request ID
A request ID is an identifier, not proof of origin. If a receiver trusts an
unsigned Stormedo-Request-Id, an attacker could send that ID first and cause
the genuine delivery to look like a duplicate.
The safe order is:
- Capture the raw request body before a parser changes it.
- Reject missing or duplicated delivery headers.
- Verify the delivery signature, expected project, public URL, HTTP method, request ID, time window, and body digest.
- Start the idempotency transaction.
- Apply or durably hand off the business operation.
- Return a successful response.
Stormedo signs each attempt with a project-scoped Ed25519 delivery token. The SDK can authenticate the request context when raw bytes are unavailable, or verify payload integrity when the receiver supplies the exact raw body. The delivery verification guide documents both modes.
Signature verification and idempotency solve different problems. Verification rejects forged or modified requests. Idempotency makes a valid repeated request safe.
A fast response is safe only after durable local work
Webhook providers usually want a prompt 2xx response. GitHub tells receivers
to respond within ten seconds and suggests handing work to an asynchronous
queue. Stripe also recommends asynchronous handling, duplicate detection, and
a quick successful response.
Sources:
The word “queue” matters. Returning 200 and then starting an in-memory task
is fast, but a process restart can lose the task after the sender has stopped
retrying.
Return success after one of these boundaries:
- The business operation and deduplication record committed atomically.
- A durable local queue or outbox accepted the work atomically with its deduplication record.
Do not hold the request open for unrelated slow work when a durable local handoff will do. Do not acknowledge first and hope an in-memory callback finishes.
Attempt history is part of reliability
A final failed status does not explain whether the destination rejected the
request, timed out, or never resolved in DNS. Operators need the history behind
the state.
For each attempt, retain enough information to answer:
- When did the attempt start and finish?
- Did the failure happen during transport or after an HTTP response?
- What response status arrived?
- Was the outcome classified as retryable?
- Why was it classified that way?
- When was the next attempt scheduled?
- Which stable request ID connected the attempts?
The request itself also needs a current status, its planned delivery time, its retry policy, and links to cancellation or replay history. Payloads and secrets need stricter retention and access rules than operational metadata. More data is not automatically better.
Stormedo exposes request state and completed attempt history through its API and dashboard. Cancellation preserves completed history. Replay creates linked new work rather than rewriting the old record.
Choose the smallest system that owns the failure
Not every HTTP call needs durable delivery.
Use an in-process call when the caller can report failure directly and losing work during a restart is acceptable.
Use a queue and worker when the task needs to execute arbitrary code, multiple consumers need the message, or the team already operates that infrastructure.
Use a workflow engine when the work spans several durable steps, waits for external events, needs compensation, or carries long-running business state.
Use a managed HTTP delivery service when the durable unit of work is an HTTP request and the application wants to hand off its timing, retry, attempt history, and final delivery outcome.
Stormedo is in the last category. It does not run customer code and does not replace a multi-step workflow engine.
How Stormedo applies the model
The quick-send API accepts raw request bytes and delivery policy in one call:
curl --request POST \ 'https://api.stormedo.com/send?url=https://your-app.example/webhooks/orders&retry=3' \ --header "Authorization: Bearer $STORMEDO_TOKEN" \ --header 'Idempotency-Key: order-ord_123-created' \ --header 'Content-Type: application/json' \ --data '{"type":"order.created","order_id":"ord_123"}'In this API, retry=3 means one initial attempt and at most three retries. The
default per-attempt timeout is 50 seconds.
The current retry policy treats these outcomes as retryable:
- DNS failures;
- connection failures;
- per-attempt timeouts;
- HTTP
408; - HTTP
429; - HTTP
5xxresponses.
Other 4xx responses are final by default. For a retryable response, Stormedo
uses a valid Retry-After value when it asks for a longer wait than the normal
backoff, bounded to 24 hours. See the retry contract
for current details.
Every automatic attempt carries the stable request ID. Receivers still own authentication, atomic deduplication, and the business result. Stormedo owns durable delivery of the HTTP request and records the transport outcome.
Production checklist
Before relying on a webhook delivery path, verify each item against a forced failure:
- The publisher uses a stable idempotency key when submitting work.
- A
202response has a documented durability meaning. - The publisher stores the returned request ID.
- Retries target temporary or ambiguous failures, not every error.
- Backoff is bounded and retry counts have a stopping rule.
- The receiver verifies the request before trusting its ID.
- Automatic attempts share one stable ID.
- Deduplication and database side effects commit atomically.
- External side effects receive a stable business key or use an outbox.
- Duplicate deliveries return success after skipping completed work.
- A quick acknowledgement happens only after a durable local boundary.
- Operators can inspect every attempt and its retry decision.
- Manual replay semantics match the receiver’s business idempotency key.
- Tests cover a lost response after the receiver commits its work.
The last test is the one most teams miss. A happy-path 200 and a clean 500
do not exercise the unknown outcome that makes webhook delivery hard.
References
- RFC 9110: HTTP Semantics
- RFC 6585: Additional HTTP Status Codes
- AWS Well-Architected: control and limit retry calls
- AWS Builders’ Library: timeouts, retries, and backoff with jitter
- GitHub webhook best practices
- Stripe webhook documentation
- Stormedo request lifecycle
- Stormedo retry behavior
- Stormedo idempotency
- Stormedo delivery verification
Send one durable request
The quickstart uses Stormedo's test destination, so you can inspect durable acceptance and the completed attempt before building a receiver.
Follow the quickstart