Skip to content
StormedoStart free
Menu

Engineering guide

Reliable webhook delivery: retries, timeouts, and idempotency

Reliable webhook delivery combines durable handoff, selective retries, stable request identity, receiver-side idempotency, signed requests, and attempt history. Retries alone cannot prevent lost work or duplicate side effects.

A webhook sender has two easy outcomes. It receives a success response, or it fails before opening a connection. The trouble lives between those outcomes. A connection can disappear after the receiver has read the body, updated its database, and started sending a response. The sender sees a timeout. The receiver has already acted.

That uncertainty changes the design. A reliable delivery system needs durable acceptance, a bounded retry policy, a stable identity for repeated attempts, an idempotent receiver, authenticated requests, and enough history to explain what happened. Leave out any one of those and “we retry webhooks” is doing too much work as a reliability claim.

Start with the delivery contract

Reliable delivery is a set of promises between the publisher, the delivery system, and the receiver. It helps to separate them before discussing retry counts.

  1. 01AcceptValidate and store the request
  2. 02AttemptOpen a connection and send it
  3. 03ObserveRecord the response or transport error
  4. 04DecideFinish or schedule another attempt
  5. 05ReconcileInspect, cancel, or replay later

Each stage answers a different question:

  • Acceptance. Who owns the work after the publisher disconnects?
  • Delivery. Did an HTTP request reach the destination?
  • Processing. Did the receiver complete the intended business operation?
  • Recovery. What happens after a temporary failure or unknown outcome?
  • Evidence. Can an operator reconstruct every attempt?

A 2xx response can confirm delivery-system success, but only the receiver can define what business success means. A payment handler might acknowledge the request after committing a database transaction. Another receiver might acknowledge after placing the event on its own durable queue.

A 202 response needs a product contract

RFC 9110 defines 202 Accepted as an intentionally noncommittal response. The server accepted the request for processing, but processing has not finished and might never happen. HTTP has no built-in way to send a later status code on the original exchange.

This means 202 alone does not promise durability. An API must say what it has done before returning the response.

A useful durable-acceptance contract should answer these questions:

  • Has the complete request definition been stored?
  • Will work survive the accepting process restarting?
  • Does the response contain an ID for later inspection?
  • Can the publisher distinguish rejection from acceptance?
  • Where can an operator read the eventual outcome?

Stormedo’s contract is explicit. It returns 202 Accepted after storing the request for durable delivery. The returned req_... value identifies the request, and delivery proceeds independently of the publisher connection. The destination has not necessarily received anything when Stormedo returns that response. See the request lifecycle for the exact statuses and final outcomes.

That distinction matters during recovery. If Stormedo returns 502 or 503 instead, it did not durably accept the submission. If the publisher loses the API response altogether, it needs publisher-side idempotency before trying the submission again.

Classify the failure before retrying

“Retry on failure” sounds reasonable until every instance retries the same struggling destination at once. A retry consumes more connections, CPU, and capacity on the system that just failed. It should happen only when another attempt might work and repeated processing is safe.

This table is a starting point for outbound HTTP. A destination’s documented contract can override it.

Observed outcome What the sender can conclude Policy starting point
DNS resolution failed No connection reached the named origin Retry with backoff
TCP or TLS connection could not be opened No HTTP response exists Retry with backoff
Connection closed while sending The receiver may have read part or all of the request Retry only with idempotency
Timeout while waiting for a response The receiver may have completed the operation Retry only with idempotency
HTTP 408 Request Timeout The server says the request did not complete in time Retry with backoff
HTTP 429 Too Many Requests The server applied a rate limit Wait, honor valid guidance, then retry
HTTP 5xx The server or an intermediary failed while handling the exchange Retry with idempotency
Other HTTP 4xx The request probably needs a caller or configuration change Stop unless the destination says otherwise
HTTP 2xx received The destination accepted the attempt under its contract Finish delivery

RFC 6585 defines 429 for rate limiting and allows the server to include Retry-After. RFC 9110 defines Retry-After as either a date or a delay in seconds. A sender should validate the value, bound it, and fall back to its own retry schedule when the header is invalid.

HTTP status is not always enough. Some APIs use a 4xx response for a temporary condition. Others use 200 while reporting a rejected operation in the body. Integrations with such APIs need an explicit destination policy. Guessing from status classes is safer than retrying everything, but it is not a substitute for the destination contract.

The timeout case is the one to design around

Here is the failure that catches otherwise careful implementations. The times are illustrative.

  1. 12:00:00.000
    The receiver reads the request.

    It authenticates the sender and starts a database transaction.

  2. 12:00:00.018
    The receiver commits the side effect.

    The order is marked paid and the request ID is stored in the same transaction.

  3. 12:00:00.021
    The receiver starts its response.

    A proxy or network connection fails before the delivery system records it.

  4. 12:00:50.000
    The delivery attempt times out.

    The sender knows that it did not receive a response. It cannot infer that the receiver did nothing.

  5. later
    The same logical request arrives again.

    The receiver recognizes its stable ID and returns success without charging the order twice.

This is an unknown outcome. Retrying is the right availability decision, but it can repeat the business effect. Refusing to retry avoids the duplicate and risks losing the operation instead.

There is no timeout value that removes the ambiguity. A shorter timeout reaches it sooner. A longer timeout occupies resources for longer. Neither lets the sender see across a broken connection.

AWS’s guidance on timeouts and retries describes the same trap: a timeout or failure does not prove that side effects did not occur. The practical answer is at-least-once delivery with idempotent processing.

Bound every retry policy

Retries can help a destination recover from a short outage. Unbounded or coordinated retries can keep it down.

A production policy needs four limits:

  1. A classification rule. Retry failures that may clear without changing the request.
  2. A delay rule. Increase the interval between attempts instead of retrying immediately.
  3. Randomization. Add jitter so many failed requests do not return on the same schedule.
  4. A stopping rule. Limit total attempts, elapsed time, or both.

AWS Well-Architected recommends exponential backoff, jitter, and a maximum retry count. It also warns against retrying errors that need a permission or configuration change.

Retries should live at one deliberate layer. If an application retries a Stormedo submission, Stormedo retries delivery, and the receiver retries its own downstream call, a single failure can expand into many calls. Count the complete path, not only the number configured in one component.

Do not make correctness depend on an exact retry timestamp. Backoff can change, a service can be busy, and a valid Retry-After can move the next attempt. The receiver should remain safe whenever the request returns.

Idempotency has two separate boundaries

Webhook systems often discuss idempotency as if it were one switch. There are two network calls and therefore two places where a response can disappear.

Publisher idempotency prevents duplicate submissions

The publisher calls the delivery API. If that response disappears, the publisher cannot tell whether the delivery system accepted the request.

Use a key that names the logical operation, such as invoice-paid-in_8Ls7, when submitting the work. Repeating the same submission with that key should return the original request rather than create a second one.

Stormedo accepts an Idempotency-Key on request creation. The reservation is scoped to the project and retained for 24 hours. The same key with the same input returns the original request. Reusing it with different input returns 409 Conflict.

Receiver idempotency prevents duplicate side effects

The delivery system calls the destination. Every automatic retry of one Stormedo request carries the same Stormedo-Request-Id header.

For a database-only operation, store that validated request ID and apply the business change in the same transaction:

BEGIN;
INSERT INTO processed_deliveries (request_id, processed_at)
VALUES ('req_33uZSQsVaf8aZjDuzkuAq', now())
ON CONFLICT (request_id) DO NOTHING;
-- Continue only if the insert created a row.
UPDATE orders
SET paid = true
WHERE id = 'ord_123';
COMMIT;

The application must check whether the insert created a row before running the update. If the ID already exists, it should skip the effect and return a successful response. The unique constraint settles concurrent duplicate attempts without a race.

Do not commit the deduplication row before the business change. A crash between those commits would mark unprocessed work as complete. One transaction avoids that gap when both changes use the same database.

An external side effect cannot join a local database transaction. If the receiver calls another API, pass a stable business idempotency key when that API supports one. Otherwise, commit an outbox record with the deduplication row and process the outbox separately. An arbitrary network call plus a local database write cannot be made exactly once by adding retries.

Automatic retry and manual replay are different work

A stable delivery ID solves duplicates among automatic attempts of one request. It does not necessarily identify the business event forever.

Stormedo replay creates a new request with a new req_... ID and records the old ID in replayed_from_request_id. That preserves the history of both executions. It also means a receiver that deduplicates only on Stormedo-Request-Id will process the replay as new work.

That behavior is useful when replay means “perform this operation again.” If replay means “recover the same payment event without repeating its effect,” put a stable business identifier in the payload and deduplicate on that value as well. The right key depends on what replay means for the application.

Read cancel and replay before building an operational replay button.

Authenticate before recording the request ID

A request ID is an identifier, not proof of origin. If a receiver trusts an unsigned Stormedo-Request-Id, an attacker could send that ID first and cause the genuine delivery to look like a duplicate.

The safe order is:

  1. Capture the raw request body before a parser changes it.
  2. Reject missing or duplicated delivery headers.
  3. Verify the delivery signature, expected project, public URL, HTTP method, request ID, time window, and body digest.
  4. Start the idempotency transaction.
  5. Apply or durably hand off the business operation.
  6. Return a successful response.

Stormedo signs each attempt with a project-scoped Ed25519 delivery token. The SDK can authenticate the request context when raw bytes are unavailable, or verify payload integrity when the receiver supplies the exact raw body. The delivery verification guide documents both modes.

Signature verification and idempotency solve different problems. Verification rejects forged or modified requests. Idempotency makes a valid repeated request safe.

A fast response is safe only after durable local work

Webhook providers usually want a prompt 2xx response. GitHub tells receivers to respond within ten seconds and suggests handing work to an asynchronous queue. Stripe also recommends asynchronous handling, duplicate detection, and a quick successful response.

Sources:

The word “queue” matters. Returning 200 and then starting an in-memory task is fast, but a process restart can lose the task after the sender has stopped retrying.

Return success after one of these boundaries:

  • The business operation and deduplication record committed atomically.
  • A durable local queue or outbox accepted the work atomically with its deduplication record.

Do not hold the request open for unrelated slow work when a durable local handoff will do. Do not acknowledge first and hope an in-memory callback finishes.

Attempt history is part of reliability

A final failed status does not explain whether the destination rejected the request, timed out, or never resolved in DNS. Operators need the history behind the state.

For each attempt, retain enough information to answer:

  • When did the attempt start and finish?
  • Did the failure happen during transport or after an HTTP response?
  • What response status arrived?
  • Was the outcome classified as retryable?
  • Why was it classified that way?
  • When was the next attempt scheduled?
  • Which stable request ID connected the attempts?

The request itself also needs a current status, its planned delivery time, its retry policy, and links to cancellation or replay history. Payloads and secrets need stricter retention and access rules than operational metadata. More data is not automatically better.

Stormedo exposes request state and completed attempt history through its API and dashboard. Cancellation preserves completed history. Replay creates linked new work rather than rewriting the old record.

Choose the smallest system that owns the failure

Not every HTTP call needs durable delivery.

Use an in-process call when the caller can report failure directly and losing work during a restart is acceptable.

Use a queue and worker when the task needs to execute arbitrary code, multiple consumers need the message, or the team already operates that infrastructure.

Use a workflow engine when the work spans several durable steps, waits for external events, needs compensation, or carries long-running business state.

Use a managed HTTP delivery service when the durable unit of work is an HTTP request and the application wants to hand off its timing, retry, attempt history, and final delivery outcome.

Stormedo is in the last category. It does not run customer code and does not replace a multi-step workflow engine.

How Stormedo applies the model

The quick-send API accepts raw request bytes and delivery policy in one call:

Terminal window
curl --request POST \
'https://api.stormedo.com/send?url=https://your-app.example/webhooks/orders&retry=3' \
--header "Authorization: Bearer $STORMEDO_TOKEN" \
--header 'Idempotency-Key: order-ord_123-created' \
--header 'Content-Type: application/json' \
--data '{"type":"order.created","order_id":"ord_123"}'

In this API, retry=3 means one initial attempt and at most three retries. The default per-attempt timeout is 50 seconds.

The current retry policy treats these outcomes as retryable:

  • DNS failures;
  • connection failures;
  • per-attempt timeouts;
  • HTTP 408;
  • HTTP 429;
  • HTTP 5xx responses.

Other 4xx responses are final by default. For a retryable response, Stormedo uses a valid Retry-After value when it asks for a longer wait than the normal backoff, bounded to 24 hours. See the retry contract for current details.

Every automatic attempt carries the stable request ID. Receivers still own authentication, atomic deduplication, and the business result. Stormedo owns durable delivery of the HTTP request and records the transport outcome.

Production checklist

Before relying on a webhook delivery path, verify each item against a forced failure:

  • The publisher uses a stable idempotency key when submitting work.
  • A 202 response has a documented durability meaning.
  • The publisher stores the returned request ID.
  • Retries target temporary or ambiguous failures, not every error.
  • Backoff is bounded and retry counts have a stopping rule.
  • The receiver verifies the request before trusting its ID.
  • Automatic attempts share one stable ID.
  • Deduplication and database side effects commit atomically.
  • External side effects receive a stable business key or use an outbox.
  • Duplicate deliveries return success after skipping completed work.
  • A quick acknowledgement happens only after a durable local boundary.
  • Operators can inspect every attempt and its retry decision.
  • Manual replay semantics match the receiver’s business idempotency key.
  • Tests cover a lost response after the receiver commits its work.

The last test is the one most teams miss. A happy-path 200 and a clean 500 do not exercise the unknown outcome that makes webhook delivery hard.

References

Send one durable request

The quickstart uses Stormedo's test destination, so you can inspect durable acceptance and the completed attempt before building a receiver.

Follow the quickstart