Choosing a Webhook Delivery Guarantee Level

Every webhook system makes a delivery guarantee, whether it is chosen deliberately or not. The guarantee is the contract you offer consumers about how often each event arrives: never more than once, at least once, or exactly once in effect. Picking the right tier is the most consequential design decision in Delivery Guarantee Levels, because it dictates your retry policy, your storage footprint, and how much work consumers must do. This guide gives you a concrete procedure and a comparison table to choose, then shows how to encode the choice in code. For the deeper theory behind why exactly-once is effectively a deduplication problem, read at-least-once vs exactly-once delivery trade-offs.

The short version: at-least-once is the right default for almost all webhooks, paired with consumer-side idempotency. At-most-once and effectively-exactly-once are specializations you reach for only when a specific requirement forces them.

Delivery guarantee decision tree Branches on whether event loss is acceptable and whether consumers can deduplicate, leading to at-most-once, at-least-once, or effectively-exactly-once. Is event loss acceptable? At-most-once no retries Can consumer dedupe? At-least-once retry + idempotency Effectively exactly-once (dedup store) yes no yes no
The decision tree: loss tolerance picks at-most-once; otherwise dedup capability decides between at-least-once and an effectively-exactly-once dedup layer.

The three guarantee levels compared

There is no true exactly-once over an unreliable network — “exactly-once” in practice means at-least-once delivery plus deduplication that makes reprocessing a no-op. The three operational tiers are:

Dimension At-most-once At-least-once Effectively-exactly-once
Duplicates Never Possible (must be tolerated) Suppressed by a dedup store
Lost events Possible on any failure Never (retried until acked) Never
Retries None Bounded retries + backoff Bounded retries + backoff
Where dedup lives N/A Consumer (its responsibility) Producer dedup store + consumer idempotency
Storage cost Lowest Low (delivery log only) Highest (durable dedup keys with long TTL)
Latency overhead Lowest Low Added lookup/write per event
Typical use Metrics, presence pings, ephemeral signals Most business events Payments, billing, inventory

The decisive trade-off is duplicates versus loss. You cannot have neither without unbounded cost; you choose which one your consumers can tolerate cheaply. Most can absorb duplicates with a small idempotency check far more easily than they can recover from a silently dropped event.

It helps to see exactly where a duplicate is born, because it is almost never a bug in the dispatcher. The canonical case is a response that never makes it back: the consumer has already applied the effect, the acknowledgement is lost in transit, and the dispatcher — which cannot distinguish a lost response from a lost request — correctly retries.

How a lost acknowledgement creates a duplicate The consumer applies an event and claims its key, its 200 response is lost in transit, the dispatcher retries with the same idempotency key, and the deduplication store turns the second delivery into a no-op. Dispatcher Consumer endpoint Dedup store POST attempt 1, key evt-9 SET NX evt-9 claimed: apply the effect 200 OK lost in transit POST attempt 2, key evt-9 SET NX evt-9 already set: skip the effect delivered twice, applied once
The duplicate is created by the network, not the dispatcher — which is why the deduplication check, wherever it lives, is the thing that actually defines your guarantee.

Step 1: Classify each event’s loss tolerance

Run the decision per event type, not per system. A single dispatcher often carries metric.sampled (loss fine) alongside invoice.paid (loss catastrophic). For each type, ask: if this event vanishes and no one notices for an hour, what breaks?

from enum import Enum

class Guarantee(Enum):
    AT_MOST_ONCE = "at_most_once"
    AT_LEAST_ONCE = "at_least_once"
    EFFECTIVELY_EXACTLY_ONCE = "eeo"

EVENT_POLICY = {
    "metric.sampled":   Guarantee.AT_MOST_ONCE,      # cheap, replaceable
    "user.updated":     Guarantee.AT_LEAST_ONCE,     # dedupe on consumer
    "invoice.paid":     Guarantee.EFFECTIVELY_EXACTLY_ONCE,  # money
}

If loss is acceptable, choose at-most-once: fire once, no retries, no delivery log. Everything else continues to Step 2.

Step 2: Decide who owns deduplication

For events that must not be lost, first make sure they cannot be lost before dispatch: at-least-once starts at the write boundary, so the event has to be persisted in the same transaction as the state change, which is what implementing the transactional outbox pattern for webhooks buys you. With that in place, the next question is whether the consumer can deduplicate. A consumer that writes through a unique idempotency key (e.g. an INSERT ... ON CONFLICT DO NOTHING keyed on the event ID) makes at-least-once safe with almost no extra machinery on the producer side. This is the sweet spot and your default:

def deliver_at_least_once(dispatcher, event, max_attempts=6):
    # Retry until acknowledged; the consumer dedupes on event["id"].
    for attempt in range(max_attempts):
        ok = dispatcher.post(event, headers={"X-Idempotency-Key": event["id"]})
        if ok:
            return "delivered"
        dispatcher.backoff(attempt)   # exponential backoff + jitter
    dispatcher.dead_letter(event)
    return "exhausted"

If you cannot rely on the consumer to dedupe — for example a third-party endpoint with side effects you don’t control — escalate to Step 3 and provide deduplication on the producer side.

Step 3: Cost the strongest tier before committing

Effectively-exactly-once is not free. It requires a durable dedup store whose key TTL must outlive your maximum retry window and any DLQ retention, so a replayed event from a dead-letter queue is still recognized as a duplicate weeks later. Concretely: with six attempts on a jittered exponential schedule the live retry window closes inside an hour, but a dead-letter queue that retains payloads for 14 days can resurface that same event on day 13. A 30-day deduplication TTL covers both with margin; a 24-hour TTL — the value most teams reach for first — expires 13 days before the dead-letter window even closes.

Sizing the deduplication TTL The retry window closes within an hour, dead-letter retention runs to fourteen days, and the deduplication key TTL must extend past both, to thirty days. Every window that can re-deliver one event live retry window: 6 attempts, under 1h dead-letter retention: 14 days deduplication key TTL: 30 days t0 1h 24h 14d 30d A 24h TTL expires 13 days before the last legitimate replay
The deduplication TTL is not a cache-sizing decision — it is the outer bound of every path that can present the same event twice, dead-letter replay included.

Budget for the storage and the per-event lookup latency:

import redis

class ExactlyOnceGate:
    def __init__(self, client: redis.Redis, ttl_days: int = 30):
        self.r = client
        self.ttl = ttl_days * 24 * 3600

    def first_delivery(self, event_id: str) -> bool:
        # SET NX is the dedup primitive; returns True only the first time.
        return bool(self.r.set(f"eeo:{event_id}", "1", nx=True, ex=self.ttl))

def deliver_eeo(dispatcher, gate: ExactlyOnceGate, event):
    if not gate.first_delivery(event["id"]):
        return "suppressed_duplicate"
    return deliver_at_least_once(dispatcher, event)

Only adopt this tier where a duplicate genuinely causes harm that the consumer cannot undo cheaply — double charges, double shipments, double-counted balances.

Step 4: Encode the choice in the dispatcher

Wire the per-event policy into one place so the guarantee is explicit and testable rather than emergent from scattered retry settings. One lookup, three code paths — and because the lookup is a plain table, the guarantee for any event type is answerable by reading a single dictionary rather than tracing retry configuration through the dispatcher.

Per-event guarantee routing A single dispatch function consults the event policy table and routes each event type into the at-most-once, at-least-once or effectively-exactly-once path. dispatch(event) one entry point EVENT_POLICY lookup default: at-least-once At-most-once path post once, no retry, no delivery log metric.sampled At-least-once path bounded retry, idempotency header user.updated Effectively-exactly-once path SET NX gate, then bounded retry invoice.paid
Mixing tiers inside one dispatcher is the correct design: the policy table is the only place the guarantee is decided, so it is also the only place it can be audited.
def dispatch(event, dispatcher, gate):
    level = EVENT_POLICY.get(event["type"], Guarantee.AT_LEAST_ONCE)
    if level is Guarantee.AT_MOST_ONCE:
        dispatcher.post(event)              # fire and forget, no retry
        return "sent_best_effort"
    if level is Guarantee.AT_LEAST_ONCE:
        return deliver_at_least_once(dispatcher, event)
    return deliver_eeo(dispatcher, gate, event)

Verification

Assert that each tier behaves as contracted under a forced failure.

def test_at_least_once_retries_then_delivers():
    calls = {"n": 0}
    class D:
        def post(self, e, headers=None):
            calls["n"] += 1
            return calls["n"] >= 3        # fail twice, then succeed
        def backoff(self, a): pass
        def dead_letter(self, e): pass
    assert deliver_at_least_once(D(), {"id": "e1"}) == "delivered"
    assert calls["n"] == 3

def test_eeo_suppresses_duplicate():
    gate = ExactlyOnceGate(fakeredis.FakeStrictRedis())
    assert gate.first_delivery("evt-9") is True
    assert gate.first_delivery("evt-9") is False   # second time: duplicate

Operationally, confirm the contract with a black-box probe: send the same event twice and inspect the consumer.

# Send a duplicate; an EEO consumer must show exactly one applied side effect.
curl -s -X POST "$ENDPOINT" -H 'X-Idempotency-Key: evt-9' -d @event.json
curl -s -X POST "$ENDPOINT" -H 'X-Idempotency-Key: evt-9' -d @event.json

Failure modes and gotchas

Frequently Asked Questions

What guarantee am I actually offering while the dedup store is down?

At-least-once, and the only real question is whether you notice. Failing open when the SET NX call errors keeps events flowing but lets duplicates reach a consumer that may not expect them; failing closed holds the event until Redis answers, which is safe but stalls exactly the money-critical types first. Decide the fallback per event type and emit a counter for gate errors, so a silent tier downgrade shows up on a dashboard instead of in a support ticket.

How do I verify a third-party consumer really deduplicates before relying on at-least-once?

Send the same event twice with the same key against their sandbox and inspect the resulting state rather than their response codes — plenty of endpoints answer 200 twice and still write two rows. If there is no sandbox, ask which uniqueness constraint they key on; a vague assurance with no named column usually means a read-then-write that races under concurrent delivery. Where you get neither, treat the consumer as non-idempotent and put the gate on your side.

If I fix a bad payload and re-emit it under the same event ID, will the gate swallow it?

Yes, and that is the intended behaviour — the key claims the event, it is not a hash of the body. Corrections need a fresh event ID and ideally a distinct type, such as an amendment or reversal, so the consumer applies a compensating change rather than quietly overwriting history. Reusing an ID to republish a corrected payload is the most common way teams discover their deduplication layer works.

Does at-most-once mean skipping the delivery log as well?

No. Dropping retries is the guarantee decision; dropping the record of what you attempted just costs you observability. Keep a compact attempt row with the status code even for fire-and-forget types, because when someone asks why a metrics dashboard has a gap that log is the only answer available. What you can skip is the durable per-event state that exists solely to drive a future retry.

Do full-state snapshot events still need a guarantee tier?

They do, but the calculus shifts: applying the same snapshot twice is naturally a no-op, so at-least-once is comfortable even against a consumer that does nothing special. The risk moves from duplication to ordering, because an older snapshot arriving after a newer one on a retry rolls state backwards. Guard that with a version or sequence number the consumer compares before writing, rather than escalating to a dedup store you do not need.

Can the effectively-exactly-once gate lose an event if the worker dies right after claiming the key?

It can, and that is the sharp edge of claiming before dispatching: the key is set, the process dies before the POST, and every later attempt sees the key and suppresses a delivery that never happened. Store a delivery record beside the key carrying its terminal state, and run a sweeper that re-drives claims which never reached one. The key then means someone owns this event, not this event was delivered.