Every webhook integration eventually receives the same event twice. Usually it happens during an incident, when retries are piling up and nobody is looking closely – and the consequence is a customer charged twice, an order shipped twice, or a record overwritten with stale data.
Idempotent webhooks are the fix: handlers built so that processing the same event once, twice or ten times produces the same result. This article explains why duplicates and out-of-order events are unavoidable, the patterns that make handlers safe, and what to get right when you are the one sending webhooks.
Why duplicates are guaranteed
Webhook providers almost universally promise at-least-once delivery, not exactly-once. That is a deliberate trade-off. A provider that cannot confirm you received an event – because your endpoint timed out, returned an error, or the network dropped the response – will send it again rather than risk you never receiving it.
So duplicates come from ordinary conditions:
- Your handler took too long, the provider timed out, and it retried an event you had already processed.
- Your endpoint returned an error after doing part of the work.
- A deployment restarted your service mid-request.
- The provider had its own incident and re-sent a batch of events.
- Someone replayed events manually after an outage.
Ordering is not guaranteed either. Events are delivered by distributed systems with retries, so an “updated” event can arrive before the “created” event, or an older state can arrive after a newer one.
The core pattern
A safe webhook handler does five things, in this order:
- Verify authenticity. Check the signature or other authentication the provider offers before trusting the payload.
- Record the event ID. Store the provider’s event ID in a table with a unique constraint. If the insert fails because the ID already exists, the event is a duplicate – acknowledge it and stop.
- Acknowledge quickly. Return success as soon as the event is safely recorded. Do the real work asynchronously from a queue. Slow handlers cause timeouts, and timeouts cause retries.
- Process idempotently. Make the business operation itself safe to repeat – for example, “set order status to paid” rather than “add a payment”, or use the event ID as a key on any record the work creates.
- Track the outcome. Mark the event processed, failed or parked, so failures can be retried and investigated.
The unique constraint in step 2 matters more than any application logic. Two copies of the same event can arrive at the same moment on two servers, and only the database can reliably decide which one wins.
Handling out-of-order events
There are two reliable approaches, and many integrations use both:
- Fetch current state. Treat the webhook as a notification that something changed, then call the provider’s API for the object’s current state and apply that. Order no longer matters, at the cost of an extra API call.
- Compare versions or timestamps. If events carry a version number or reliable timestamp, store the latest one applied for each object and ignore anything older.
What does not work is assuming events arrive in the order they happened.
Queues, retries and dead letters
Once events are recorded and acknowledged, processing should run from a queue with:
- Retries with backoff for transient failures such as a downstream API being briefly unavailable.
- A dead-letter queue for events that keep failing, so they are parked for a person rather than retried forever or silently dropped.
- Alerting on dead-letter volume and processing lag.
- Replay – the ability to re-run events for a time window after an outage or a fix. Because handlers are idempotent, replay is safe.
Reconcile anyway
Even a well-built handler can miss events if the provider never sends them, or if your endpoint was unreachable for longer than the provider’s retry window. Schedule a periodic reconciliation that lists recent objects from the provider’s API and compares them with your records. It catches what webhooks missed, and it proves the integration is working rather than assuming it.
When you are the one sending webhooks
If your platform sends webhooks to partners or customers, the same principles apply from the other side:
- Give every event a unique, stable ID and include it in every retry.
- Sign payloads and publish how to verify them.
- Retry with backoff over a documented period, and say so.
- Include a version or timestamp so consumers can handle ordering.
- Offer an API to fetch current state and an event history consumers can replay from.
- Document the delivery guarantees honestly: at-least-once, no ordering guarantee.
Consumers will build better integrations against a provider that tells them the truth about delivery.
Illustrative scenario: duplicate fulfillments under load
This scenario is a composite. It is representative of the webhook remediation work we do, but it does not describe a specific client, and nothing in it is a reported result.
A retailer’s order service receives webhooks from its commerce platform and a payment provider. During a promotion, handler latency rises, the providers retry, and the warehouse receives some orders twice. The handler processes everything inline and has no record of which events it has seen.
The remediation adds an event table with a unique constraint on provider and event ID, moves processing to a queue, and changes fulfillment creation to use the order ID as an idempotency key. Status updates fetch the current order state from the platform before applying it. A nightly job compares the day’s orders on the platform with the order service and lists anything missing or duplicated. The existing duplicates are identified by replaying the event history against the new logic.
What a good outcome looks like: the next promotion produces retries but no duplicate shipments, and the team learns about webhook failures from an alert rather than from the warehouse.
Related reading
- Stripe integration beyond checkout
- Integration reconciliation reports: proving two systems agree
- Shopify ERP integration at peak
Working with eProxim
Idempotent processing, retry with dead-letter handling, audit trails and replay are defaults in the integrations eProxim builds, not upgrades. We have built event-driven systems where duplicates have real cost – including a telecom usage pipeline designed so replayed records never double-charged – and we fix webhook implementations that lose or double-process events.
Have webhooks that occasionally do things twice? See our enterprise integration services, or start a partner conversation.