0% completed
Webhooks
On This Page
- The Incident
- The Obvious Fixes, and Why They Fail
- The Pattern
- Walkthrough with Numbers
- Trade-offs
- When Not to Use It
- Webhooks vs. Polling
- Real-World Examples
- In the Interview
- Check Yourself
- Related Patterns
- TL;DR
1. The Incident
Your payment provider reports successful charges through a webhook. For each payment, it sends an HTTP POST to https://api.yourstore.com/webhooks/payments. The handler receives the event, marks the order paid, reserves inventory, notifies the warehouse, sends a receipt, and finally returns 200. Under heavy traffic, this work takes about 7.5 seconds.
The provider waits 5 seconds for a response and then closes the connection.
From the provider's perspective, 40% of deliveries fail because they exceed that deadline, so it retries them. Your handler usually completes after the provider stops waiting, and the retry repeats every action, causing two shipments for deliveries whose processing actually succeeded.
The slow handler continues increasing the measured failure rate until it crosses the provider's health threshold. According to its documented policy, the provider then disables your endpoint.
The resulting failure is silent because no requests arrive and your logs contain no new errors. Customers continue paying for three hours while their orders remain "pending," and support tickets reveal the problem before monitoring does.
Unlike earlier examples, the other side of this connection belongs to another company. A webhook applies pub/sub across that boundary, where the sender controls its timeout, retry schedule, and endpoint-disable policy. Your system must operate correctly under those fixed rules.
2. The Obvious Fixes, and Why They Fail
"Make the handler faster." Reducing processing to 4 seconds satisfies today's limit, but one slow warehouse call, database problem, or traffic spike exceeds it again. The same sequence of delay, retry, duplicate work, and endpoint disablement remains possible because the structure has not changed.
"Ask the provider for a longer timeout, or fewer retries." The provider's short timeout protects its service from your delays, as the timeout lesson explains. Its retries prevent brief failures in your system from losing financial events. Removing either protection transfers more risk to both companies.
"Forget webhooks. We'll poll their API instead." Polling every minute wastes rate-limit capacity, introduces up to one minute of delay, and usually discovers no changes. However, polling does guarantee eventual completeness, while webhooks provide speed. A reliable design uses both benefits instead of choosing only one.
3. The Pattern
A webhook is an event delivered as an HTTP POST from another organization. The receiving rule: check it is authentic, save it, respond "received" within milliseconds, and do the actual work later, behind your own queue.
Apply five rules in the order that each request encounters them.
- Verify before you believe. The webhook URL is public, so anyone can send
{"type": "payment.succeeded"}to it. Providers attach an HMAC, which is a hash computed with a secret key shared only by both companies. Recompute that HMAC and compare it before accepting the event. Also check the provider's timestamp to prevent replay of an old, valid message. Without these checks, an attacker can mark unpaid orders as paid. - Save it and acknowledge. Nothing else. Store the raw event, place it on your message queue, and return 200 in about 15 ms. 200 means "safely received," never "fully processed." Inventory, warehouse notification, and receipts run later behind your queue, using your retry and failure policies. The provider's 5-second deadline no longer affects business processing.
- Deduplicate on the event ID. Providers document at-least-once delivery because they resend whenever success is uncertain. Use each unique event ID as an idempotency key so repeated processing has only one effect. This rule prevents duplicate shipments.
- Do not trust the order, or the payload's freshness. A "refund updated" event may arrive before its related "charge succeeded" event. For state-dependent decisions, treat the webhook as notice of a change and fetch current state from the provider's API. That API is the source of truth, or the authoritative copy, while a retried payload may be hours old. Stripe recommends this approach.
- Reconcile, because silence is not success. Events may never arrive because an endpoint was disabled, an outage exceeded the retry window, or a network path failed silently. Periodically request everything changed after a stored marker, often called a watermark. The webhook tells you quickly. The sweep tells you about everything, including the events that never arrived.
If you are ever the sender (your product delivering webhooks to your customers), these rules become your responsibilities. Sign payloads, provide unique event IDs, publish the retry schedule, keep delivery timeouts short, and disable endpoints that fail repeatedly. A dashboard should expose every delivery and provide a "redeliver" action.
The sender must keep enough delivery history for customer debugging by recording the event ID, destination, attempt number, response code, response time, and next retry without exposing secrets. Customers can then distinguish an event that was never attempted from one their server rejected.
4. Walkthrough with Numbers
Rebuild the incident using these five rules.
The handler now takes about 15 ms. Signature verification uses about 1 ms, while raw-event storage and queue insertion use about 10 ms. The handler then returns 200. Because no slow dependency remains in the request, the provider observes a 100% success rate through load spikes and deployments.
The real work proceeds through your queue as described in the queue lesson. Spikes create backlog, worker crashes cause redelivery, and events that repeatedly fail move to the dead letter queue.
A duplicate delivery arrives 90 seconds later because the first 200 response was lost in transit. The worker finds the event ID in its processed set and stops, producing one shipment. This applies the idempotency lesson across a company boundary.
The hourly reconciliation job asks for changes after its bookmark and usually finds none. During a provider delivery incident, it discovers 217 missing confirmations and places them on the same queue without requiring an emergency response.
One new alarm exists for webhook silence. If an endpoint that normally receives steady traffic becomes quiet for 15 minutes, the system alerts an operator. This detects failures that produce no local errors.
5. Trade-offs
You gain:
- Near-real-time updates from other companies, without constant polling.
- The provider's retry machinery working for you, now that your handler responds quickly.
- One consistent entry point for external events, feeding queue-and-worker machinery you already know how to run.
You pay:
You inherit their rules. Each provider defines its own at-least-once delivery, ordering, retry schedule, disable policy, and payload versioning. Every integration therefore has a separate contract to understand.
A public endpoint to defend. Signature validation, timestamp windows, and secret rotation become part of your security boundary. Anyone with a leaked signing secret can forge "payment succeeded" events.
Debugging across a boundary. Determining whether the sender failed to send or your service dropped an event requires the provider's dashboard, your raw-event records, and comparable timestamps. Saving the raw request makes this investigation possible.
Use one correlation record for every event ID across receipt, queue processing, business effects, and reconciliation, allowing an operator to determine whether an event arrived more than once and whether downstream work succeeded. Without this record, each team sees only one part of the transaction.
The polling loop never goes away. Reconciliation provides completeness, while webhooks provide timely notification. Production systems operate both mechanisms permanently.
Store the reconciliation bookmark only after recovered events enter durable processing and record the covered time range for complete later audits. Advancing it earlier can create another silent gap if the job stops between reading the provider's response and saving its events. Replaying an earlier bookmark is safe when event IDs provide reliable deduplication across both webhook and polling paths, even after a partial reconciliation failure.
Monitor reconciliation lag as well.
6. When Not to Use It
Between your own services. A webhook recreates broker behavior through HTTP but lacks durable subscriptions, replay, and managed redelivery. When you control both services, use pub/sub with stronger guarantees.
When completeness matters more than speed. Billing reconciliation, financial close, and compliance exports should begin with polling or bulk export APIs. Webhooks can add freshness without becoming the only data path.
High-frequency streams. Thousands of individual HTTP POSTs per second are inefficient. Use the provider's streaming or bulk interface when available.
When someone is waiting for the answer. An interactive question needs request-response. Adding a webhook round trip increases latency and failure modes.
The incident began by treating the webhook request as the place to complete business work. It should only accept and store the notification. Timeouts, duplicates, endpoint disablement, and silent gaps all become harder when that distinction is ignored.
Never let a provider choose your internal database transaction through an arbitrary payload: validate the event type and object before enqueueing, then authorize state transitions in the worker. A valid signature proves who sent the message, not that every requested change is appropriate for your system.
7. Webhooks vs. Polling
These mechanisms complement each other in production:
| Webhooks (push) | Polling (pull) | |
|---|---|---|
| Freshness | Seconds | However often you poll |
| Cost when nothing is happening | Nearly zero | Constant "anything new?" requests |
| Completeness | Best effort: gaps happen | Guaranteed, eventually |
| Who controls delivery | The sender | You |
| Worst failure looks like | A quiet day | Stale data (visible, at least) |
Polling costs requests and introduces delay. Webhooks have a more subtle weakness because a delivery failure can look identical to no new activity. Use webhooks for speed, polling for truth, with a bookmark that lets polling recover anything push delivery misses.
For continuous streams to browsers or applications you control, persistent connections replace separate POST requests. Server-Sent Events, covered next, provides one form of this approach. Webhooks specifically address server-to-server push across company boundaries.
8. Real-World Examples
- Stripe is the reference implementation of everything above. It has signed payloads with timestamp checks, event IDs for dedupe, retries spread over days, and endpoint auto-disable. Its official advice is to fetch current state instead of trusting payload freshness.
- GitHub signs deliveries with
X-Hub-Signature-256and gives you a delivery log with a redeliver button: the sender-side obligations, done well. - Slack requires a 200 within 3 seconds on event deliveries. Nobody's business logic fits in 3 seconds, and that is the point: the contract forces accept-first, work-later.
- Twilio, Shopify, PayPal: every integration-heavy platform converges on the same contract terms, because the failure modes converge.
For AI engineers: asynchronous AI services often use webhooks to report completion. Batch inference and long fine-tuning jobs can notify clients this way. The receiver still verifies, stores, acknowledges quickly, and processes behind a queue. Agent products also accept webhook triggers from email, calendars, and CRMs. Deduplication is especially important because a repeated event can start an expensive agent run twice.
9. In the Interview
An interviewer may ask directly about Stripe integration or indirectly ask how payment settlement reaches your system. Expect follow-ups about work performed before returning 200, duplicate handling, and detection of events that never arrive. The completeness question is often the most revealing.
The 30-second answer: "My webhook handler does three things only: verify the signature and timestamp, save the raw event, and enqueue it. It returns 200 in about 15 milliseconds, because 200 means received, not processed. Workers do the real work behind my own queue. They deduplicate on the event ID, since delivery is at-least-once. For state-based decisions they fetch current state from the provider's API instead of trusting a possibly stale payload. Then two protections. An hourly reconciliation poll from a bookmark, because webhooks are best-effort and the worst failure is silence. And an alarm if a normally busy endpoint goes quiet. Speed from the push, truth from the poll."
Likely follow-ups:
"How do you know you missed one?" The webhook path alone cannot identify an absent event. Reconciliation eventually recovers every gap, while a silence alarm detects large gaps quickly. A lack of traffic must be treated as a monitored state.
"Secure the endpoint." Verify the HMAC over the raw body with the shared secret, using a constant-time comparison whose duration does not reveal how closely values match. Check the timestamp to prevent replay, rotate the secret safely, and reject invalid requests before parsing them. IP allowlists can provide an additional layer, but signatures provide authentication.
A mistake that fails candidates: completing business processing before responding or failing to address duplicate delivery. Either design works only while every dependency behaves normally.
10. Check Yourself
Q1 (recall). List the receiver's five rules in the order a request meets them, and state precisely what the 200 response means and does not mean.
Q2 (trade-off). The provider's payload already contains everything your worker needs. This lesson still recommends fetching current state from their API before acting. Defend the extra call, and name the cases where you would skip it.
Q3 (scenario). Payment confirmations have quietly stopped. Logs show no handler errors for three weeks. The provider's dashboard shows your endpoint was auto-disabled 20 days ago, after a failure spike during a deploy. Reconstruct the chain of events, then redesign so this class of failure cannot happen silently again.
<details> <summary>Answers</summary>A1. Verify the signature and timestamp, save the raw event durably, acknowledge with 200 in milliseconds, process later through a queue with event-ID deduplication, and reconcile from a bookmark. The 200 response means the event is safely in your possession, not that its business effects are complete.
A2. A payload reflects the time its delivery attempt was created and may become stale after retries or arrive out of order. The provider API contains current authoritative state, so fetch before making state-dependent decisions. Skip the call for self-contained append-only events, when volume would exceed the API limit, or when the action is idempotent and inexpensive to correct.
A3. A deployment caused slow responses or errors, retries encountered the same problem, and the measured failure rate triggered automatic disablement. Delivery then stopped without producing local errors. A 15 ms save-and-acknowledge handler removes slow business work from the request. A silence alarm identifies future disablement within minutes, hourly reconciliation bounds any data gap, and deployment checks verify endpoint health. Completeness should not depend entirely on another company's retry window.
</details>11. Related Patterns
- Message Queue: the separation between accepting the news and doing the work; the handler's only job is feeding it.
- Idempotency: event-ID dedupe is the cross-company version of idempotency keys; at-least-once is written into every provider's contract.
- Publish-Subscribe: what a webhook is, structurally, once carried over HTTP across a company boundary, with the broker's guarantees renegotiated as contract terms.
- Retry with Exponential Backoff: running on the sender's side, both for you and against you; worth understanding from both sides.
12. TL;DR
| Problem | External events arrive as HTTP POSTs governed by the sender's timeout, retries, and disable policy; do the work inline and you get duplicates, a disabled endpoint, and a silent three-hour gap discovered by support tickets. |
| Mechanism | Verify the signature and timestamp, save the raw event, return 200 in milliseconds, process behind your own queue with event-ID dedupe, fetch current state at decision time, and reconcile with a bookmark-based poll. |
| Costs | You inherit the sender's delivery rules, defend a public endpoint, debug across a company boundary, and run the polling check permanently, because webhooks provide speed, not completeness. |
| Skip it when | Both ends are yours (use pub/sub), completeness is the requirement (poll first), or the volume calls for a stream instead of POSTs. |
| 30-second answer | Treat each webhook as a signed, at-least-once notification: save and acknowledge in milliseconds, work behind your own queue with dedupe, fetch truth from the API before acting, and back it all with a reconciliation poll plus a silence alarm, because the worst webhook failure looks exactly like a quiet day. |
Flashcards Review
What is a webhook?
Reading Progress
0%
On This Page
- The Incident
- The Obvious Fixes, and Why They Fail
- The Pattern
- Walkthrough with Numbers
- Trade-offs
- When Not to Use It
- Webhooks vs. Polling
- Real-World Examples
- In the Interview
- Check Yourself
- Related Patterns
- TL;DR