System Design Patterns: From Fundamentals to Real Systems
Vote

0% completed

​

Publish-Subscribe

  1. The Incident
  1. The Obvious Fixes, and Why They Fail
  1. The Pattern
  1. Walkthrough with Numbers
  1. Trade-offs
  1. When Not to Use It
  1. Choreography vs. Orchestration
  1. Real-World Examples
  1. In the Interview
  1. Check Yourself
  1. Related Patterns
  1. TL;DR

1. The Incident

The message queue solved the flash-sale failure: checkout now queues fulfillment, confirms each order immediately, and produces data that other teams want to consume.

First, the email team requests order events for receipts. Analytics, fraud, and loyalty follow. Each request adds another "send to this team's queue" call inside checkout. After eighteen months, checkout writes sequentially to five queues, creating three persistent problems:

1. Every new consumer means a checkout deploy. Checkout is the company's most important revenue service, yet another team's data request now requires changing and redeploying it. Its release schedule limits every team that needs order data.

2. Crashes leave teams with different data. Last Tuesday, checkout stopped after writing to three of five queues. Fulfillment received and shipped the order, but fraud never saw it. Five separate queue writes cannot succeed or fail as one operation, so a crash can leave any subset of teams without an event.

3. Other teams' problems become checkout's problems. When loyalty's queue slowed for an hour, checkout's five sequential writes also slowed. Checkout latency now depends on a queue owned by a different team.

The data science team then requests historical order events for model training, but must open a checkout pull request and wait for its next release. The producer must know about and deliver to every consumer, so each new consumer adds engineering work, risk, and coordination that cannot grow safely across teams.

2. The Obvious Fixes, and Why They Fail

"It works. Keep adding queues." The mechanism continues to function, but checkout owns every connection, retry policy, and delivery guarantee. Every new write also increases the chance that a crash leaves consumers with inconsistent data.

"Let teams read the orders database directly." Five teams then depend on the exact table structure, so a future schema migration can silently break dashboards and models the checkout team does not know exist. Repeated polling, which means continually asking whether data changed, also loads the production database. Change data capture offers a disciplined alternative by turning the database change log into an event stream. Direct, ungoverned table access provides none of its safety guarantees.

"One shared queue that everyone reads." A queue sends each message to exactly one worker. If fraud, email, and analytics share that queue, each service receives only a random portion of orders. The broker reports no error, so every team quietly operates with incomplete data. The self-check returns to this failure.

These teams do not need five separate jobs. They need independent access to one fact: an order was placed. The producer should publish that fact once for every interested consumer instead of delivering five separate copies itself.

3. The Pattern

The producer publishes a fact to a named channel, called a topic, exactly once. It does not know who is listening. The broker then delivers a separate copy to every subscription. Each subscriber processes its copy at its own pace, with its own failure handling.

Publish Subscribe Pattern
Publish Subscribe Pattern

The pattern contains four parts.

1. The topic is a named stream of facts such as order.placed. Checkout publishes once to that location and then continues. The broker performs one durable write, meaning it stores the event on disk before confirming, regardless of how many teams listen. Because checkout writes only once, a crash cannot leave different destinations with different results.

2. A subscription is one team's independent copy of the stream. Each subscription behaves like that team's own message queue, with its own backlog, stream position, acknowledgments, redeliveries, and dead letter queue. If fraud remains unavailable for an hour, only its backlog grows. Email and checkout continue normally.

3. Joining is configuration, not code. Data science creates a subscription without changing or deploying checkout. Brokers that retain history, including Kafka, also allow a new subscriber to replay earlier events and train a model on months of data before switching to live traffic.

4. The event's format is now a public contract. Unknown consumers treat event fields as an API. Events should state facts such as "an order was placed, here are its details," rather than commands such as "send a receipt." The publisher reports the fact, and each consumer chooses its response. Changes should also remain additive: new optional fields are safe, while renamed or removed fields can break unseen subscribers. Teams commonly use a schema registry, which checks event formats before release.

The central idea is simple: a producer states what happened, and every consumer determines what the fact means. Checkout does not need to know that receipts exist.

4. Walkthrough with Numbers

Assume the topic receives 200 orders per second, then compare the old and new designs.

Publishing cost. The old design makes five sequential queue writes that take about 25 ms, and each new team increases this cost. The new design makes one durable write in about 5 ms, independent of subscriber count.

The data science request. Data science creates its subscription in one afternoon without changing checkout. The topic retains 30 days of history, allowing the team to replay roughly 500 million past events for training before consuming live traffic. Previously, the same request required a pull request and release cycle but provided no history.

The failure drill. Fraud remains unavailable for 40 minutes, so its subscription accumulates 200 x 2,400 seconds = 480,000 events. The service processes this backlog after recovery, while email sends every receipt on time and checkout remains unaffected. In the earlier design, fraud's outage either slowed checkout or caused 40 minutes of missing events.

The audit that now passes. Every subscription can report whether it processed order X, still has it in backlog, or moved it to a dead letter queue.

5. Trade-offs

You gain:

  • The producer's cost and risk stay constant no matter how many consumers exist.
  • Teams ship on their own schedules. This organizational decoupling is the biggest benefit in practice.
  • Failures are isolated per subscription. With history retention, new consumers can start with the full past.

You pay:

Your event format is a public API now. A renamed field can break unknown consumers. Deploy-time coupling is replaced by contract discipline through schema registries, additive-only changes, and deprecation windows. A deprecation window is a defined period during which the old version continues to work.

Duplicates, multiplied. Each subscription independently provides at-least-once delivery, meaning it may deliver an event more than once but never zero times. Five subscribers create five separate places that require idempotency. Idempotent processing produces the same result whether an event is handled once or twice.

Invisible coupling. Teams may lose track of which services consume a topic. The dependency still exists, although it is no longer visible in producer code, so mature organizations maintain a subscriber catalog.

No "all done" signal. Each subscription knows only its own progress, so pub/sub cannot confirm that every required action for order X has finished. Workflows that require this answer need explicit coordination.

This limitation affects user-visible completion because publishing one event cannot prove that an order is ready. When inventory, payment, and shipping must finish before success, a coordinator tracks those steps, while pub/sub handles independent reactions after each one.

Slow consumers can lose data. Brokers retain history for a limited retention window. When a subscriber falls farther behind than that window, it permanently loses events rather than merely processing them late. Monitor the age of every subscription backlog.

Set that alert well before retention expires so the team has time to recover, add capacity, or rebuild the subscription from another authoritative copy before the broker permanently deletes history that the consumer has not processed.

6. When Not to Use It

1. It is a job, not a fact. "Resize this image" describes one task that one worker should perform, so use a message queue.

2. The producer needs the result. Checkout cannot continue until it knows whether payment was authorized. Publishing payment.requested without coordinating the result does not solve that dependency. Use request-response or an orchestrated workflow when the next step requires the outcome.

3. Two services, low traffic. One producer and one consumer usually need only a direct queue, which is simpler to operate and understand.

4. Request-response rebuilt from events. Pairing user.lookup.requested with user.lookup.completed recreates a synchronous call with greater latency and harder debugging. An event name containing "request" often indicates that a direct call is more suitable.

At organizational scale, unmanaged pub/sub creates topic sprawl: hundreds of topics without owners, schema rules, or visible consumers. Removing mandatory coordination between teams requires replacing it with explicit contracts and a reliable catalog.

The catalog should name the topic owner, the schema, the retention period, the data classification, and every subscription owner. With that written down, a schema change, an incident, or a deletion request can be handled without searching through logs. It also identifies abandoned consumers before their backlogs exceed retention.

7. Choreography vs. Orchestration

Pub/sub supports choreography, in which services react independently to events and their combined behavior forms the overall process. Orchestration uses one coordinator to execute a workflow step by step. These approaches suit different dependencies.

Choreography (pub/sub)Orchestration (a coordinator)
Who is in controlNobody central; behavior emergesThe coordinator, explicitly
Adding a stepNew subscription, no coordinationUpdate the coordinator's workflow
"Is order X fully processed?"No single place knowsThe coordinator knows, by design
Failure handlingEach consumer handles its ownThe coordinator retries, undoes, escalates
Best forIndependent reactions: analytics, email, indexingSequences where each step depends on the last: inventory, payment, shipping

Choose based on whether the steps are truly independent. Email and analytics do not depend on their relative order, so choreography works well. However, reserving inventory, charging payment, shipping, and undoing earlier work after a failure form a dependent workflow. The saga lesson covers orchestration for that case. Many systems orchestrate the order workflow and publish facts from each stage for independent consumers.

8. Real-World Examples

  • AWS SNS feeding SQS queues is this pattern built from managed services. SNS is the topic, and each team's SQS queue is its subscription. Broadcast first, then each team gets its own task list.
  • Kafka treats each consumer group as an independent subscription over a stored log. That is why it dominates this space: broadcast, replay, and per-team queue behavior in one system.
  • Google Cloud Pub/Sub is the same model, fully managed. Subscriptions are objects you create and manage directly.
  • Redis Pub/Sub is the cautionary example. It is fire-and-forget: if a subscriber is offline, the message is gone permanently. "Pub/sub" without durable subscriptions is a much weaker promise with the same name. Always check which one you are being offered.

For AI engineers: an interaction-event topic often provides the central stream for an ML platform. One user.interacted stream can feed an analytics warehouse, a feature store, an embedding refresh job, and a recommender's real-time features. The same replayable history supports future models. Agent platforms increasingly publish every tool call and decision once, allowing observability, evaluation, and audit systems to consume them independently without delaying the agent.

9. In the Interview

An interviewer may ask directly how five teams should receive order data, or indirectly ask what happens when a sixth service needs the same information. Adding another producer arrow each time indicates growing coupling. A different question may test whether you incorrectly choreograph a dependent order, payment, and shipping workflow that needs orchestration.

The 30-second answer: "Checkout publishes an order.placed fact to a topic, once. Every interested team gets its own subscription with its own backlog, retries, and dead letter queue. So a slow or crashed consumer affects nobody else. Adding a consumer is configuration, not a checkout deploy. With a log-based broker like Kafka, the new consumer can replay history. The costs I manage: the event schema is now a public API, so changes are additive-only and checked against a registry. Every subscription needs idempotent processing. And I alert on backlog age so a slow consumer never falls past the retention window. If the steps are really a workflow that needs completion tracking, I orchestrate that part instead."

Likely follow-ups:

  1. "One subscriber processes at half the speed of the others. What happens?" Only that subscription's backlog grows, while other consumers remain unaffected. If the backlog exceeds the retention window, events are permanently lost. Monitor backlog age and scale the slow subscription with competing consumers, applying the queue pattern inside that one subscription.
  2. "How do you change an event's schema without breaking consumers you don't know about?" Add optional fields without renaming or removing existing ones, and use a schema registry to reject breaking builds. A genuinely breaking change requires publishing version 2 alongside version 1 for a defined deprecation window. Maintain a consumer catalog so dependencies are known.

A mistake that fails candidates: connecting several services to one shared queue and describing it as broadcast, even though each service receives only a sample. Another mistake is rebuilding request-response with paired events. Both indicate confusion between jobs and facts.

10. Check Yourself

Q1 (recall). Explain the difference between a topic and a subscription. Then state exactly what it costs the producer when a sixth consumer team is added.

Q2 (trade-off). "Pub/sub decouples services." Argue the opposite for two minutes: what coupling survives, and what new discipline replaces the old deploy-time coordination?

Q3 (scenario). Fraud reports scoring only about 20% of orders. Email reports sending about 30% of receipts. The broker's metrics show nothing was lost, and checkout's publish count matches order volume exactly. Diagnose the setup and fix it.

<details> <summary>Answers</summary>

A1. A topic stores one durable stream of published facts. A subscription gives one consumer an independent feed with its own position, backlog, acknowledgments, and dead letter queue. Adding a sixth consumer costs the producer no code, deployment, latency, or awareness. The new consumer must provide idempotent processing and monitor its backlog, while the overall system gains another dependency on the event schema.

A2. Consumers remain coupled to the event schema, so changing one field can break systems the producer cannot identify. Pub/sub therefore moves coupling from the release schedule into a data contract. A schema registry, additive changes, deprecation windows, a consumer catalog, and a clear owner for each subscription must govern that contract.

A3. The three services share one subscription or queue, which makes them competing consumers. The broker sends each event to exactly one service, so each sees a random subset based on worker count and speed. Create a separate subscription for fraud, email, and fulfillment so every service receives every event. Use competing consumers only among workers inside one service. The rule in one line: broadcast across services, compete within a service.

</details>
  • Message Queue: each subscription effectively is one. The broadcast-into-queues setup (SNS into SQS) composes the two patterns directly.
  • Idempotency: every subscription is at-least-once on its own, so N consumers means N chances to double-apply.
  • Event-Driven Architecture: the next lesson. Pub/sub is the transport; that lesson decides what should travel over it.
  • Change Data Capture: turns a database's own change log into the event stream, the disciplined version of "just read my tables."
  • Webhooks: the same idea applied across company boundaries, over plain HTTP, with the guarantees renegotiated.

12. TL;DR

ProblemWhen the producer must personally deliver to every consumer, each new team means producer deploys, slower fan-out, and partial crashes where some teams get the event and others do not.
MechanismPublish each fact once to a topic; the broker gives every subscription its own full copy, with its own backlog, retries, and dead letter queue. Subscribing is configuration, and log-based brokers add replay.
CostsThe event schema becomes a public API (additive-only, registry-enforced), duplicates per subscription, coupling that goes invisible without a catalog, no completion signal, and retention as a hard limit for slow consumers.
Skip it whenIt is a job for one worker (queue), the producer needs the result (request-response or orchestration), or the "events" are really a step-by-step workflow (saga).
30-second answerPublish facts once to a topic; every subscriber consumes independently, at its own pace, with its own failure handling. Adding consumers costs the producer nothing, and history replays for newcomers. Govern the schema like the API it now is, deduplicate per subscription, and orchestrate anything that is really a workflow.
Flashcards Review

What is the publish-subscribe pattern?

1 / 19
General
Test Your Knowledge
Check your understanding and reinforce the key concepts covered in this section with a short, targeted assessment.
9 Questions
~14 mins
Your progress is saved automatically

Reading Progress

0%


Vote for new content

On This Page

  1. The Incident
  1. The Obvious Fixes, and Why They Fail
  1. The Pattern
  1. Walkthrough with Numbers
  1. Trade-offs
  1. When Not to Use It
  1. Choreography vs. Orchestration
  1. Real-World Examples
  1. In the Interview
  1. Check Yourself
  1. Related Patterns
  1. TL;DR