0% completed
Event-Driven Architecture
On This Page
- The Incident
- The Obvious Fixes, and Why They Fail
- The Pattern
How much data the event carries
- Walkthrough with Numbers
- Trade-offs
- When Not to Use It
- Notification Events vs. Full-State Events
- Real-World Examples
- In the Interview
- Check Yourself
- Related Patterns
- TL;DR
1. The Incident
The pub/sub migration succeeded: checkout publishes order.placed once, five teams subscribe independently, and checkout deployments are stable again.
However, each event contains only {event: "order.placed", order_id: "ord_7f3a"}. This notification tells consumers that an order exists but requires them to request its details.
Fraud receives the event and calls GET /orders/ord_7f3a for the amount and items. Email, analytics, loyalty, and fulfillment call the same endpoint, with fulfillment accidentally calling twice. As a result, every event you publish comes back to you as five API calls.
During the next flash sale, 5,000 orders per second produce 5,000 events per second. Consumers then generate 25,000 GET requests per second hitting checkout's API, precisely during the traffic spike the queue was intended to absorb.
Checkout's read service fails, consumer calls time out, and retries increase the load further, while growing backlogs delay fraud scoring by several minutes during the year's highest-fraud period.
Although communication became asynchronous, the original load returned. Events reported that something happened without describing what happened, so every consumer still depended on checkout for details. You added asynchronous delivery on top of a synchronous dependency. That is not an event-driven architecture yet.
2. The Obvious Fixes, and Why They Fail
"Cache the lookups." A cache before checkout can reduce some traffic, but every consumer requests each new order immediately. Because that order did not exist one second earlier, all five consumers miss the cache together and call the API. They also continue depending on checkout's availability.
"Slow the consumers down." Rate limiting protects checkout by making every downstream team process at checkout's allowed speed. Backlogs then grow during the most important bursts, and fraud detection depends on checkout's capacity. The queue delays the dependency without removing it.
"Could we just put the details in the event?" Yes, and that is the pattern. Copying data introduces genuine costs, which the self-check examines. However, the first two proposals only reduce the cost of callbacks, while complete events can make callbacks unnecessary.
3. The Pattern
In an event-driven architecture, services talk to each other through events. One decision decides whether it works: how much data each event carries. Carry enough, and a consumer can finish its job on its own. Carry too little, and it has to call you back for the rest, which is the storm described above.
How much data the event carries
The earlier lessons covered how an event gets delivered. This one is about what goes inside it. There are three levels.
Level 1: Event Notification "Order ord_7f3a was placed."
Consumers must call back for details.
|
Level 2: Event-Carried State "Order ord_7f3a: 3 items, $142,
Transfer customer c_91, ships to Austin."
Consumers act without calling anyone.
|
Level 3: Event Sourcing The events ARE the data.
Current state is rebuilt by replaying them.
Level 1, notification. The event says that something happened and leaves every consumer to ask for the details. Events remain small and data is not copied, but this level creates the callback traffic from the incident.
Level 2, event-carried state transfer. The event includes the facts consumers need. Each service stores a local view, which is a small table containing only relevant fields and updated from events. Fraud stores the order fields required for scoring and reads them locally in under a millisecond, even when checkout is unavailable. Each consumer now owns a copy it can delete and rebuild, instead of calling checkout for every order.
Level 3, event sourcing. The event log is the system of record, meaning the official data source, and tables are rebuilt from that log. This storage decision has its own lesson and is unnecessary for solving a callback storm. Level 2 is sufficient.
Most practical event-driven systems use Level 2 with three rules.
Events are facts, never commands. Publish that an order was placed rather than instructing a specific consumer to send a receipt. A command ties the publisher back to what one consumer is supposed to do, which is the coupling you were trying to remove.
Local views are copies, never the truth. Every view can be deleted and rebuilt by replaying topic history. Checkout remains the official data owner, while other services keep convenient copies.
Freshness becomes a "when," not an "if." A local view always follows the source by some delay, usually milliseconds but sometimes minutes during a backlog. This behavior is called eventual consistency: all copies become consistent after updates finish propagating. Products must account for that delay explicitly.
That delay should be visible and measurable by recording the source version or event timestamp in each local view and exposing its age through monitoring. A support screen can display an "as of" time, while an alert fires when a view exceeds its agreed freshness limit. Eventual consistency is manageable when the product and operators know exactly how far behind each copy is.
4. Walkthrough with Numbers
Consider the flash sale again, now using full-state events.
The event grows from about 100 bytes to about 2 KB. At 5,000 events per second, the broker handles 10 MB per second, which is modest for a production broker. In return, the system eliminates 25,000 GETs per second by exchanging a small amount of bandwidth for service independence.
Checkout's read traffic during the sale: flat. Consumers make no callbacks, so the API serves users instead of internal lookups.
Fraud reads locally in about 1 ms instead of making a 50 ms network call with timeout and retry logic. Its performance and availability no longer depend on checkout during the busiest hour.
A new consumer replays history. With 30 days of retention, the topic contains roughly 500 million events. A new service replays them in several hours to build its local view and then begins consuming live events. Without replay, the team would backfill through checkout's API and add significant load.
5. Trade-offs
You gain:
- Independence: consumers read locally, stay up when the producer is down, and never overload anyone with callbacks.
- Producer load that does not depend on how many consumers exist or how much data they need.
- Views that can be rebuilt from history, which makes recovery and onboarding the same operation: replay.
You pay:
Copies everywhere. Six services may store order data in six different shapes. Storage is inexpensive, but every team must remember that these copies are disposable and the owner's store remains authoritative.
Eventual consistency, always. Every view has some delay, so the product must handle cases such as a support dashboard showing data that is 40 seconds old.
Personal data spreads. Including a customer's email in a full-state event copies it into every subscriber and the retained topic. A deletion request must then clean N+1 locations, as the self-check demonstrates.
This privacy cost should influence the schema before publication by limiting it to widely needed facts, classifying sensitive fields, and recording every subscribing owner. Encryption protects stored data but does not remove the need to delete or restrict unnecessary copies, so a convenient event field can become a long-term compliance obligation.
The schema matters even more. Larger events expose more public fields. Schema registries, additive-only changes, and a consumer catalog become mandatory rather than optional.
Harder debugging. Each service has its own timestamped answer for order X. Correlation IDs, which identify the same request across events and logs, and distributed tracing become essential.
Operators also need replay procedures that are safe to repeat. Rebuilding one view should not resend emails, charge payments, or trigger other external effects. Separate state-building consumers from side-effecting consumers, and test replay against production-sized history before relying on it during recovery.
6. When Not to Use It
Inside one service. Functions inside a monolith already share consistency and transactions. Event-driven architecture addresses boundaries between services and teams, so applying it without those boundaries adds unnecessary complexity.
When the read must be exactly current. Selling the last inventory unit or approving a withdrawal requires the latest value. Route these reads synchronously to the owning service and use local views for data that can tolerate a delay.
Request-response, rebuilt from events. Full-state events do not improve a workflow that remains a synchronous dependency. The warning from the pub/sub lesson still applies.
Before the team is ready. Without schema enforcement, tracing, and backlog alerts, event-driven systems are difficult to operate. A smaller number of services using direct calls can be the safer design until those capabilities exist.
Another misuse is introducing events to avoid assigning data ownership. If six services keep views but none owns the fact, their copies can diverge without any authoritative answer. Every fact needs one owning service, while events distribute that service's data.
Ownership also determines how corrections work: a consumer should not repair an order field locally and publish a competing truth. It sends a command to the owning service, which validates the change and publishes a new fact that makes every view converge through the same ordered path.
7. Notification Events vs. Full-State Events
The two practical levels have different strengths:
| Notification (Level 1) | State transfer (Level 2) | |
|---|---|---|
| The event contains | Just an ID: "go look it up" | The facts themselves |
| When a consumer needs data | Calls the producer | Reads its own local view |
| Producer's read load | Grows with consumer demand | Zero |
| Copies of data / privacy exposure | Minimal | Every subscriber, plus retained history |
| Freshness | The lookup is always current | Views lag slightly behind |
| Best when | Few consumers need details; data is sensitive | Many consumers need the same data; independence matters |
A practical middle option includes commonly used fields in the event and an ID for uncommon detailed lookups. Mature systems often decide field by field by asking who needs the value and what widespread copies would cost. Frequently used fields belong in the event, while sensitive fields remain with their owner.
8. Real-World Examples
- LinkedIn built Kafka for exactly this: activity events as one large stream that every team reads to maintain its own view. The "central nervous system" nickname for Kafka started there.
- Uber: a trip is a stream of events long before it is a database row. Pricing, ETA prediction, fraud, and driver payments each consume it independently.
- Bank ledgers are an early form of the pattern: transactions are immutable facts, and your balance is a view computed from them. Event-driven design predates software.
- Netflix runs viewing events through the same kind of central stream into recommendations, artwork personalization, and capacity planning.
For AI engineers: full-state events keep model features current without API callbacks. A feature store's online values update from interaction events, while a vector index re-embeds content from a content.updated event carrying the content. Training pipelines use retained history as a dataset, applying the same replay process used to initialize a microservice. This pattern lets a recommender observe a purchase within 200 ms without calling checkout.
9. In the Interview
An interviewer may ask how service B obtains service A's data without calling it. More often, after you draw a topic, they ask what the event contains. That decision determines callback traffic, consumer independence, and the amount of personal data copied through the system.
The 30-second answer: "I treat the event's contents as the main design decision. Thin notification events force every consumer to call the producer back, so the coupling and the load return. For widely consumed facts I put the state in the event. Each consumer maintains its own local view: updated from the stream, rebuildable by replay, readable in microseconds, and available even when the producer is down. The costs I manage: every view is eventually consistent, so reads that must be exactly current go synchronously to the owning service. Sensitive fields stay out of the event or get encrypted. The larger schema is a public API under registry control. Ownership stays with one service: views are copies, never the truth."
Likely follow-ups:
"A support agent updates an order and doesn't see the change on their own dashboard. Why, and what do you do?" The dashboard reads a local view that has not processed the write yet. The interface can show an "as of 12:04:31" timestamp, route the writer's next read to the source through read-your-writes behavior, or wait until the view reaches the write's version. Choose the appropriate guarantee for each screen.
"A user invokes their right to be deleted. Their data sits in full-state events across a 30-day retained topic and six local views. Go." Delete the record from the owning store and publish a tombstone, which is a special event instructing every view to purge that user. If the retained history must be removed before normal expiration, use crypto-shredding by encrypting personal fields with a per-user key and destroying that key. Every retained copy, replay, and backup then becomes unreadable.
A mistake that fails candidates: describing notification events as decoupled while every consumer still calls the producer. The opposite error is using full-state events without explaining staleness, deletion, or ownership. Both ignore the architectural effect of event contents.
10. Check Yourself
Q1 (recall). Name the three levels of event design, and for each, say in one line where the truth lives.
Q2 (trade-off). The fraud team asks you to add customer.lifetime_value to the full order.placed event, "since it's already flowing." Argue both sides in four sentences, then decide.
Q3 (scenario). Eight months after moving to full-state events, legal forwards a deletion request. The customer's email address exists in three places: checkout's database, the 30-day retained order.placed topic, and the local views of six consumer services. Two of those views have no listed owner. Design the deletion, and name the process failure that made it hard.
A1. Notification events contain an ID, and the producer retains the full truth. Event-carried state transfer includes the facts, but the producer remains authoritative while consumers keep disposable local copies. With event sourcing, the event log itself is the truth and tables are rebuilt from it.
A2. Fraud needs the value, and adding a field is a non-breaking change that avoids an API callback. However, lifetime value describes the customer rather than this order, becomes stale quickly, and exposes a sensitive business measure to every subscriber. Keep order.placed focused on order facts. Publish lifetime value separately or let fraud read it from the feature store, then combine both inputs in fraud's local view.
A3. Delete the user from checkout, publish a tombstone, and verify that every consumer purges its view. Confirm whether 30-day expiration meets the legal deadline; otherwise, use per-user encryption and destroy the key through crypto-shredding. The two unowned views reveal the larger process failure: subscribing to personal data did not require a named owner, catalog entry, and data-handling approval. Deletion is only manageable when every copy is known, and the long-term repair may be removing the email address from the event.
</details>11. Related Patterns
- Publish-Subscribe: the transport this pattern uses; this lesson decides what travels over it.
- Event Sourcing: the third level, where the log no longer feeds the truth but is the truth.
- Change Data Capture: produces full-state events directly from a database's change log, when the producer cannot or will not publish.
- Saga: the pattern for multi-step workflows, which events should not be used to imitate.
- Idempotency: view updates arrive at-least-once, meaning the same event can be delivered twice; applying it twice must leave the view correct.
12. TL;DR
| Problem | Notification-only events send every consumer back to the producer's API for details, so the coupling and the load return asynchronously. |
| Mechanism | Put the facts in the event; each consumer maintains a local, replay-rebuildable view and acts on its own. Facts have one owner; views are copies, never the truth. |
| Costs | Data copied into every subscriber, views that always lag slightly, personal-data spread that makes deletion a project, and an even heavier public-schema contract. |
| Skip it when | You are inside one service, the read must be exactly current, it is secretly a synchronous call or a workflow, or the team cannot yet support schema rules, tracing, and alerts. |
| 30-second answer | The event's contents are the architecture: carry enough state that consumers act from their own replayable local views instead of calling back, route exactly-current reads to the single owning service, keep sensitive fields out of the event, and govern the schema like the public API it is. |
Flashcards Review
What is event-driven architecture, and what is its central design decision?
rkrbommanapally
· 6 days ago
The over explanation of the topic is complex, its not simple. LIke the topic "The pattern", its difficult to understand and corelate, i already worked in kafka but this explanation is so complex i feel
Reading Progress
0%
On This Page
- The Incident
- The Obvious Fixes, and Why They Fail
- The Pattern
How much data the event carries
- Walkthrough with Numbers
- Trade-offs
- When Not to Use It
- Notification Events vs. Full-State Events
- Real-World Examples
- In the Interview
- Check Yourself
- Related Patterns
- TL;DR