0% completed
Resilience and Error Handling
An online store shows product reviews from a separate reviews service. One afternoon, the reviews service becomes slow. Each call to it now takes 30 seconds instead of 50 ms.
A few minutes later, the whole store stops loading. Even the checkout page fails, and checkout does not use reviews at all.
This lesson answers two questions. How does a system keep working when some of its parts fail? And how do we stop one failure from spreading to everything else?
Failures Are Normal
In a large system, some part is always failing
.....
.....
.....
sidpssp
· 2 years ago
notes
Aastha Bist
· 2 years ago
Under retry and backoff strategies, I did not understand this particular statement: "This can increase the likelihood of successful operation completion while preventing excessive load on the system during failure scenarios."
Can someone please provide more context/examples related to what this means?
Reading Progress
0%