{"article":{"slug":"retry-vs-circuit-breaker-vs-fallback","title":"Retry vs Circuit Breaker vs Fallback","subtitle":null,"summary":"Retries, circuit breakers, and fallbacks solve different failure modes in backend integrations. Using the wrong one can amplify outages—here is when each pattern helps and when it hurts.","content_type":"tutorial","language":"en","canonical_url":"https://kkbhati07.github.io/notes/retry-circuit-breaker-fallback.html","author":{"name":"Kanishk Bhati","url":"https://kkbhati07.github.io/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Kanishk Bhati","url":"https://kkbhati07.github.io/","listing_slug":null,"listing":null},"topics":[{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Infrastructure","slug":"infrastructure","url":"https://listedarticles.com/topics/infrastructure"},{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1469,"reading_minutes":6,"published_at":"2026-09-24T00:00:00.000Z","added_at":"2026-09-25T09:17:00.098Z","updated_at":"2026-09-25T09:17:00.098Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/retry-vs-circuit-breaker-vs-fallback","markdown_url":"https://listedarticles.com/articles/retry-vs-circuit-breaker-vs-fallback.md","example":false,"citation":"Kanishk Bhati, Kanishk Bhati. \"Retry vs Circuit Breaker vs Fallback.\" 24 Sept 2026. https://kkbhati07.github.io/notes/retry-circuit-breaker-fallback.html (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://kkbhati07.github.io/notes/retry-circuit-breaker-fallback.html"},"body_markdown":"## The problem\n\nOne thing I have been paying more attention to in backend systems is how they behave when something they depend on stops behaving normally.\n\nAn API call to an external service can fail for many reasons. The dependency might be temporarily unavailable, slow to respond, returning errors, or failing consistently.\n\nThe first instinct is often simple:\n\n“Try again.”\n\nSometimes that is exactly what the system needs.\n\nBut retrying is not a complete failure strategy.\n\nIf the dependency is already struggling, repeatedly sending the same request can create even more traffic. A temporary failure can turn into a larger problem if every caller keeps retrying without any limits.\n\nThat is where I started thinking about retries, circuit breakers, and fallbacks as three different tools rather than three interchangeable patterns.\n\nA retry asks:\n\n“Should I try this operation again?”\n\nA circuit breaker asks:\n\n“Should I keep calling this dependency at all right now?”\n\nA fallback asks:\n\n“If the dependency cannot complete this operation, what can I do instead?”\n\nThose are different questions.\n\n## Retry: dealing with transient failures\n\nRetries are useful when a failure is likely to be temporary.\n\nFor example, a network request might fail because of a short-lived connection problem or a temporary interruption in the dependency.\n\nInstead of immediately failing the request, the application can try the operation again.\n\nConceptually:\n\nRetry (conceptual)\n\nThe important part is that retries need boundaries.\n\nA retry policy should consider things such as:\n\n- How many times should the operation be retried?\n- Which failures are actually retryable?\n- Should there be a delay between attempts?\n- Should the delay increase between attempts?\n- Is the operation safe to repeat?\n\nThe last question is particularly important.\n\nRetrying a read is generally a different problem from retrying an operation that changes state. If an operation creates something, charges something, sends something, or otherwise has side effects, blindly repeating it can create a second operation rather than simply recovering the first one.\n\nSo the lesson for me is not “add retries.”\n\nIt is:\n\nRetry when the failure looks transient and the operation can be safely retried.\n\n## Why retries alone are not enough\n\nImagine an external dependency is consistently failing.\n\nWithout a limit, every incoming request can trigger multiple attempts.\n\nOne request becomes several requests to the failing service.\n\nAt scale, that can look something like:\n\nUnbounded retries (conceptual)\n\nThe application is spending more work trying to reach something that is already unavailable.\n\nThis is where retries can become part of the problem.\n\nA retry mechanism should therefore be designed with the possibility that the dependency may remain unavailable for longer than expected.\n\nThat led me to the next question:\n\nWhat happens when continuing to call the dependency is no longer useful?\n\n## Circuit breaker: stopping repeated failure\n\nA circuit breaker introduces another decision.\n\nInstead of allowing every request to keep calling a failing dependency, the application can temporarily stop making those calls after failures cross a defined threshold.\n\nConceptually:\n\nCircuit breaker (conceptual)\n\nYes\n\nNo\n\nThe useful part is not the name “circuit breaker.”\n\nThe useful part is the change in behavior.\n\nInstead of repeatedly discovering that the dependency is unavailable, the application can recognize the failure pattern and stop sending unnecessary traffic for a period of time.\n\nThat gives the dependency some breathing room and prevents the application from spending resources repeatedly waiting for failures.\n\nEventually, the circuit can allow requests through again to determine whether the dependency has recovered.\n\nThe important idea is that a circuit breaker is about controlling calls to an unhealthy dependency.\n\nA retry is about giving an individual operation another chance.\n\nThey solve different problems.\n\n## Retry vs Circuit Breaker\n\nI think about the distinction like this:\n\n### Retry\n\n“Maybe this particular failure is temporary.”\n\nThe system makes another attempt.\n\n### Circuit breaker\n\n“This dependency appears unhealthy right now.”\n\nThe system stops repeatedly calling it.\n\nThat means they can work together.\n\nA request might retry a small number of times for a transient failure, while the circuit breaker prevents the service from continuing to hammer the dependency when failures become persistent.\n\nThe exact thresholds and timing depend on the system. There is no universal number that makes a circuit breaker correct.\n\nThe important part is understanding what decision the mechanism is making.\n\n## Fallback: doing something useful when the dependency fails\n\nRetries and circuit breakers answer questions about calling the dependency.\n\nA fallback answers a different question:\n\n“What should the application do if that dependency cannot provide the result?”\n\nSometimes the answer is to return a controlled error.\n\nSometimes the application can use previously available data.\n\nSometimes a default response is acceptable.\n\nSometimes the operation simply cannot continue and the system should communicate that clearly to the caller.\n\nConceptually:\n\nFallback (conceptual)\n\nA fallback should not pretend that the dependency succeeded when it did not.\n\nThat distinction matters.\n\nReturning stale or partial information might be acceptable for one type of operation and completely wrong for another.\n\nThe fallback has to match the business meaning of the operation.\n\n## Putting the three together\n\nThe most useful mental model for me is not:\n\n“Which one should I use?”\n\nIt is:\n\n“What kind of failure am I dealing with, and what should happen next?”\n\nA simplified flow looks like this:\n\nFailure path (conceptual)\n\nThis is only a conceptual model. Real systems can have more states and more detailed policies.\n\nThe important part is that each mechanism has a different responsibility.\n\n- **Retry:** Give a potentially transient operation another chance.\n- **Circuit breaker:** Stop repeatedly calling a dependency that appears unhealthy.\n- **Fallback:** Provide an alternative behavior when the dependency cannot complete the operation.\n\n## The part I find easy to get wrong\n\nIt is tempting to think of reliability patterns as things you simply add to a service.\n\nAdd retries.\n\nAdd a circuit breaker.\n\nAdd a fallback.\n\nDone.\n\nBut each one changes system behavior.\n\nA retry changes how many times an operation can execute.\n\nA circuit breaker changes whether a dependency is called at all.\n\nA fallback changes what the user or calling service receives when something fails.\n\nThose decisions should therefore be connected to the actual operation.\n\nFor example, retrying a request that is safe to repeat is a different problem from retrying an operation with side effects.\n\nSimilarly, returning cached data as a fallback might be useful for a read-heavy endpoint but unacceptable when the caller requires the latest authoritative state.\n\nReliability is therefore not only about keeping requests alive.\n\nIt is also about deciding which behavior is safe when the normal path stops working.\n\n## Failure handling is also about protecting the system\n\nOne thing I increasingly appreciate about these patterns is that they protect more than the individual request.\n\nA retry can help recover from a short-lived problem.\n\nA circuit breaker can prevent a struggling dependency from receiving an increasing number of repeated requests.\n\nA fallback can allow the rest of the system to degrade in a controlled way instead of failing unpredictably.\n\nThat means failure handling is partly about protecting the dependency and partly about protecting the application itself.\n\nThe goal is not necessarily:\n\n“Never show an error.”\n\nSometimes the correct behavior is a clear, controlled failure.\n\nA reliable system is not one where failures never happen.\n\nIt is one where failures have predictable behavior.\n\n## Trade-offs\n\n### Retries\n\nUseful for:\n\n- Transient failures\n- Short-lived network problems\n- Operations that are safe to repeat\n\nTrade-offs:\n\n- Additional latency\n- Additional traffic\n- Risk of duplicate side effects if the operation is not safe to repeat\n- Can amplify an existing outage if poorly bounded\n\n### Circuit breakers\n\nUseful for:\n\n- Consistently failing dependencies\n- Preventing repeated calls to an unhealthy service\n- Failing fast when continuing to call is unlikely to help\n\nTrade-offs:\n\n- Additional state and configuration\n- Temporary failures can cause calls to be rejected while the circuit is open\n- Thresholds and recovery behavior need to match the system\n\n### Fallbacks\n\nUseful for:\n\n- Graceful degradation\n- Returning alternative or previously available information\n- Providing controlled behavior when a dependency is unavailable\n\nTrade-offs:\n\n- The fallback may provide less complete information\n- Stale or partial data can be misleading if used incorrectly\n- Not every operation has a meaningful fallback\n\nNone of these patterns is universally required.\n\nThe architecture should reflect the dependency, the operation being performed, and what failure means to the application.\n\n## What I took away\n\nThe biggest shift for me was moving away from thinking about retries as the default answer to every external-service failure.\n\nSometimes a retry is exactly right.\n\nSometimes continuing to retry is the wrong thing to do.\n\nSometimes the system should stop calling the dependency for a while.\n\nSometimes the right answer is to return something different.\n\nRetries, circuit breakers, and fallbacks are therefore not competing solutions. They are different decisions in the failure path.\n\nThe useful question is not:\n\n“Which pattern should I add?”\n\nIt is:\n\n“What should this system do when the normal path stops working?”\n\nThat question leads to much better backend designs.","body_html":"<h2 id=\"the-problem\">The problem</h2>\n<p>One thing I have been paying more attention to in backend systems is how they behave when something they depend on stops behaving normally.</p>\n<p>An API call to an external service can fail for many reasons. The dependency might be temporarily unavailable, slow to respond, returning errors, or failing consistently.</p>\n<p>The first instinct is often simple:</p>\n<p>“Try again.”</p>\n<p>Sometimes that is exactly what the system needs.</p>\n<p>But retrying is not a complete failure strategy.</p>\n<p>If the dependency is already struggling, repeatedly sending the same request can create even more traffic. A temporary failure can turn into a larger problem if every caller keeps retrying without any limits.</p>\n<p>That is where I started thinking about retries, circuit breakers, and fallbacks as three different tools rather than three interchangeable patterns.</p>\n<p>A retry asks:</p>\n<p>“Should I try this operation again?”</p>\n<p>A circuit breaker asks:</p>\n<p>“Should I keep calling this dependency at all right now?”</p>\n<p>A fallback asks:</p>\n<p>“If the dependency cannot complete this operation, what can I do instead?”</p>\n<p>Those are different questions.</p>\n<h2 id=\"retry-dealing-with-transient-failures\">Retry: dealing with transient failures</h2>\n<p>Retries are useful when a failure is likely to be temporary.</p>\n<p>For example, a network request might fail because of a short-lived connection problem or a temporary interruption in the dependency.</p>\n<p>Instead of immediately failing the request, the application can try the operation again.</p>\n<p>Conceptually:</p>\n<p>Retry (conceptual)</p>\n<p>The important part is that retries need boundaries.</p>\n<p>A retry policy should consider things such as:</p>\n<ul><li>How many times should the operation be retried?</li><li>Which failures are actually retryable?</li><li>Should there be a delay between attempts?</li><li>Should the delay increase between attempts?</li><li>Is the operation safe to repeat?</li></ul>\n<p>The last question is particularly important.</p>\n<p>Retrying a read is generally a different problem from retrying an operation that changes state. If an operation creates something, charges something, sends something, or otherwise has side effects, blindly repeating it can create a second operation rather than simply recovering the first one.</p>\n<p>So the lesson for me is not “add retries.”</p>\n<p>It is:</p>\n<p>Retry when the failure looks transient and the operation can be safely retried.</p>\n<h2 id=\"why-retries-alone-are-not-enough\">Why retries alone are not enough</h2>\n<p>Imagine an external dependency is consistently failing.</p>\n<p>Without a limit, every incoming request can trigger multiple attempts.</p>\n<p>One request becomes several requests to the failing service.</p>\n<p>At scale, that can look something like:</p>\n<p>Unbounded retries (conceptual)</p>\n<p>The application is spending more work trying to reach something that is already unavailable.</p>\n<p>This is where retries can become part of the problem.</p>\n<p>A retry mechanism should therefore be designed with the possibility that the dependency may remain unavailable for longer than expected.</p>\n<p>That led me to the next question:</p>\n<p>What happens when continuing to call the dependency is no longer useful?</p>\n<h2 id=\"circuit-breaker-stopping-repeated-failure\">Circuit breaker: stopping repeated failure</h2>\n<p>A circuit breaker introduces another decision.</p>\n<p>Instead of allowing every request to keep calling a failing dependency, the application can temporarily stop making those calls after failures cross a defined threshold.</p>\n<p>Conceptually:</p>\n<p>Circuit breaker (conceptual)</p>\n<p>Yes</p>\n<p>No</p>\n<p>The useful part is not the name “circuit breaker.”</p>\n<p>The useful part is the change in behavior.</p>\n<p>Instead of repeatedly discovering that the dependency is unavailable, the application can recognize the failure pattern and stop sending unnecessary traffic for a period of time.</p>\n<p>That gives the dependency some breathing room and prevents the application from spending resources repeatedly waiting for failures.</p>\n<p>Eventually, the circuit can allow requests through again to determine whether the dependency has recovered.</p>\n<p>The important idea is that a circuit breaker is about controlling calls to an unhealthy dependency.</p>\n<p>A retry is about giving an individual operation another chance.</p>\n<p>They solve different problems.</p>\n<h2 id=\"retry-vs-circuit-breaker\">Retry vs Circuit Breaker</h2>\n<p>I think about the distinction like this:</p>\n<h3 id=\"retry\">Retry</h3>\n<p>“Maybe this particular failure is temporary.”</p>\n<p>The system makes another attempt.</p>\n<h3 id=\"circuit-breaker\">Circuit breaker</h3>\n<p>“This dependency appears unhealthy right now.”</p>\n<p>The system stops repeatedly calling it.</p>\n<p>That means they can work together.</p>\n<p>A request might retry a small number of times for a transient failure, while the circuit breaker prevents the service from continuing to hammer the dependency when failures become persistent.</p>\n<p>The exact thresholds and timing depend on the system. There is no universal number that makes a circuit breaker correct.</p>\n<p>The important part is understanding what decision the mechanism is making.</p>\n<h2 id=\"fallback-doing-something-useful-when-the-dependency-fails\">Fallback: doing something useful when the dependency fails</h2>\n<p>Retries and circuit breakers answer questions about calling the dependency.</p>\n<p>A fallback answers a different question:</p>\n<p>“What should the application do if that dependency cannot provide the result?”</p>\n<p>Sometimes the answer is to return a controlled error.</p>\n<p>Sometimes the application can use previously available data.</p>\n<p>Sometimes a default response is acceptable.</p>\n<p>Sometimes the operation simply cannot continue and the system should communicate that clearly to the caller.</p>\n<p>Conceptually:</p>\n<p>Fallback (conceptual)</p>\n<p>A fallback should not pretend that the dependency succeeded when it did not.</p>\n<p>That distinction matters.</p>\n<p>Returning stale or partial information might be acceptable for one type of operation and completely wrong for another.</p>\n<p>The fallback has to match the business meaning of the operation.</p>\n<h2 id=\"putting-the-three-together\">Putting the three together</h2>\n<p>The most useful mental model for me is not:</p>\n<p>“Which one should I use?”</p>\n<p>It is:</p>\n<p>“What kind of failure am I dealing with, and what should happen next?”</p>\n<p>A simplified flow looks like this:</p>\n<p>Failure path (conceptual)</p>\n<p>This is only a conceptual model. Real systems can have more states and more detailed policies.</p>\n<p>The important part is that each mechanism has a different responsibility.</p>\n<ul><li><strong>Retry:</strong> Give a potentially transient operation another chance.</li><li><strong>Circuit breaker:</strong> Stop repeatedly calling a dependency that appears unhealthy.</li><li><strong>Fallback:</strong> Provide an alternative behavior when the dependency cannot complete the operation.</li></ul>\n<h2 id=\"the-part-i-find-easy-to-get-wrong\">The part I find easy to get wrong</h2>\n<p>It is tempting to think of reliability patterns as things you simply add to a service.</p>\n<p>Add retries.</p>\n<p>Add a circuit breaker.</p>\n<p>Add a fallback.</p>\n<p>Done.</p>\n<p>But each one changes system behavior.</p>\n<p>A retry changes how many times an operation can execute.</p>\n<p>A circuit breaker changes whether a dependency is called at all.</p>\n<p>A fallback changes what the user or calling service receives when something fails.</p>\n<p>Those decisions should therefore be connected to the actual operation.</p>\n<p>For example, retrying a request that is safe to repeat is a different problem from retrying an operation with side effects.</p>\n<p>Similarly, returning cached data as a fallback might be useful for a read-heavy endpoint but unacceptable when the caller requires the latest authoritative state.</p>\n<p>Reliability is therefore not only about keeping requests alive.</p>\n<p>It is also about deciding which behavior is safe when the normal path stops working.</p>\n<h2 id=\"failure-handling-is-also-about-protecting-the-system\">Failure handling is also about protecting the system</h2>\n<p>One thing I increasingly appreciate about these patterns is that they protect more than the individual request.</p>\n<p>A retry can help recover from a short-lived problem.</p>\n<p>A circuit breaker can prevent a struggling dependency from receiving an increasing number of repeated requests.</p>\n<p>A fallback can allow the rest of the system to degrade in a controlled way instead of failing unpredictably.</p>\n<p>That means failure handling is partly about protecting the dependency and partly about protecting the application itself.</p>\n<p>The goal is not necessarily:</p>\n<p>“Never show an error.”</p>\n<p>Sometimes the correct behavior is a clear, controlled failure.</p>\n<p>A reliable system is not one where failures never happen.</p>\n<p>It is one where failures have predictable behavior.</p>\n<h2 id=\"trade-offs\">Trade-offs</h2>\n<h3 id=\"retries\">Retries</h3>\n<p>Useful for:</p>\n<ul><li>Transient failures</li><li>Short-lived network problems</li><li>Operations that are safe to repeat</li></ul>\n<p>Trade-offs:</p>\n<ul><li>Additional latency</li><li>Additional traffic</li><li>Risk of duplicate side effects if the operation is not safe to repeat</li><li>Can amplify an existing outage if poorly bounded</li></ul>\n<h3 id=\"circuit-breakers\">Circuit breakers</h3>\n<p>Useful for:</p>\n<ul><li>Consistently failing dependencies</li><li>Preventing repeated calls to an unhealthy service</li><li>Failing fast when continuing to call is unlikely to help</li></ul>\n<p>Trade-offs:</p>\n<ul><li>Additional state and configuration</li><li>Temporary failures can cause calls to be rejected while the circuit is open</li><li>Thresholds and recovery behavior need to match the system</li></ul>\n<h3 id=\"fallbacks\">Fallbacks</h3>\n<p>Useful for:</p>\n<ul><li>Graceful degradation</li><li>Returning alternative or previously available information</li><li>Providing controlled behavior when a dependency is unavailable</li></ul>\n<p>Trade-offs:</p>\n<ul><li>The fallback may provide less complete information</li><li>Stale or partial data can be misleading if used incorrectly</li><li>Not every operation has a meaningful fallback</li></ul>\n<p>None of these patterns is universally required.</p>\n<p>The architecture should reflect the dependency, the operation being performed, and what failure means to the application.</p>\n<h2 id=\"what-i-took-away\">What I took away</h2>\n<p>The biggest shift for me was moving away from thinking about retries as the default answer to every external-service failure.</p>\n<p>Sometimes a retry is exactly right.</p>\n<p>Sometimes continuing to retry is the wrong thing to do.</p>\n<p>Sometimes the system should stop calling the dependency for a while.</p>\n<p>Sometimes the right answer is to return something different.</p>\n<p>Retries, circuit breakers, and fallbacks are therefore not competing solutions. They are different decisions in the failure path.</p>\n<p>The useful question is not:</p>\n<p>“Which pattern should I add?”</p>\n<p>It is:</p>\n<p>“What should this system do when the normal path stops working?”</p>\n<p>That question leads to much better backend designs.</p>","headings":[{"level":2,"text":"The problem","id":"the-problem"},{"level":2,"text":"Retry: dealing with transient failures","id":"retry-dealing-with-transient-failures"},{"level":2,"text":"Why retries alone are not enough","id":"why-retries-alone-are-not-enough"},{"level":2,"text":"Circuit breaker: stopping repeated failure","id":"circuit-breaker-stopping-repeated-failure"},{"level":2,"text":"Retry vs Circuit Breaker","id":"retry-vs-circuit-breaker"},{"level":3,"text":"Retry","id":"retry"},{"level":3,"text":"Circuit breaker","id":"circuit-breaker"},{"level":2,"text":"Fallback: doing something useful when the dependency fails","id":"fallback-doing-something-useful-when-the-dependency-fails"},{"level":2,"text":"Putting the three together","id":"putting-the-three-together"},{"level":2,"text":"The part I find easy to get wrong","id":"the-part-i-find-easy-to-get-wrong"},{"level":2,"text":"Failure handling is also about protecting the system","id":"failure-handling-is-also-about-protecting-the-system"},{"level":2,"text":"Trade-offs","id":"trade-offs"},{"level":3,"text":"Retries","id":"retries"},{"level":3,"text":"Circuit breakers","id":"circuit-breakers"},{"level":3,"text":"Fallbacks","id":"fallbacks"},{"level":2,"text":"What I took away","id":"what-i-took-away"}]}}