Fault-Tolerant System
Updated 2026-08-15
INTRODUCTION
English translation pending.
CORE DEFINITION
Fault tolerance designs a system so that a failure in one component degrades performance rather than causing collapse, using redundancy, error detection and isolation, graceful degradation, and defined recovery rules. It accepts that failures will occur and invests in containing them instead of preventing every one. The qualification is that redundancy adds cost and complexity, and that tolerance can mask an underlying cause if failures are absorbed without diagnosis.
SCAFFOLDING EFFECT
Reduce cognitive load
- Use Failure Containment: define how each failure is isolated before it can spread to the whole. - Use Degradation Planning: decide in advance which functions are shed first when capacity drops.
Anchor fast decisions
Any single component has a nonzero failure rate, so prevention alone leaves the whole system exposed to the weakest part. Adding independent paths and detection means a failure removes capability rather than availability, and the system continues at reduced capacity while the fault is handled.
MINIMUM ACTION
In progress 0/1Practice this model in one real situation:
account_treeGenealogyexpand_more
menu_bookReferencesexpand_more
Source support: Explicit
- en.wikipedia.orghttps://en.wikipedia.org/wiki/Fault_toleranceverified
PRIVATE NOTES · Only visible to you
SAVED Q&A
ENTRY Q&A · Private saving available
Ask with a clear boundary
thinkingmodels answers from published entry context only.
Your question is sent to thinkingmodels. The answer uses public entry context only.
RELATED MODELS