a practice that allows systems to provide basic functionality even when core processes are not available; this is in contrast to being completely unavailable when one function fails.
— DevOps for the Modern Enterprise, Mirco Hering
1 of our books defines this term.
a practice that allows systems to provide basic functionality even when core processes are not available; this is in contrast to being completely unavailable when one function fails.
— DevOps for the Modern Enterprise, Mirco Hering
See also blameless postmortem · concept of error budgets · incident · mean time to discovery (MTTD) · mean time to recovery (MTTR) · site reliability engineering
Filed under Reliability & Incidents
Every book and audiobook, the video library, the courses, and the MCP. The papers stay free.