Reliability & Incidents
Keeping it up, and learning when it is not: SRE vocabulary, error budgets, incident swarms, the blunt end and the sharp end, and resilience engineering's view of how systems really fail.
56 terms filed here — every definition a credited primary source.
- abstract incident metricsFindings From The Field: Two Years of Studying Incidents Closely
- AdaptabilitySignals and Levers
- ambulance diversionLeadership Lessons Learned From Improving Flow In Hospital Settings using Theory of Constraints
- anti-fragile philosophyRadical Ideas Enterprises Can Learn From The Cloud
- availabilityIndustrial DevOps
- blameless postmortemDevOps for the Modern Enterprise
- blunt endFindings From The Field: Two Years of Studying Incidents Closely
- break glassDevOps at Capital One: Focusing on Pipeline and Measurement
- business-critical applicationsEdge of the Present: Pursuing AI’s Frontier in the Enterprise
- CapacitySignals and Levers
- Chaos Engineering5 books · 2 talks
- code yellowHow Google SRE and Developers Work Together
- concept of error budgetsDevOps for the Modern Enterprise
- CountermeasuresMaking Work Visible, Second Edition
- destructive testingThe DevOps Handbook
- DRIMoving to 1ES at Microsoft
- embracing the redVisual Studio Online: The Inside Story from COTS to Cloud
- error budgetShifting Left on Production Excellence with Observability
- error budget policy1 book · 1 talk
- Escaped defects2 books define this
- frontline-affected downtimeAgile/Dawie Ruined My Life
- going solidFeatures Versus Futures
- governance toilDear Security, Compliance, and Auditors, We’re Sorry. Love, DevOps.
- graceful degradationDevOps for the Modern Enterprise
- impact minutesDevOps Lessons from Executive Leadership: Why DevOps Matters to Corporate Performance
- incidentDevOps for the Modern Enterprise
- Incident Response TeamsHow Google SRE and Developers Collaborate
- incident swarm modelMore Culture, More Engineering, Less Duct Tape v3.0
- latencyIndustrial DevOps
- LATTESThe Ops in DevOps is a VERB. It is not a Noun!
- mean time to discovery (MTTD)DevOps for the Modern Enterprise
- mean time to recovery (MTTR)DevOps for the Modern Enterprise
- messy detailsLearning Effectively From Incidents: The Messy Details
- microfracture repairWorking at the Center of the Cyclone
- mission-criticalDigital Transformation - Thriving Through the Transition
- MonstersWiring the Winning Organization
- MTTUVendorDome: Can You Accelerate Delivery Without Compromising Security?
- performance environmentWiring the Winning Organization
- recalibration problemWorking at the Center of the Cyclone
- RetrospectivesAgile Conversations
- revenue protection work1 book · 1 talk
- rule of threeBeyond the Culture Deck What You Don’t Already Know About Netflix
- Runtime couplingWiring the Winning Organization
- senior engineerWhat Handcrafted Servers Taught Us About Handcrafted Code
- service handback mechanismThe DevOps Handbook
- service operability assessmentOperability and You Build It You Run It at John Lewis & Partners
- severity one swarmHow Swarming Transforms Enterprise Support to Work Better with DevOps
- site reliability engineering3 books · 1 talk
- SLOEngineering ITSM Through Site Reliability Engineering
- SNAFUWorking at the Center of the Cyclone
- stabilizationWiring the Winning Organization
- stop-work authorityBridging the Gap: Empowering Civil Engineering Projects through DevOps and Fusion Teams
- swarm modelMore Culture, More Engineering, Less Duct-Tape
- telematicsDevOps for the Modern Enterprise
- unknown-unknownsSooner Safer Happier
- wall of failIn Search of DevOps - the Evolution of a Data Warehouse Team Towards DevOps