site reliability engineering

3 of our books and 1 talk define this term — and their definitions are worth reading side by side.

SRE (site reliability engineering) adopts software engineering principles and practices and applies them to infrastructure and operations problems. The main goal is to create highly reliable and scalable software systems.

Investments Unlimited, in “Chapter 13”, Helen Beal, Bill Bensing, Jason Cox, Michael Edenzon, Tapabrata Pal, Caleb Queern, John Rzeszotarski, Andres Vega, and John Willis

Site reliability engineering is Google's take to delivering value to customers. It talks about building end-to-end reliability at site. Site reliability engineering is more a post-production set of processes and activities for systems at scale, and it operates on the principles of prevent, recover, and optimize

Meenal Meenaakshi · DevOps SRE or ITIL – Know Before You Leap!, US 2021 · at 12:58

See also blameless postmortem · concept of error budgets · graceful degradation · incident · mean time to discovery (MTTD) · mean time to recovery (MTTR)

Filed under Reliability & Incidents

← All glossary terms

One membership. The whole library.

Every book and audiobook, the video library, the courses, and the MCP. The papers stay free.

Start membership