Capability 10 / 18
Site Reliability
Keeping production healthy with SLOs, on-call discipline, and fast recovery.
Site reliability is where architecture meets 2 a.m. — the discipline of keeping production trustworthy under real, unpredictable conditions. That means SLOs that reflect what users actually experience, on-call rotations that don't burn people out, and incident response that finds root cause instead of just restoring service.
It's end-to-end platform accountability: capacity and risk stewardship, change management that doesn't treat every deploy as a coin flip, and a postmortem discipline that actually changes what happens next time.
What this looks like in practice
- •SLO/SLA targets grounded in real user impact, not arbitrary round numbers
- •Incident command and postmortem discipline that produces real follow-through, not just a document
- •Change management and rollout safety built to make deploys boring
- •Capacity and risk stewardship reviewed before it becomes an incident