Resilience management - Measure
MTTR
(Mean Time to Recovery)
The average time from a failure until the service is working again
What's it for?
Shows how quickly a failed service usually becomes usable again.
For example…
Over a quarter, three production incidents took 40, 20, and 30 minutes to restore; MTTR is about half an hour
If ten outages take an average of 35 minutes from failure to restored service, the service's MTTR for recovery is 35 minutes
Think of it like…
Average time to get the shop open again after a power cut — not how often the power fails