Resilience management - Measure
MTBF
(Mean Time Between Failures)
The average time a system runs successfully between one failure and the next
What's it for?
Shows how long something usually runs between failures, helping teams judge reliability.
For example…
A disk array that fails twice a year has a much lower MTBF than one that fails once every five years.
If a device typically runs for about 1,000 hours before failing, its MTBF is roughly 1,000 hours
Think of it like…
How many miles a car averages between breakdowns—longer gaps mean more trustworthy kit