MTBF, MTTR and availability: three numbers, three different questions
MTBF says how often a machine fails, MTTR how quickly it comes back, availability what share of the time the line can produce. We calculate all three from one example, show which decision each one supports, and list the four most common mistakes.
In a maintenance meeting, "our MTBF is 500 hours" says little on its own. The same machine can have an MTTR of 3 hours or of 30; availability is very different in the two cases. Three numbers answer three separate questions: how often does it fail, how quickly does it recover, how much time is left for production.
1. Definitions
- MTBF (mean time between failures) = total operating time / number of operating intervals. Repair time is not counted as operating time.
- MTTR (mean time to repair) = total repair time / number of failures.
- Availability = MTBF / (MTBF + MTTR). This is inherent availability, which counts only failures and repairs; planned maintenance and waiting for material are not included.
2. Example
Assume a filling line ran 2,000 hours in a quarter, had 4 unplanned failures in that time, and the repairs took 12 hours in total.
- MTBF = 2,000 / 4 = 500 hours
- MTTR = 12 / 4 = 3 hours
- Availability = 500 / (500 + 3) = 99.40%
The failure rate λ = 1 / MTBF = 0.002 failures per hour. Under a constant failure rate, the probability of getting through a 100-hour block of shifts without a failure is exp(−100/500) = 81.9%; the probability of 500 failure-free hours is exp(−1) = 36.8%. So MTBF does not mean "one failure every 500 hours"; the probability of surviving one MTBF without a failure is only 36.8%.
3. Which number supports which decision?
- A low MTBF points to design, operating conditions or part quality; preventive and predictive maintenance work on this.
- A high MTTR points to organization: spare parts, diagnosis time, reaching a qualified technician. In the example, cutting MTTR from 3 hours to 1.5 raises availability from 99.40% to 99.70%. The same gain comes from halving the failures from 4 to 2 (MTBF 1,000 hours). Both give the same result, but one is an investment in the maintenance plan, the other in spares and diagnostics.
- On a serial line availabilities multiply: three of the same machine in sequence give a line availability of 0.9940³ = 98.22%. Two machines in parallel with redundancy (one running is enough) give 1 − (1 − 0.9940)² = 99.996%. This assumes independent failures and a perfect switchover; in reality the switchover takes time.
4. Four common mistakes
- Counting repair time as operating time. MTBF is inflated and availability looks better than it is.
- Counting work orders instead of failures. One failure may have three work orders; MTBF drops to a third.
- Deriving MTBF from a single quarter. The 500 hours calculated from 4 failures become 1,000 hours next quarter with 2 failures; follow the trend, not the number. For equipment with few failures, at least 10–12 months of records are needed.
- Applying the constant-failure-rate assumption to wearing parts. For bearings, belts and seals the failure rate rises with time; there exp(−t/MTBF) is optimistic and a Weibull analysis is needed.
Try it with your own records
The MTBF / MTTR calculator works from summary figures or from a failure log (start–end date-times); the data stays in your browser. To have your failure log interpreted, get in touch.
Let's discuss this for your plant
More notes
All notes →- 3 min readBearing life L10: what it says and what it does not, with load, speed and reliabilityFrom the catalogue dynamic load rating to the basic rating life, then to the reliability and operating-condition adjustment: we calculate a bearing's expected life step by step, show why a small increase in load shortens life so much, and list the questions L10 does not answer.
- 3 min readPareto of downtime causes: 80/20 is not a law, it is a starting pointWe rank one month of downtime records by cause and work out the cumulative share. The example deliberately does not follow 80/20: five of eight causes make up 86.7% of the total. We show how to choose the threshold, how record quality distorts the result, and what comes after Pareto.
- 3 min readThe true cost of unplanned downtime: how to calculate lost hours, repair and the payback of predictive maintenanceThe cost of one stoppage is more than the repair invoice. We walk through lost production hours, repair and a realistic estimate of what early warning saves in one worked example, and show which assumption moves the result most.