The bathtub curve is the graph that describes how an equipment's failure rate evolves across its life: high at the beginning, low and constant in the middle, rising at the end. The shape recalls a bathtub in cross-section, and each of the three phases calls for a different maintenance response. Understanding the curve avoids the most expensive strategy mistake: applying the wrong phase's tool.
Phase 1: infant mortality
A high failure rate right after installation or intervention, falling with time. The causes have nothing to do with wear: manufacturing defects, assembly errors, contamination during installation, wrong startup settings. The right response is not more maintenance, it is more quality at the entrance: rigorous commissioning, assembly procedures, checked torque, measured alignment. An uncomfortable detail: every intervention restarts a small infant mortality, which is why too much preventive can worsen reliability instead of improving it.
Phase 2: useful life, the random failures
A low, approximately constant failure rate. The failures here do not announce themselves by the calendar: momentary overload, contamination, operating error, latent defects. Against random failure, time-based preventive is spend without return (replacing a good component does not change tomorrow's failure probability). What works: inspection and condition monitoring, predictive maintenance, which catches degradation as it starts, regardless of age.
Phase 3: wear-out
The failure rate growing with age: fatigue, abrasion, corrosion, insulation aging. Here age informs, and preventive by time or usage makes sense again, with the interval positioned before the curve's knee. Reading the history (MTBF and the failure distribution, as in the MTTR and MTBF guide) is what reveals where that knee is.
What the curve does not tell: not every asset follows the bathtub
The classic aviation reliability studies that gave birth to RCM (reliability centered maintenance) showed that most complex equipment has no dominant wear-out phase: failure patterns are mostly random, and only a minority of items benefit from age-based replacement. The practical consequence is direct: the curve applies per failure mode, not per whole equipment. The same gearbox has components in the wear-out phase (seals, bearings) and purely random modes (jamming by contamination). A good strategy mixes the tools per mode, which is exactly the job of reliability engineering.
The classic mistake: preventive against random failure
The symptom: a dense preventive plan, high cost, and the breakdowns continue. The frequent diagnosis: the asset's dominant failure modes are random (phase 2), and periodic replacement does not reach them, sometimes even feeding them via post-intervention infant mortality. The fix: migrate those modes to inspection and condition, and reserve time-based replacement for what actually wears. It is the kind of revision that the offender ranking and a well-kept failure history let you do with criteria.
Frequently Asked Questions
What is the bathtub curve?
It is the graph of the failure rate across an equipment's life, with three phases: infant mortality (high and falling), useful life (low and constant, random failures) and wear-out (rising with age).
What causes equipment infant mortality?
Manufacturing defects, assembly or installation errors, contamination and startup adjustments. The answer is quality in commissioning and interventions, not more periodic maintenance.
Does every equipment follow the bathtub curve?
No. The studies that founded RCM showed most complex equipment fails predominantly at random, with no dominant wear-out phase. The curve applies per failure mode, not per whole asset.
Which strategy fits each phase of the curve?
Infant mortality: assembly and commissioning quality. Random failures: inspection and condition monitoring. Wear-out: preventive by time or usage, with the interval placed before the failure rate takes off.
