Reliability engineering is the discipline that studies why assets fail and uses that knowledge to decide each asset's maintenance strategy. While the maintenance routine solves today's failure, reliability works so tomorrow's failure does not happen, or happens without consequence. This guide organizes the discipline's concepts, numbers and methods, and the practical path to start.
The base trio: reliability, availability and maintainability
Reliability is the probability that the asset performs its function, without failing, for a defined period under the expected operating conditions. Maintainability is how easily it returns to operation when it fails: access, component standardization, procedures, diagnosis. Availability is the result of both: the fraction of time the asset is fit to operate. The relationship is direct: high reliability spaces failures out (high MTBF), good maintainability shortens repairs (low MTTR), and availability collects both effects. Improving maintainability is usually the cheapest and least remembered path: you can gain availability without touching the failure rate, just by removing obstacles from the repair's way.
The numbers of reliability
The discipline speaks in numbers, and the central ones are MTBF (mean time between failures), the failure rate (its inverse) and MTTR. The formulas and measurement traps are in the MTTR and MTBF guide and in the maintenance KPIs guide. What reliability engineering adds: aggregates hide patterns. MTBF per equipment class shows where to invest; MTBF per individual asset points at the offender; and the distribution of failures over time (not just the mean) reveals which phase of life the asset is in, which is the bathtub curve's subject.
The bathtub curve: the three phases of failure
The curve describes how the failure rate evolves across an asset's life: high at the start (infant mortality, from manufacturing, assembly or startup errors), low and constant in the middle (random failures) and rising at the end (wear-out). Each phase calls for a different strategy response, and many assets do not even follow the full curve; the bathtub curve article breaks down the phases and what each changes in practice.
The methods: RCM, FMEA and root cause analysis
Three methods carry the discipline day to day. RCM (reliability centered maintenance) defines each asset's strategy from its failure modes and their consequences: the decision method. FMEA structures the survey of failure modes with prioritization by severity, occurrence and detection: the mapping method. And root cause analysis (5 whys, Ishikawa) dissects the failure that already happened so it does not repeat: the learning method. All three feed on the same raw material: a well-made failure record, with mode, cause and component cataloged on the notification.
The role in the structure: who does reliability
In larger plants, reliability engineering is a dedicated role, separate from the routine: while planning and control secures the week, the reliability engineer looks at the quarter and the year, hunting patterns in recurring failures and revising strategy. In smaller plants, it is a hat the maintenance engineer wears part of the time. The structural mistake is giving the hat to no one: without an owner, failure analysis only happens after disasters, and strategy freezes. The concept applied on the shop floor is in why reliability matters.
Where to start, without an academic project
The path of least resistance has four steps. One: guarantee decent failure records from now on (mode and cause catalogs on the notification, component identified); without that, any future analysis is born blind. Two: build the ranking of the 10 worst offenders by downtime and by cost over the last 12 months; the top of the list concentrates a disproportionate share of the pain. Three: run real root cause analysis on the top 3, with verified countermeasures. Four: revise those assets' strategy (plan, frequency, applicable predictive techniques) and measure before and after by MTBF. Reliability starts small and concrete, not with statistical software.
Frequently Asked Questions
What does reliability engineering do?
It studies asset failure patterns (history, failure modes, distribution over time) and uses that knowledge to define and revise maintenance strategy, prioritizing the equipment that hurts most in downtime and cost.
What is the difference between reliability and availability?
Reliability is the probability of not failing during a period; availability is the fraction of time fit to operate, resulting from reliability (spacing failures) combined with maintainability (repairing fast).
What is maintainability?
It is how easily an asset returns to operation: physical access, diagnosis, component standardization, procedures and spare parts logistics. Improving it reduces MTTR without touching the failure rate.
Does a high MTBF mean good reliability?
It is a good sign, but the mean hides patterns: two assets with the same MTBF can have completely different failure distributions, calling for different strategies. Reliability engineering looks at the distribution, not just the mean.
