Back
Maintenance Productivity

Reliability engineering in maintenance: concepts, curves and methods

P
PM Run Team
Published
Updated

Maintenance reliability engineering is the discipline that studies why assets fail and uses that knowledge to decide the maintenance strategy of each asset: which failures to prevent, which to detect early and which to accept. While the maintenance routine solves today's failure, reliability engineering works so that tomorrow's failure does not happen, or happens without consequence.

The discipline rests on a few concepts, a handful of numbers and three methods, and it only works when the failure record that feeds all of them is trustworthy. The root cause failure analysis (RCFA) method is where most plants see the discipline for the first time.

The base trio: reliability, availability and maintainability

Reliability is the probability that the asset performs its function, without failing, for a defined period under the expected operating conditions. Maintainability is how easily it returns to operation when it fails: access, component standardization, procedures, diagnosis. Availability is the result of both: the fraction of time the asset is fit to operate. The relationship is direct: high reliability spaces failures out (high MTBF), good maintainability shortens repairs (low MTTR), and availability collects both effects. Improving maintainability is usually the cheapest and least remembered path: you can gain availability without touching the failure rate, just by removing obstacles from the repair's way.

The numbers of reliability

The discipline speaks in numbers, and the central ones are MTBF (mean time between failures), the failure rate (its inverse) and MTTR. The formulas and measurement traps are in the MTTR and MTBF guide and in the maintenance KPIs guide. What reliability engineering adds: aggregates hide patterns. MTBF per equipment class shows where to invest; MTBF per individual asset points at the offender; and the distribution of failures over time (not just the mean) reveals which phase of life the asset is in, which is the bathtub curve's subject.

The bathtub curve: the three phases of failure

The curve describes how the failure rate evolves across an asset's life: high at the start (infant mortality, from manufacturing, assembly or startup errors), low and constant in the middle (random failures) and rising at the end (wear-out). Each phase calls for a different strategy response, and many assets do not even follow the full curve; the bathtub curve article breaks down the phases and what each changes in practice.

The methods: RCM, FMEA and root cause analysis

Three methods carry the discipline day to day. RCM (reliability centered maintenance) defines each asset's strategy from its failure modes and their consequences: the decision method. FMEA structures the survey of failure modes with prioritization by severity, occurrence and detection: the mapping method. And root cause analysis (5 whys, Ishikawa, fault tree analysis) dissects the failure that already happened so it does not repeat: the learning method. All three feed on the same raw material: a well-made failure record, with mode, cause and component cataloged on the notification.

The role in the structure: who does reliability

In larger plants, reliability engineering is a dedicated role, separate from the routine: while planning and control secures the week, the reliability engineer looks at the quarter and the year, hunting patterns in recurring failures and revising strategy. In smaller plants, it is a hat the maintenance engineer wears part of the time. The structural mistake is giving the hat to no one: without an owner, failure analysis only happens after disasters, and strategy freezes. The sections below show what the concept looks like on the shop floor.

Where to start, without an academic project

The path of least resistance has four steps. One: guarantee decent failure records from now on (mode and cause catalogs on the notification, component identified); without that, any future analysis is born blind. Two: build the ranking of the 10 worst offenders by downtime and by cost over the last 12 months; the top of the list concentrates a disproportionate share of the pain. Three: run real root cause analysis on the top 3, with verified countermeasures. Four: revise those assets' strategy (plan, frequency, applicable predictive techniques) and measure before and after by MTBF. Reliability starts small and concrete, not with statistical software.

Maintenance and reliability engineering: how the two work together

Maintenance and reliability engineering is usually one area with two time horizons. Maintenance executes: it plans, schedules and performs the work that keeps assets running this week. Reliability engineering decides: it looks at the failure history of months and years to define what work should exist, on which assets and at what interval. The Society for Maintenance & Reliability Professionals (SMRP) treats both as one body of knowledge, with equipment reliability and work management as separate pillars that depend on each other.

Maintenance vs reliability: where the roles differ

QuestionMaintenanceReliability engineering
HorizonToday, this week, this shutdownThe quarter, the year, the asset life
Main outputExecuted work orders, restored assetsRevised strategies, plans and designs
Typical questionHow do we fix this failure fast and safely?Why does this failure keep coming back, and what stops it?
Core indicatorsMTTR, schedule compliance, backlogMTBF, recurrence, failure distribution
Raw materialNotifications, orders, confirmationsThe history those same records build

The last row explains why the two cannot be separated. Reliability engineering has no data of its own: every analysis it runs uses the notifications and confirmations that maintenance produces. If execution records late or badly, the reliability engineer analyzes noise.

Reliability and maintainability engineering, the pair often abbreviated as R&M, adds the design side: the same principles applied before the asset exists, when access, standardization and diagnosis can still be designed in. The international vocabulary for these terms is IEC 60050-192, the dependability part of the International Electrotechnical Vocabulary, which defines reliability, maintainability and availability as the components of dependability.

Maintenance reliability analysis starts with failure capture

Maintenance reliability is the ability to prove, with field records, which asset failed, when it stopped, why it stopped and how the team responded. When a failure is logged after the shift with a generic cause and a reconstructed timeline, MTBF, MTTR and availability stop guiding decisions and start reflecting reporting lag. The dashboard can look organized while it measures a rebuilt version of the operation.

Minimum elements of a reliable failure record

  • Correct technical object: the equipment or functional location where the failure occurred, not where it was easiest to open the order.
  • Work order tied to the event: scope, labor, materials, times and status carried by the order that executed the repair.
  • Cause recorded consistently: object part, damage and cause from the notification catalogs, which is what makes recurrence analysis possible. ISO 14224 is the usual reference for separating these categories.
  • Traceable downtime: malfunction start and end with the breakdown indicator, coherent with the time the asset really stopped.
  • Related preventive plan: when a failure hits an asset covered by a plan, the record must feed the review of that plan.
  • Measuring points: readings outside the limit recorded as measurement documents, not remembered later.

A reliability routine that does not overload planning

Discipline here means capturing the right data at the moment of execution and keeping the routine light for whoever executes and useful for whoever plans:

  • Standardize field reporting so that the order captures asset, symptom, cause, times, material, stop condition and evidence while the work happens.
  • Separate mandatory data from supporting detail. The technician fills in the minimum needed for MTBF, MTTR and recurrence, not a long report on every order.
  • Organize the analysis by work center and criticality, so planning sees where recurrence consumes capacity and not only how many orders were closed.
  • Hold a short review cadence in which repeated failures feed plan review, preventive scope, material needs and scheduling priority.
  • Remove parallel compilations. When the indicator depends on copying data between list transactions, spreadsheets and a dashboard assembled afterwards, planning spends its time closing numbers instead of acting on causes.

Why reliability matters in maintenance

Poor reliability becomes downtime cost long before it becomes a KPI. It shows up first as a stopped line, an emergency order that jumps the schedule, overtime and a planner rebuilding the day. Only later does it appear as a lower MTBF or a drop in availability. A failure that returns before its root cause is treated takes predictability away from everyone: the technician works under pressure, the supervisor chases field feedback and planning replaces scheduling with containment.

Completing preventive work on the calendar is not the same as being reliable

Completing the preventive plan on time helps, but it does not sustain reliability when execution becomes a mechanical routine without evidence. A closed order without proper confirmation looks like backlog reduction and still leaves open questions: what condition was the asset in, what caused the failure, how long did the intervention take, and does the plan need revision? Increasing preventive frequency without fixing that flow only creates more tasks, and every extra intervention carries its own risk of assembly error, as the bathtub curve shows for the early-failure phase.

Reliability takes shape when there is traceability between maintenance plan, order, notification, equipment, functional location, measuring point, work center and confirmation. With that chain in place, the team can tell a random failure from a repeated failure with the same operating pattern, and five conditions become visible:

  • a maintenance plan aligned with the failure mode, not only with the interval;
  • compliant execution, performed according to procedure, priority and asset condition;
  • a technical history that connects symptom, cause, intervention and confirmation;
  • consistent measurements, recorded at the right point and close to execution;
  • root cause analysis that stops the same problem from returning under another order number.

Data quality is where most programs stall. At Citrosuco, centralizing execution in SAP brought work order confirmation to 95% in the evaluated plant, and that level of confirmation is what makes the history usable for reliability analysis.

Reliability decisions are only as good as the confirmations behind them. See how PM Run delivers MTTR and MTBF data per equipment and class, ready for your own reliability panel.

Frequently asked questions about maintenance reliability

What does reliability engineering do?

It studies asset failure patterns (history, failure modes, distribution over time) and uses that knowledge to define and revise maintenance strategy, prioritizing the equipment that hurts most in downtime and cost.

What is the difference between reliability and availability?

Reliability is the probability of not failing during a period; availability is the fraction of time fit to operate, resulting from reliability (spacing failures) combined with maintainability (repairing fast).

What is maintainability?

It is how easily an asset returns to operation: physical access, diagnosis, component standardization, procedures and spare parts logistics. Improving it reduces MTTR without touching the failure rate.

Does a high MTBF mean good reliability?

It is a good sign, but the mean hides patterns: two assets with the same MTBF can have completely different failure distributions, calling for different strategies. Reliability engineering looks at the distribution, not just the mean.

What is a maintenance reliability engineer?

It is the engineer who owns failure analysis and maintenance strategy for a group of assets: ranks the worst offenders, leads root cause analysis, reviews plans and frequencies with the planner and checks by MTBF and recurrence whether the changes worked. In smaller plants, the maintenance engineer covers the role part of the time.

PM Run

Built for SAP.Not just adapted. Native.

PM Run connects planning, field execution, and supervision through native SAP integration, with no parallel spreadsheets, no re-entry at end of shift, and no data loss.

Used by leading operations in their sectors

Logo Volkswagen
Logo Eurofarma
Logo Saint-Gobain
Logo Marcopolo
Logo Moura
Logo Alpargatas

Back to blog