Back
Manutenção Industrial

RCFA: from failure evidence to verified corrective action

P
PM Run Team
August 23, 2026

RCFA is a structured analysis that investigates a significant failure, identifies causal factors and controllable root causes, and carries actions through effectiveness verification. The acronym means Root Cause Failure Analysis. In industrial maintenance, its value is preventing a component replacement, the strongest opinion or the first plausible explanation from ending the investigation.

RCFA requires more than a meeting after downtime or a one-word root-cause field. The process starts by preserving evidence, reconstructs the event, tests hypotheses, identifies barrier and control deficiencies, defines proportionate actions and confirms whether recurrence and risk actually change. The analysis presents the full flow through a completed didactic case.

When to open an RCFA

Not every occurrence needs the same depth. The trigger should follow consequence and recurrence rather than the position of the requester. A plant may open RCFA for safety or environmental events, relevant production loss, material damage, recurrence, failure of a critical asset, quality deviation, unexplained degradation or a previous action that failed.

The rule must exist before the event. Without a criterion, similar occurrences receive different treatment. A simple screening can record consequence, criticality, repetition, uncertainty and future exposure. The technical owner decides depth and team. Administrative closure of the work order must not erase the need for investigation.

RCA, RCFA and analytical tools

Root cause analysis, RCA, is the broad category. RCFA applies that reasoning to a failure and requires a complete technical deliverable. The Five Whys may explore a relatively linear chain. FTA decomposes combinations leading to a top event. FMEA anticipates failure modes and effects. No tool establishes cause by itself. The team selects one according to question, evidence and complexity.

A defensible RCFA separates three levels. Physical cause describes the mechanism that produced loss of function. Human and task factors explain local decisions and conditions without blaming the person. Organizational causes explain why controls, criteria, resources or governance allowed the event. Actions must correspond to the level identified.

A nine-step method

  1. Classify and stabilize. Protect people, process, environment and asset before investigating.
  2. Preserve evidence. Retain condition, parts, permitted photographs, trends, parameters, documents and interviews without altering originals.
  3. Define event and scope. State loss of function, consequence, boundary, period and questions the analysis must answer.
  4. Build the team. Include operations, maintenance, engineering, process and system knowledge, with a facilitator and decision owner.
  5. Construct the timeline. Order confirmed facts, sources and gaps. Keep interpretation separate from the event.
  6. Analyze barriers and change. Examine what should prevent, detect or mitigate and what changed.
  7. Test hypotheses. Seek supporting and contradicting evidence. An untested hypothesis remains a hypothesis.
  8. Define causes and actions. Link every action to a cause, owner, due date and completion criterion.
  9. Verify effectiveness. Observe a sufficient window and reopen analysis when the control does not deliver the expected result.

Worked case: conveyor CV-17

Every asset, number and fact below is didactic. It is not a customer result, benchmark or PM Run performance claim. Conveyor CV-17 feeds raw material into a continuous process stage. On August 4, 2025, the drive shut down on high temperature at the coupled-side bearing. The line was unavailable for 3 hours and 40 minutes. No injury, environmental impact or fire occurred.

Definition and questions

ElementCase definition
EventLoss of conveying function after a high-bearing-temperature trip
Consequence3 h 40 min of line unavailability
ScopeBearing, lubrication, alignment, condition monitoring, anomaly response and recent history
Outside scopeFull line redesign and economic replacement of the conveyor
Question 1Which mechanism led to temperature rise and damage?
Question 2Why did earlier signals not lead to effective intervention?
Question 3Which controls must change to reduce recurrence?

Preserved evidence

  • Temperature and vibration history for the previous 90 days, exported with date and source.
  • Notifications, orders, confirmations and attachments from the prior 12 months.
  • Bearing, lubricant residue, guard, base and coupling condition before cleaning.
  • Removed parts identified and retained for inspection.
  • Lubrication instruction, task list and monitoring criteria effective on the event date.
  • Separate interviews with operator, inspector, planner and two technicians.
  • Alignment record from the previous replacement, found without an attached acceptance result.

Case vibration values rose from 4.2 mm/s to 8.6 mm/s over six weeks. Temperature relative to a comparable bearing increased by 19 °C. These are didactic data and not universal limits. The plant had an internal alert criterion, but the document did not state who opened a notification, within which time or with which fields.

Factual timeline

DateConfirmed factSource
Jan 12, 2025Bearing replaced after noise; order closed without an alignment resultOrder and interview
Jun 18, 2025First persistent vibration increase above the local baselineArchived trend
Jul 2, 2025Route recorded an attention condition in an external worksheetRoute file
Jul 16, 2025New reading confirmed growth; no notification was locatedTrend and SAP search
Jul 31, 2025Lubrication performed; order only said lubricate, without quantity or initial conditionOrder and task list
Aug 4, 2025Temperature protection tripped, line stopped and bearing was replacedHistorian and emergency order

Expected barriers and performance

BarrierExpected functionObserved result
Post-work alignmentConfirm assembly conditionNo acceptance evidence at completion
Condition routeDetect degradation and trigger treatmentDetected trend but remained outside notification flow
Lubrication planApply defined product, quantity and methodProduct identified, quantity and purge criterion absent
Planning reviewConvert a valid anomaly into prioritized workNo item reached the formal backlog
Temperature protectionMitigate damage and consequenceTripped and limited the consequence

Hypotheses tested

The team considered process overload, misalignment, inadequate lubrication, incorrect assembly, component defect and false measurement. Load remained within the declared didactic range. The protection sensor was compared with a reference instrument and did not explain the event. Part inspection found a pattern compatible with accelerated degradation under combined alignment and lubrication conditions, but the component alone did not prove the organizational chain.

The incomplete previous replacement record prevented confirmation of initial condition. The lubrication instruction did not state quantity or response to the observed condition. The route detected growth but used an external worksheet and had no conversion rule from alert to notification. The three gaps were independent and converged on missing control of the cycle from detection through decision and verification.

Causal factors and controllable root cause

  • Most likely physical cause: bearing degradation under combined unverified alignment and lubrication without a complete parameter.
  • Causal factor 1: closure of the prior order did not require alignment evidence.
  • Causal factor 2: lubrication task did not include quantity, method and response criteria.
  • Causal factor 3: condition alert lacked an owner, response time and integration with formal backlog.
  • Controllable root cause: maintenance-strategy governance did not define acceptance criteria and a closed loop among intervention, monitoring and SAP PM for this asset family.

The wording was tested with a practical question: if the root cause is corrected, will the organization reduce the likelihood of this and similar events? The answer was yes, provided all three barriers were addressed together. Replacing the bearing again would restore immediate condition but not the system that allowed degradation to advance.

Actions linked to causes

Didactic actionOwnerDueCompletion evidence
Revise job plan with alignment and acceptance testMaintenance engineering30 daysApproved list applied to a pilot order
Define lubricant, quantity, method and responseEngineering and lubrication30 daysRevised plan and qualified performer
Create alert-to-notification ruleReliability and planning15 daysCriterion, owner, response time and simulated test
Review seven equivalent assetsPlanning45 daysOrders or recorded justification per asset
Audit completion evidenceSupervision90 daysMonthly sample with action on deviations

Train the team was rejected as a broad action that did not correspond to the causes by itself. Training enters only when expected behavior, revised material and a competence check exist. Be more careful was also rejected because it changes neither barrier nor process.

Path into SAP PM

In the case, the occurrence receives a notification with technical object, malfunction start, symptom, failure mode, effect, detection method and permitted evidence according to configuration. The emergency order preserves actual operations and materials. The RCFA is referenced in history. Physical actions and plan revisions create their own orders instead of remaining as report text.

SAP documentation on failure data describes failure mode, effect and detection method in notifications. Those fields can support RCM and FMEA when configured and governed. The system does not confirm cause automatically. A well-used catalog organizes evidence; technical investigation decides what it supports.

PM Run can support execution, mobility, planning and return of data over SAP PM. It does not perform RCFA, diagnose failure or create an engineering action on its own. SAP remains the system of record while the operation defines criteria and approves change.

Effectiveness indicators and window

The team set 180 days or three complete inspection cycles, whichever was longer. A rare event would not be validated only by the absence of a new trip. The review followed:

  • percentage of valid alerts converted into notifications within the internal response time;
  • percentage of family orders with alignment and test evidence;
  • adherence to the revised lubrication task;
  • vibration and temperature trend against the local baseline;
  • recurrence of the family within 30, 90 and 180 days;
  • overdue actions and actions closed without evidence;
  • similar failures across seven equivalent assets.

Didactic criteria were 100% treatment of critical alerts under the rule, at least 95% complete order evidence and no repeat linked to the treated barriers. These values are not benchmarks. If the asset operates infrequently, the window must consider exposure rather than calendar alone.

Implementation result

Implemented controlMilestoneObserved didactic evidenceStatus
Alert-to-notification ruleD+124 of 4 simulated scenarios created a notification within 24 hComplete
Task list with alignment and acceptanceD+281 of 1 pilot order recorded values before and after assemblyComplete
Revised lubrication taskD+28Product, quantity, method, and purge defined; 3 of 3 performers qualifiedComplete
Review of equivalent assetsD+437 of 7 assets reviewed; two orders opened for condition gapsComplete
Completion evidence auditD+90Three monthly samples completed; deviations treated before closureComplete

All five planned actions were implemented within their didactic due dates, with completion evidence accepted for 5 of 5 actions by the accountable roles. The original notification, emergency order, action orders, and revised task-list version remained linked in SAP PM history.

Observed effectiveness-window result

Observation ran from August 5, 2025 through January 31, 2026, totaling 180 days and six complete inspection cycles. The population included CV-17 and seven equivalent assets. The physical ceiling was 8 assets × 180 days × 24 h = 34,560 calendar hours. After excluding 5,120 hours of planned outage, standby, and noncomparable load, 29,440 comparable operating hours remained.

IndicatorNumeratorDenominatorResultDidactic criterion
Critical alerts treated within the response time99 valid alerts100%100% within 24 h
Orders with complete evidence2020 applicable orders100%At least 95%
Executions adhering to the lubrication task1818 applicable executions100%100%
Recurrence linked to the treated barriers029,440 comparable h0.000 per 1,000 hNo repeat
Critical actions with accepted evidence55 planned actions100%100%

CV-17 completed six cycles in the defined load band. The highest recorded vibration was 4.9 mm/s, and the largest temperature difference from the comparable bearing was 5 °C, using the same didactic method that had recorded 8.6 mm/s and 19 °C before the event. These values demonstrate the case comparison and do not establish universal limits.

At the D+180 review, the team technically closed the RCFA because all five actions were implemented, all three effectiveness criteria were met, and no recurrence associated with the treated barriers appeared in the defined exposure. Monthly monitoring remained in the plan. The analysis must reopen if a valid alert exceeds 24 h without a notification, complete evidence falls below 95% in the rolling sample, or the same mechanism recurs. Closure validates control of the barriers in this case without claiming that every bearing failure has been eliminated.

Limits and safeguards

  • Temporal correlation does not prove a physical mechanism.
  • A damaged part may lose evidence of its initial condition after the event.
  • Interviews are important sources and must be checked against records.
  • No recurrence in a short window does not prove effectiveness.
  • One root cause may require several actions and change control.
  • Blaming the last person in the chain ends analysis before organizational control.

Evidence preservation and confidence levels

The CV-17 team assigned a confidence level to each causal statement. Direct physical evidence, verified control history, and repeatable tests received the highest weight. Interviews and operator recollections were preserved with time and role, but they were not treated as equivalent to an instrument record. Missing evidence remained visible in the analysis.

Before repair changed the scene, the coordinator defined what had to be photographed, measured, tagged, and retained. Bearing orientation, lubrication condition, damaged surfaces, fastener state, and protection trips were documented under plant rules. Chain of custody was used where material would be sent for laboratory examination.

The team also recorded disconfirming evidence. The drive current had remained within its operating band, which weakened the hypothesis of a persistent overload. Alignment after controlled disassembly was outside the internal criterion, which supported a mechanical contributor but did not prove the complete chain by itself. This discipline prevented the most vivid clue from becoming the conclusion.

Verification ownership

Each action received an owner for implementation and another role for effectiveness verification. The planner could confirm that the task list changed, while reliability engineering checked whether lubrication condition and bearing recurrence improved. Operations verified the revised startup condition. Closing the action required evidence from the defined window rather than a completion date alone.

Technical references

Frequently asked questions

What is the difference between RCA and RCFA?

RCA is the broad root-cause-analysis category. RCFA is a structured application to a failure, with an event, evidence, causal factors, controllable causes, actions and verification. Organizations use terms differently, so the internal procedure should state the expected deliverable.

Does RCFA always find one root cause?

No. Industrial systems can require a combination of causes and barriers. Forcing one cause oversimplifies. The criterion is to identify controllable causes that explain the evidence and whose correction reduces recurrence or consequence.

When is an RCFA closed?

The report may be approved when evidence, causes and actions are coherent. Technical closure comes only after critical actions are implemented and the effectiveness window is observed. If indicators do not change, the analysis must reopen.

To connect work orders, planning and field feedback to SAP PM, learn about the PM Run operational layer. Investigation, diagnosis and action approval remain the technical responsibility of the company.

RCFA
Root cause analysis
Reliability engineering
Failure analysis
SAP PM
Industrial maintenance
PM Run

Built for SAP.Not just adapted. Native.

PM Run connects planning, field execution, and supervision through native SAP integration, with no parallel spreadsheets, no re-entry at end of shift, and no data loss.

Used by leading operations in their sectors

Logo Volkswagen
Logo Eurofarma
Logo Saint-Gobain
Logo Marcopolo
Logo Moura
Logo Alpargatas

Back to blog