RCFA is a structured analysis that investigates a significant failure, identifies causal factors and controllable root causes, and carries actions through effectiveness verification. The acronym means Root Cause Failure Analysis. In industrial maintenance, its value is preventing a component replacement, the strongest opinion or the first plausible explanation from ending the investigation.
RCFA requires more than a meeting after downtime or a one-word root-cause field. The process starts by preserving evidence, reconstructs the event, tests hypotheses, identifies barrier and control deficiencies, defines proportionate actions and confirms whether recurrence and risk actually change. The analysis presents the full flow through a completed didactic case.
When to open an RCFA
Not every occurrence needs the same depth. The trigger should follow consequence and recurrence rather than the position of the requester. A plant may open RCFA for safety or environmental events, relevant production loss, material damage, recurrence, failure of a critical asset, quality deviation, unexplained degradation or a previous action that failed.
The rule must exist before the event. Without a criterion, similar occurrences receive different treatment. A simple screening can record consequence, criticality, repetition, uncertainty and future exposure. The technical owner decides depth and team. Administrative closure of the work order must not erase the need for investigation.
RCA, RCFA and analytical tools
Root cause analysis, RCA, is the broad category. RCFA applies that reasoning to a failure and requires a complete technical deliverable. The Five Whys may explore a relatively linear chain. FTA decomposes combinations leading to a top event. FMEA anticipates failure modes and effects. No tool establishes cause by itself. The team selects one according to question, evidence and complexity.
A defensible RCFA separates three levels. Physical cause describes the mechanism that produced loss of function. Human and task factors explain local decisions and conditions without blaming the person. Organizational causes explain why controls, criteria, resources or governance allowed the event. Actions must correspond to the level identified.
A nine-step method
- Classify and stabilize. Protect people, process, environment and asset before investigating.
- Preserve evidence. Retain condition, parts, permitted photographs, trends, parameters, documents and interviews without altering originals.
- Define event and scope. State loss of function, consequence, boundary, period and questions the analysis must answer.
- Build the team. Include operations, maintenance, engineering, process and system knowledge, with a facilitator and decision owner.
- Construct the timeline. Order confirmed facts, sources and gaps. Keep interpretation separate from the event.
- Analyze barriers and change. Examine what should prevent, detect or mitigate and what changed.
- Test hypotheses. Seek supporting and contradicting evidence. An untested hypothesis remains a hypothesis.
- Define causes and actions. Link every action to a cause, owner, due date and completion criterion.
- Verify effectiveness. Observe a sufficient window and reopen analysis when the control does not deliver the expected result.
Worked case: conveyor CV-17
Every asset, number and fact below is didactic. It is not a customer result, benchmark or PM Run performance claim. Conveyor CV-17 feeds raw material into a continuous process stage. On August 4, 2025, the drive shut down on high temperature at the coupled-side bearing. The line was unavailable for 3 hours and 40 minutes. No injury, environmental impact or fire occurred.
Definition and questions
| Element | Case definition |
|---|---|
| Event | Loss of conveying function after a high-bearing-temperature trip |
| Consequence | 3 h 40 min of line unavailability |
| Scope | Bearing, lubrication, alignment, condition monitoring, anomaly response and recent history |
| Outside scope | Full line redesign and economic replacement of the conveyor |
| Question 1 | Which mechanism led to temperature rise and damage? |
| Question 2 | Why did earlier signals not lead to effective intervention? |
| Question 3 | Which controls must change to reduce recurrence? |
Preserved evidence
- Temperature and vibration history for the previous 90 days, exported with date and source.
- Notifications, orders, confirmations and attachments from the prior 12 months.
- Bearing, lubricant residue, guard, base and coupling condition before cleaning.
- Removed parts identified and retained for inspection.
- Lubrication instruction, task list and monitoring criteria effective on the event date.
- Separate interviews with operator, inspector, planner and two technicians.
- Alignment record from the previous replacement, found without an attached acceptance result.
Case vibration values rose from 4.2 mm/s to 8.6 mm/s over six weeks. Temperature relative to a comparable bearing increased by 19 °C. These are didactic data and not universal limits. The plant had an internal alert criterion, but the document did not state who opened a notification, within which time or with which fields.
Factual timeline
| Date | Confirmed fact | Source |
|---|---|---|
| Jan 12, 2025 | Bearing replaced after noise; order closed without an alignment result | Order and interview |
| Jun 18, 2025 | First persistent vibration increase above the local baseline | Archived trend |
| Jul 2, 2025 | Route recorded an attention condition in an external worksheet | Route file |
| Jul 16, 2025 | New reading confirmed growth; no notification was located | Trend and SAP search |
| Jul 31, 2025 | Lubrication performed; order only said lubricate, without quantity or initial condition | Order and task list |
| Aug 4, 2025 | Temperature protection tripped, line stopped and bearing was replaced | Historian and emergency order |
Expected barriers and performance
| Barrier | Expected function | Observed result |
|---|---|---|
| Post-work alignment | Confirm assembly condition | No acceptance evidence at completion |
| Condition route | Detect degradation and trigger treatment | Detected trend but remained outside notification flow |
| Lubrication plan | Apply defined product, quantity and method | Product identified, quantity and purge criterion absent |
| Planning review | Convert a valid anomaly into prioritized work | No item reached the formal backlog |
| Temperature protection | Mitigate damage and consequence | Tripped and limited the consequence |
Hypotheses tested
The team considered process overload, misalignment, inadequate lubrication, incorrect assembly, component defect and false measurement. Load remained within the declared didactic range. The protection sensor was compared with a reference instrument and did not explain the event. Part inspection found a pattern compatible with accelerated degradation under combined alignment and lubrication conditions, but the component alone did not prove the organizational chain.
The incomplete previous replacement record prevented confirmation of initial condition. The lubrication instruction did not state quantity or response to the observed condition. The route detected growth but used an external worksheet and had no conversion rule from alert to notification. The three gaps were independent and converged on missing control of the cycle from detection through decision and verification.
Causal factors and controllable root cause
- Most likely physical cause: bearing degradation under combined unverified alignment and lubrication without a complete parameter.
- Causal factor 1: closure of the prior order did not require alignment evidence.
- Causal factor 2: lubrication task did not include quantity, method and response criteria.
- Causal factor 3: condition alert lacked an owner, response time and integration with formal backlog.
- Controllable root cause: maintenance-strategy governance did not define acceptance criteria and a closed loop among intervention, monitoring and SAP PM for this asset family.
The wording was tested with a practical question: if the root cause is corrected, will the organization reduce the likelihood of this and similar events? The answer was yes, provided all three barriers were addressed together. Replacing the bearing again would restore immediate condition but not the system that allowed degradation to advance.
Actions linked to causes
| Didactic action | Owner | Due | Completion evidence |
|---|---|---|---|
| Revise job plan with alignment and acceptance test | Maintenance engineering | 30 days | Approved list applied to a pilot order |
| Define lubricant, quantity, method and response | Engineering and lubrication | 30 days | Revised plan and qualified performer |
| Create alert-to-notification rule | Reliability and planning | 15 days | Criterion, owner, response time and simulated test |
| Review seven equivalent assets | Planning | 45 days | Orders or recorded justification per asset |
| Audit completion evidence | Supervision | 90 days | Monthly sample with action on deviations |
Train the team was rejected as a broad action that did not correspond to the causes by itself. Training enters only when expected behavior, revised material and a competence check exist. Be more careful was also rejected because it changes neither barrier nor process.
Path into SAP PM
In the case, the occurrence receives a notification with technical object, malfunction start, symptom, failure mode, effect, detection method and permitted evidence according to configuration. The emergency order preserves actual operations and materials. The RCFA is referenced in history. Physical actions and plan revisions create their own orders instead of remaining as report text.
SAP documentation on failure data describes failure mode, effect and detection method in notifications. Those fields can support RCM and FMEA when configured and governed. The system does not confirm cause automatically. A well-used catalog organizes evidence; technical investigation decides what it supports.
PM Run can support execution, mobility, planning and return of data over SAP PM. It does not perform RCFA, diagnose failure or create an engineering action on its own. SAP remains the system of record while the operation defines criteria and approves change.
Effectiveness indicators and window
The team set 180 days or three complete inspection cycles, whichever was longer. A rare event would not be validated only by the absence of a new trip. The review followed:
- percentage of valid alerts converted into notifications within the internal response time;
- percentage of family orders with alignment and test evidence;
- adherence to the revised lubrication task;
- vibration and temperature trend against the local baseline;
- recurrence of the family within 30, 90 and 180 days;
- overdue actions and actions closed without evidence;
- similar failures across seven equivalent assets.
Didactic criteria were 100% treatment of critical alerts under the rule, at least 95% complete order evidence and no repeat linked to the treated barriers. These values are not benchmarks. If the asset operates infrequently, the window must consider exposure rather than calendar alone.
Implementation result
| Implemented control | Milestone | Observed didactic evidence | Status |
|---|---|---|---|
| Alert-to-notification rule | D+12 | 4 of 4 simulated scenarios created a notification within 24 h | Complete |
| Task list with alignment and acceptance | D+28 | 1 of 1 pilot order recorded values before and after assembly | Complete |
| Revised lubrication task | D+28 | Product, quantity, method, and purge defined; 3 of 3 performers qualified | Complete |
| Review of equivalent assets | D+43 | 7 of 7 assets reviewed; two orders opened for condition gaps | Complete |
| Completion evidence audit | D+90 | Three monthly samples completed; deviations treated before closure | Complete |
All five planned actions were implemented within their didactic due dates, with completion evidence accepted for 5 of 5 actions by the accountable roles. The original notification, emergency order, action orders, and revised task-list version remained linked in SAP PM history.
Observed effectiveness-window result
Observation ran from August 5, 2025 through January 31, 2026, totaling 180 days and six complete inspection cycles. The population included CV-17 and seven equivalent assets. The physical ceiling was 8 assets × 180 days × 24 h = 34,560 calendar hours. After excluding 5,120 hours of planned outage, standby, and noncomparable load, 29,440 comparable operating hours remained.
| Indicator | Numerator | Denominator | Result | Didactic criterion |
|---|---|---|---|---|
| Critical alerts treated within the response time | 9 | 9 valid alerts | 100% | 100% within 24 h |
| Orders with complete evidence | 20 | 20 applicable orders | 100% | At least 95% |
| Executions adhering to the lubrication task | 18 | 18 applicable executions | 100% | 100% |
| Recurrence linked to the treated barriers | 0 | 29,440 comparable h | 0.000 per 1,000 h | No repeat |
| Critical actions with accepted evidence | 5 | 5 planned actions | 100% | 100% |
CV-17 completed six cycles in the defined load band. The highest recorded vibration was 4.9 mm/s, and the largest temperature difference from the comparable bearing was 5 °C, using the same didactic method that had recorded 8.6 mm/s and 19 °C before the event. These values demonstrate the case comparison and do not establish universal limits.
At the D+180 review, the team technically closed the RCFA because all five actions were implemented, all three effectiveness criteria were met, and no recurrence associated with the treated barriers appeared in the defined exposure. Monthly monitoring remained in the plan. The analysis must reopen if a valid alert exceeds 24 h without a notification, complete evidence falls below 95% in the rolling sample, or the same mechanism recurs. Closure validates control of the barriers in this case without claiming that every bearing failure has been eliminated.
Limits and safeguards
- Temporal correlation does not prove a physical mechanism.
- A damaged part may lose evidence of its initial condition after the event.
- Interviews are important sources and must be checked against records.
- No recurrence in a short window does not prove effectiveness.
- One root cause may require several actions and change control.
- Blaming the last person in the chain ends analysis before organizational control.
Evidence preservation and confidence levels
The CV-17 team assigned a confidence level to each causal statement. Direct physical evidence, verified control history, and repeatable tests received the highest weight. Interviews and operator recollections were preserved with time and role, but they were not treated as equivalent to an instrument record. Missing evidence remained visible in the analysis.
Before repair changed the scene, the coordinator defined what had to be photographed, measured, tagged, and retained. Bearing orientation, lubrication condition, damaged surfaces, fastener state, and protection trips were documented under plant rules. Chain of custody was used where material would be sent for laboratory examination.
The team also recorded disconfirming evidence. The drive current had remained within its operating band, which weakened the hypothesis of a persistent overload. Alignment after controlled disassembly was outside the internal criterion, which supported a mechanical contributor but did not prove the complete chain by itself. This discipline prevented the most vivid clue from becoming the conclusion.
Verification ownership
Each action received an owner for implementation and another role for effectiveness verification. The planner could confirm that the task list changed, while reliability engineering checked whether lubrication condition and bearing recurrence improved. Operations verified the revised startup condition. Closing the action required evidence from the defined window rather than a completion date alone.
Technical references
- U.S. Department of Energy, DOE-NE-STD-1004-92, Root Cause Analysis Guidance Document, accessed August 23, 2026.
- SAP Help Portal, Failure Data, accessed August 23, 2026.
- SAP Help Portal, Process Maintenance Notification, accessed August 23, 2026.
Frequently asked questions
What is the difference between RCA and RCFA?
RCA is the broad root-cause-analysis category. RCFA is a structured application to a failure, with an event, evidence, causal factors, controllable causes, actions and verification. Organizations use terms differently, so the internal procedure should state the expected deliverable.
Does RCFA always find one root cause?
No. Industrial systems can require a combination of causes and barriers. Forcing one cause oversimplifies. The criterion is to identify controllable causes that explain the evidence and whose correction reduces recurrence or consequence.
When is an RCFA closed?
The report may be approved when evidence, causes and actions are coherent. Technical closure comes only after critical actions are implemented and the effectiveness window is observed. If indicators do not change, the analysis must reopen.
To connect work orders, planning and field feedback to SAP PM, learn about the PM Run operational layer. Investigation, diagnosis and action approval remain the technical responsibility of the company.
