Equipment failures are an unavoidable part of industrial operations, but repeated failures do not have to be. When the same pump, motor, gearbox, compressor, or production machine fails again and again, simply repairing or replacing the damaged component may only provide a temporary solution.
This is where Root Cause Analysis (RCA) becomes important.
Root Cause Analysis is a structured problem-solving approach used by maintenance and reliability teams to identify the underlying reasons behind equipment failures. Instead of asking only “What failed?”, RCA asks a more important question:
“Why did the failure happen in the first place?”
By identifying and eliminating the underlying cause, organizations can reduce repeat failures, improve equipment reliability, lower maintenance costs, and reduce unplanned downtime.
What Is Root Cause Analysis in Maintenance?
Root Cause Analysis is a systematic process used to identify the fundamental cause or causes of an equipment failure or recurring maintenance problem.
A failure may have several levels of causes.
For example:
Bearing failed → Poor lubrication → Incorrect lubrication interval → Inadequate maintenance strategy
Replacing the bearing fixes the immediate problem, but it does not necessarily prevent another failure.
RCA attempts to move deeper into the failure chain until the organization identifies a cause that can be effectively controlled or eliminated.
The goal is not to find someone to blame.
The goal is to improve the equipment and the maintenance system.
Why Root Cause Analysis Is Important
Repeated failures can consume significant maintenance resources.
Consider a pump that fails four times in one year.
Each failure may require:
- Emergency technician response
- Spare parts
- Production downtime
- Overtime
- Equipment isolation
- Troubleshooting
- Repair
- Testing
If the underlying cause is never corrected, the organization continues paying for the same problem.
RCA changes this approach.
Instead of:
Failure → Repair → Failure → Repair → Failure
the objective becomes:
Failure → Investigation → Root Cause → Corrective Action → Verification → Prevention
This can create long-term reliability improvement.
When Should Maintenance Teams Perform RCA?
Not every minor failure requires a formal RCA.
A detailed investigation is usually most valuable when a failure:
- Causes significant production downtime
- Creates a safety risk
- Causes environmental impact
- Results in expensive repairs
- Happens repeatedly
- Affects a critical asset
- Creates major quality problems
- Has an unknown cause
- Produces secondary equipment damage
Organizations should establish clear criteria for determining which failures require formal RCA.
The Difference Between Symptoms, Causes and Root Causes
One of the most important concepts in RCA is understanding the difference between a symptom, an immediate cause, and a root cause.
For example:
Symptom
The pump stopped operating.
Immediate Cause
The motor protection system tripped.
Physical Cause
The motor was drawing excessive current.
Root Cause
The pump was operating under abnormal mechanical loading caused by persistent misalignment.
If the maintenance team only resets the protection system, the problem will return.
If they replace the motor without investigating the loading condition, the new motor may eventually experience the same problem.
RCA aims to identify the deeper cause that can be addressed permanently.
Step 1: Define the Problem
The first step is to clearly describe what happened.
Avoid vague statements such as:
“Pump failed.”
A better problem statement might be:
“Pump P-101 stopped unexpectedly during normal production operation after motor current increased above the protection limit.”
A good problem statement should identify:
- What failed?
- Where did it happen?
- When did it happen?
- Under what conditions?
- What was the consequence?
A clear problem statement prevents the investigation from becoming too broad.
Step 2: Collect Evidence
RCA should be based on evidence rather than assumptions.
Maintenance engineers should collect information such as:
- Equipment history
- Previous failure reports
- Work orders
- Inspection records
- Sensor data
- Vibration measurements
- Temperature readings
- Oil-analysis results
- Operating conditions
- Maintenance procedures
- Operator observations
- Photographs
- Damaged components
The more accurate the information, the stronger the investigation.
You may also like: Net Positive Suction Head in Centrifugal Pump Example | PDF
Step 3: Reconstruct the Failure
Determine what happened before, during, and after the failure.
A timeline can be useful.
For example:
08:00 — Pump operating normally
09:30 — Operator noticed abnormal noise
10:00 — Vibration increased
10:20 — Motor current increased
10:30 — Motor protection activated
10:45 — Pump stopped
This timeline can reveal important relationships between symptoms and failure events.
Step 4: Identify Possible Causes
Once the evidence has been collected, identify potential causes.
For a motor-driven pump, possible causes could include:
- Misalignment
- Imbalance
- Bearing deterioration
- Poor lubrication
- Mechanical looseness
- Cavitation
- Overloading
- Electrical problems
- Installation errors
- Process abnormalities
Do not immediately select the first explanation that appears reasonable.
A good RCA considers multiple possibilities and uses evidence to eliminate unsupported causes.
Step 5: Use the 5 Whys Method
The 5 Whys is one of the simplest RCA techniques.
It involves repeatedly asking “Why?” until the investigation reaches a controllable underlying cause.
Example:
Problem:
Pump bearing failed.
Why?
Because the bearing overheated.
Why did it overheat?
Because lubrication was inadequate.
Why was lubrication inadequate?
Because the bearing was not receiving the required lubrication quantity.
Why was the quantity incorrect?
Because the lubrication procedure did not specify the correct amount.
Why did the procedure not specify the correct amount?
Because the maintenance standard had never been updated after the equipment modification.
The investigation has now moved beyond the failed bearing to a maintenance-system issue.
The 5 Whys should not always involve exactly five questions. The objective is to continue asking meaningful questions until the underlying cause is sufficiently understood.
Step 6: Use a Fishbone Diagram
A Fishbone Diagram, also called an Ishikawa or Cause-and-Effect Diagram, helps teams organize possible causes.
Common categories include:
- People
- Equipment
- Methods
- Materials
- Measurement
- Environment
For an industrial failure, the team might investigate:
People
Was the equipment operated or maintained correctly?
Equipment
Was there a design or component problem?
Methods
Were maintenance procedures adequate?
Materials
Were correct parts and lubricants used?
Measurement
Were condition-monitoring measurements accurate?
Environment
Did temperature, contamination, humidity, or other environmental conditions contribute?
This approach encourages teams to consider the entire system rather than focusing only on the failed component.
Step 7: Verify the Root Cause
A suspected cause is not necessarily a confirmed root cause.
Maintenance engineers should look for evidence that supports the conclusion.
For example, if misalignment is suspected, engineers may inspect:
- Shaft alignment
- Coupling condition
- Vibration spectrum
- Bearing wear
- Installation records
The evidence should support the relationship between the suspected cause and the actual failure.
A strong RCA should be able to answer:
“How do we know this is the root cause?”
Step 8: Develop Corrective Actions
Once the root cause has been confirmed, develop corrective actions.
Corrective actions should address the actual cause rather than simply replacing the failed component.
Examples include:
- Correcting shaft alignment
- Improving lubrication procedures
- Modifying equipment design
- Changing operating practices
- Updating maintenance procedures
- Improving inspection frequency
- Installing condition monitoring
- Training maintenance personnel
- Changing component specifications
The best corrective action is one that prevents recurrence.
Step 9: Assign Responsibility and Deadlines
RCA recommendations should become actionable tasks.
Each action should have:
- A responsible person
- A completion date
- Required resources
- Priority
- Verification criteria
For example:
Action: Update pump lubrication procedure.
Responsible: Maintenance Engineer.
Deadline: 30 days.
Verification: Updated procedure approved and implemented.
This prevents RCA reports from becoming documents that are completed but never acted upon.
Step 10: Verify Effectiveness
Corrective action is not complete simply because the task has been performed.
The organization should verify whether the problem has actually been eliminated.
For example:
Before corrective action:
Bearing failures: 4 per year
After corrective action:
Bearing failures: 0 over the following 12 months
This provides stronger evidence that the corrective action was effective.
Condition-monitoring data can also help verify improvement.
Common Root Cause Analysis Tools
Maintenance teams can use several RCA tools depending on the complexity of the problem.
5 Whys
Simple and effective for relatively straightforward problems.
Fishbone Diagram
Useful for brainstorming and organizing potential causes.
Fault Tree Analysis
Useful for complex systems with multiple failure pathways.
FMEA
Helps identify potential failure modes, effects, causes, and controls.
Pareto Analysis
Helps identify the small number of failure causes responsible for a large percentage of problems.
Event and Causal Factor Analysis
Useful for reconstructing complex events and identifying relationships between causes.
No single RCA tool is appropriate for every situation.
Common RCA Mistakes
Blaming Individuals
RCA should focus on system and process improvement rather than automatically blaming an operator or technician.
Stopping at the First Cause
The first explanation may be an immediate cause rather than the root cause.
Using Opinions Instead of Evidence
Conclusions should be supported by equipment data, inspection results, records, and physical evidence.
Correcting Only the Symptom
Replacing a failed part without addressing why it failed can lead to recurrence.
Creating Weak Corrective Actions
“Tell technicians to be more careful” is usually weaker than improving the procedure, design, training, or control system that allowed the failure to occur.
Failing to Follow Up
An RCA is incomplete until corrective actions have been implemented and their effectiveness verified.
How RCA Supports Reliability Engineering
Root Cause Analysis is an important part of a proactive reliability strategy.
RCA information can be used to identify:
- Recurring failure modes
- Poor equipment designs
- Maintenance strategy weaknesses
- Operating problems
- Training gaps
- Spare-parts problems
- Installation issues
This information can then be used to improve:
- Preventive maintenance
- Predictive maintenance
- Reliability-Centered Maintenance
- Equipment design
- Maintenance procedures
- Training programs
- Asset-management strategies
In this way, RCA turns individual failures into opportunities for organizational learning.










