Article

Root Cause Analysis in Refining: Why Finding the True Cause Still Takes Too Long

Scroll

Hasma Habibiy

August 8, 2026

Overview

A crude unit starts corroding. Wash water goes up, inhibitor goes up, monitoring goes up, and six months later throughput comes down while the investigation is still open. The effort was never the problem. The true cause was sitting in a historian, a LIMS record, and an incident report from three years ago, and nobody had the time to connect them.

The Cost of Getting RCA Wrong

Every refinery has experienced it.

A crude unit begins showing elevated overhead corrosion rates. Corrosion coupons indicate increasing metal loss. Inspection reports flag areas of concern. Operations adjusts wash water rates. Corrosion inhibitors are increased. Additional monitoring is implemented.

Yet the problem continues.

Months later, the refinery is forced to reduce crude throughput to manage risk while engineering and reliability teams continue investigating the issue. Production losses accumulate, operating costs rise, and confidence in the corrective actions begins to erode.

The challenge is not a lack of effort. Most refineries have highly capable engineers, operators, inspectors, and reliability specialists. The challenge is identifying the true root cause hidden within thousands of process variables, maintenance records, laboratory results, operating logs, and historical reports.

In many cases, recurring problems persist not because organizations fail to respond, but because they respond to symptoms rather than causes.

The financial consequences can be significant. Reduced crude throughput, increased maintenance spending, accelerated equipment degradation, unplanned outages, and safety exposure can easily result in millions of dollars of lost margin annually. Finding the wrong cause is often nearly as costly as finding no cause at all.

Why Traditional RCA Struggles in Modern Plants

Root Cause Analysis has long been a cornerstone of refinery reliability programs. Methodologies such as Five Whys, Cause-and-Effect Analysis, Fault Tree Analysis, and TapRooT® provide valuable frameworks for investigation.

However, the complexity of modern refining operations creates new challenges.

A typical refinery investigation requires information from numerous systems:

  • Process historians

  • Laboratory Information Management Systems (LIMS)

  • Computerized Maintenance Management Systems (CMMS)

  • Inspection databases

  • Reliability systems

  • Incident reports

  • Shift logs

  • Operating procedures

  • Engineering studies

The data exists, but it is rarely connected.

As a result, investigation teams often spend the majority of their time collecting and organizing information before meaningful analysis can even begin. By the time the investigation team assembles the required evidence, critical context may already have been lost.

At the same time, a growing percentage of refinery expertise resides in the experience of senior engineers and operators approaching retirement. Many of the lessons learned from previous incidents are buried within reports, emails, spreadsheets, or simply human memory.

The result is a process that is increasingly difficult to execute consistently and efficiently.

The Missing Piece: Context

The industry often talks about having too much data.

In reality, the problem is not data volume. It is context.

Consider an overhead corrosion event in a crude unit.

A process historian may show:

  • Rising overhead pressure

  • Reduced reflux flow

  • Increasing chloride concentration

  • Falling overhead pH

Viewed independently, these appear to be separate operating conditions.

An experienced refinery engineer, however, immediately recognizes a potential cause-and-effect relationship. They understand how desalter performance, crude quality, wash water effectiveness, overhead chemistry, and corrosion mechanisms interact.

This is context.

Context connects individual data points into a coherent operational story.

Without context, investigators see trends.

With context, they see causality.

The difference between the two often determines whether an investigation successfully identifies the true root cause.

How AI Changes Root Cause Analysis

Recent advances in industrial AI are transforming how investigations can be performed.

Rather than replacing engineering expertise, AI augments it by dramatically reducing the effort required to gather, organize, and interpret information.

Step 1: Gather Information Automatically

Instead of manually searching multiple systems, AI can collect relevant information from:

  • Historian trends

  • Maintenance work orders

  • Inspection findings

  • Laboratory data

  • Operator logs

  • Previous RCA reports

  • Engineering documents

Information that previously required days of effort can be assembled in minutes.

Step 2: Build Event Timelines

One of the most challenging aspects of RCA is reconstructing what happened and when.

AI can automatically create a timeline such as:

  • Desalter efficiency begins declining

  • Chloride concentrations increase

  • Overhead pH decreases

  • Corrosion rates accelerate

  • Throughput restrictions are implemented

This provides investigators with a structured view of events rather than requiring manual reconstruction.

Step 3: Identify Causal Relationships

Traditional analytics excel at identifying correlations.

Effective RCA requires understanding cause and effect.

By incorporating process knowledge and equipment relationships, AI can evaluate how process changes may have contributed to an event rather than simply identifying variables that changed at the same time.

Step 4: Generate Investigative Hypotheses

Rather than beginning with a blank sheet of paper, engineers can start with ranked hypotheses such as:

  1. Desalter performance degradation

  2. Increased chloride content in crude feed

  3. Wash water system effectiveness loss

  4. Corrosion inhibitor underperformance

  5. Instrumentation error

The engineer remains responsible for validation, but the investigation starts from a position of insight rather than uncertainty.

Example Use Case: Crude Unit Overhead Corrosion

Imagine a refinery experiencing recurring overhead corrosion despite multiple corrective actions.

Historically, investigators would need to manually review:

  • Crude blend changes

  • Desalter performance trends

  • Chloride measurements

  • Wash water rates

  • Corrosion inhibitor injections

  • Inspection records

  • Previous incident reports

This process could take weeks.

An AI-enabled investigation could immediately identify that:

  • Corrosion rates began increasing shortly after a crude slate change

  • Desalter salt removal efficiency had gradually deteriorated

  • Chloride carryover increased over several months

  • Similar conditions existed during a previous corrosion event three years earlier

Instead of spending valuable engineering time locating information, investigators can focus on validating causes and implementing corrective actions.

The result is faster resolution, greater confidence, and improved organizational learning.

What Makes Refining RCA Different from Generic AI

Many organizations are experimenting with general-purpose AI tools.

These tools can summarize documents, answer questions, and search information.

However, refining Root Cause Analysis requires something more.

It requires an understanding of process relationships.

A refinery-specific AI understands concepts such as:

  • Desalter efficiency impacts overhead corrosion

  • Reformer severity influences gasoline octane

  • Excess furnace oxygen affects efficiency and emissions

  • Tower flooding impacts fractionation performance

  • Pump operating conditions influence cavitation risk

This process understanding allows investigations to move beyond document retrieval and toward engineering reasoning.

The difference is similar to comparing a search engine with a seasoned refinery engineer. One finds information. The other understands what the information means.

Capturing Expert Knowledge Before It Walks Out the Door

One of the greatest challenges facing the refining industry is workforce transition.

Many organizations are losing decades of operating experience as senior personnel retire.

Historically, knowledge transfer has relied on mentoring, procedures, and lessons-learned databases. While valuable, these approaches often fail to preserve the full context behind engineering decisions.

AI provides an opportunity to capture and operationalize institutional knowledge.

Past incident investigations, engineering studies, reliability assessments, and operating experiences can become accessible and searchable sources of expertise.

Instead of repeating investigations that were performed years earlier, engineers can build upon previous knowledge and accelerate decision-making.

Every investigation becomes an opportunity to strengthen the organization’s collective intelligence.

The Future: Continuous RCA

Traditional RCA is reactive.
An event occurs.
An investigation begins.
A report is generated.
Corrective actions are implemented.
The future is different.

AI enables organizations to move toward continuous root cause analysis.

Rather than waiting for failures, systems can continuously monitor operational behavior, detect abnormal patterns, identify emerging causal relationships, and generate early warnings before significant consequences occur.

The objective shifts from understanding failures after they happen to preventing them from occurring in the first place.

For refinery operators, this represents a fundamental change in how reliability and operational excellence are achieved.

Conclusion

Refineries do not suffer from a shortage of data.

They suffer from a shortage of contextual understanding across thousands of disconnected information sources.

Root Cause Analysis remains one of the most important tools for improving reliability, safety, and profitability, yet traditional approaches often require significant time and effort simply to assemble the necessary information.

AI-powered RCA offers a new path forward.

By combining process knowledge, engineering expertise, operational data, and institutional memory, organizations can dramatically reduce investigation time, improve the quality of corrective actions, and prevent recurring failures.

The goal is not to replace engineers.

The goal is to give engineers the context they need to find the true cause faster—and to ensure that critical knowledge is never lost.

In an industry where a single recurring problem can cost millions of dollars per year, finding the right cause has never been more important.

Give us your worst engineering task today

Give us your
worst engineering
task today

Give us your worst
engineering task today

Get in touch

Create a free website with Framer, the website builder loved by startups, designers and agencies.