Learn

The operations leader's guide to root cause analysis.

Root cause analysis for IT incidents is the evidenced determination of why an outage happened and what stops it recurring. This hub covers the definitions, the metrics such as Time to RCA, the ITIL vocabulary of problem management, and the practical guides for delivering an RCA inside a contractual window. Each entry leads with the answer, then the detail.

Definition

RCA as a Service

RCA as a Service is a managed root-cause service for enterprise IT operations: after a Sev1 or Sev2 incident, an outside team investigates the failure on the customer's existing observability data and delivers a finished, evidence-linked root cause analysis inside a contractual SLA. The customer buys the deliverable and the turnaround commitment, not another tool to staff. Onepane is RCA as a Service: AI agents investigate inside your VPC, engineers sign off, and pricing is per service under coverage plus accepted RCAs.

Read the full definition at What is RCA as a Service?, or see how the service runs at How it works.

Glossary

A-Z of root cause analysis terms.

A quick reference for the vocabulary of incident and problem management, RCA artifacts, the metrics that measure them, and the deployment terms that come up in a security review. Short, specific and quotable.

C

CAPA

artifact

Corrective and preventive action. A CAPA record ties each root cause to the action that fixes it and the action that stops it recurring, with an owner, a due date and closure evidence. Borrowed from quality management, CAPA is the artifact auditors ask for in healthcare and financial services when they want to know an outage was actually followed up.

Causal chain

process

The ordered sequence of events linking the root cause to the customer-visible failure, with each step supported by evidence. A complete causal chain answers the question at every hop: this happened, which caused that. Gaps in the chain are where confidence drops, and an honest RCA marks them rather than filling them with assumption.

Cause code

ITIL

A structured classification of an incident or problem's root cause, chosen from a controlled list such as change-induced, capacity, dependency failure, configuration or defect. Cause codes make root causes reportable across many incidents. They lose value when the list is too long, when engineers default to Other, or when the code is not linked to evidence.

Change attribution

process

Establishing which deployment, configuration change, infrastructure-as-code apply, feature flag or cloud-provider event is implicated in an incident, with the diff and its approval record. Change attribution is a statement of what changed, held distinct from who approved it. It is the single most useful step in most investigations because most outages follow a change.

CMDB

ITIL

Configuration management database: the ITSM record of configuration items, their attributes and their relationships. A CMDB is meant to answer what depends on what and who owns it. In practice most are stale because they are updated by hand, which is why RCA methods that require an accurate CMDB stall and why observed topology is used instead.

Confidence statement

artifact

The section of an RCA that says how sure the investigator is of the stated cause, which alternative hypotheses were considered, and what evidence was missing or ambiguous. Written in plain terms rather than as a bare percentage. The confidence statement is what lets a reviewer decide how much weight to put on the conclusion before signing.

Contributing factor

process

A condition that made an incident more likely, longer or worse without being its cause. Missing alerts, stale runbooks, an unowned service and a slow rollback are typical contributing factors. They belong in the RCA with their own corrective actions, but listing them in place of a root cause is a sign the analysis stopped early.

FAQ

Common questions.

What is RCA as a Service?

RCA as a Service is a managed root-cause service for enterprise IT operations: after a Sev1, an outside team investigates on your existing observability data and delivers a finished, evidence-linked root cause analysis inside a contractual SLA. You buy the deliverable and the turnaround, not another tool to staff.

How long should a root cause analysis take?

Enterprise contracts commonly commit to a written RCA within three to five business days of a Sev1; some split it into a preliminary RCA at 24 to 48 hours and a longer final. The metric to track is Time to RCA: elapsed time from Sev1 closure to a signed, accepted RCA.

What is the difference between alert correlation and root cause analysis?

Alert correlation groups related alerts into one incident to reduce noise. Root cause analysis explains why the incident happened, with evidence, which change is implicated, who owns the failing service and what to fix. Correlation is a step; it does not establish causation or produce a document.

Can you outsource root cause analysis?

Yes. A managed provider can connect to your existing telemetry, change and ticketing systems, investigate each qualifying incident and return an evidence-linked RCA within an agreed SLA. Your service owner still reviews and signs, and your team still owns the corrective actions. In regulated estates the provider should run inside your own VPC.