How to evaluate the accuracy of AI root cause analysis
What to measure, why abstention matters, and the replay method
The accuracy of AI root cause analysis cannot be judged from a demo, a vendor's published figure or a benchmark on someone else's incidents. It is judged by running the system on your own historical Sev1s, blind, and scoring the output against what your engineers actually found, while giving credit for honest abstention. This guide sets out what to measure and how to run that evaluation.
A confidently wrong root cause is worse than none. Teams act on it. Any evaluation that scores only hit rate, and not the rate and quality of abstention, will favour the system that guesses most confidently.
Why are vendor accuracy numbers not comparable?
Published accuracy figures for AI root cause analysis are measured on different incident sets, with different definitions of correct, different telemetry coverage and different amounts of human help. A figure computed on a vendor's curated set, or on a cloud-native estate with full tracing, says little about a hybrid estate with a mainframe, an Oracle database and a stale CMDB. This guide does not cite any competitor's accuracy figures for that reason. The only number that predicts your outcome is the one measured on your incidents.
What should you measure?
Score every historical incident on all of these, not just the first.
Cause accuracy
Does the identified root cause match what the human investigation established? Score as match, partial (right area, wrong specific), or miss. Decide the rubric before you see results.
Abstention rate and quality
How often does the system say it cannot determine the cause? When it abstains, is the stated missing evidence real? Correct abstention on a genuinely undeterminable incident is a pass, not a fail.
Confident-wrong rate
The share of incidents where the system stated a cause with high confidence and was wrong. This is the most important number and the one to weight hardest.
Time to cause
Elapsed time from incident data being available to a stated cause, compared with how long the human investigation took.
Evidence quality
Does every claim in the output link to real telemetry, log lines or change records that a reviewer can open? Prose without evidence scores zero regardless of accuracy.
Change attribution and ownership
When the cause was a change, did the system identify the specific change? Did it resolve the owning team correctly?
How do you run the replay?
The replay bake-off is the evaluation method: the system works your last 90 days of Sev1s blind and is scored against the human record. It replaces the demo.
- 1
Select the incident set
All Sev1s and Sev2s from the last 90 days, not a hand-picked subset. Include the ones your team never fully explained; those test abstention.
- 2
Fix the rubric before you start
Define match, partial and miss for cause accuracy. Define what counts as a valid abstention. Decide how confident-wrong is weighted. Write it down so it cannot drift once results arrive.
- 3
Withhold the human RCA
The system gets the incident ticket, the time window and access to the telemetry, change and topology data as it existed. It does not get the postmortem, the chat transcript conclusions or the human RCA.
- 4
Run inside your own environment
If the system needs your production telemetry to leave your perimeter to run the evaluation, that is a finding about the deployment model before it is a finding about accuracy. Prefer an evaluation that runs in your VPC.
- 5
Score blind, then compare
Have engineers who ran the original incidents score each output against the rubric without knowing which incidents the system flagged as high confidence. Then compare accuracy, abstention, confident-wrong, time to cause and evidence quality.
- 6
Read the abstentions carefully
Each abstention names missing evidence. Check whether the gap is real. Where it is, you have found a telemetry coverage gap that also caps your human investigators; where it is not, you have found a system weakness.
- 7
Ask what the SLA would be
The replay tells you the accuracy and time to cause the system can sustain on your estate. A credible provider will set the Time to RCA and abstention terms in the contract from this result, not from a slide.
What does good look like?
A system that matches or beats the human record on the well-instrumented incidents, abstains on the ones your team also could not explain, states the missing evidence accurately, links every claim to something a reviewer can open, and has a confident-wrong rate close to zero. That is a system a sceptical SRE lead can trust and a regulated buyer can accept with human sign-off. Raw hit rate on its own is not the goal.
The 90-day replay is Onepane's only evaluation.
Onepane does not run demos. Send your last 90 days of Sev1 tickets; the replay runs inside your VPC and returns a Replay Bake-off Report scoring accuracy, abstention and time to cause against what your engineers actually wrote. Accuracy, abstention rate and Time to RCA are then published monthly. Start at /replay.
The bottom line: Evaluate AI root cause analysis on your own historical incidents, blind, with a rubric fixed in advance, and weight confident-wrong hardest. Abstention with accurately named missing evidence is a pass. The replay result, not a vendor figure, is what should set the SLA.
Onepane's 90-day replay runs blind on your own Sev1s inside your VPC and scores accuracy, abstention and time to cause against the human record; that result sets the SLA.