Learn · Guide · Time to RCA

RCA SLA benchmarks: how fast should a root cause analysis be delivered?

Benchmarks, contract norms and how to set a window you can meet

Most enterprise contracts expect a written root cause analysis within a few business days of a Sev1, and most operations teams have no idea whether they meet that. This guide defines Time to RCA as a metric, sets out what MSAs and vendor SLAs commonly commit to, what a missed RCA costs, and how to set and instrument an RCA SLA your team can actually keep.

Time to RCA: the elapsed time from closure of a Sev1 incident to delivery of a signed, accepted root cause analysis. It measures the document, not the moment someone first suspected the cause. It is the number an RCA SLA tests, and almost nobody tracks it.

Why is Time to RCA the metric that matters?

MTTD, MTTI and MTTR all end before the RCA is written. A team can restore service in forty minutes, identify the cause in two hours, and still deliver the written RCA eleven days late, because the person who has to write it is also the person running the next incident. Customers, auditors and problem-review boards do not consume MTTI. They consume the document. Time to RCA measures the gap between the incident being over and the deliverable being in the recipient's hands, and it is the only one of the four that maps directly onto a contract clause.

What do RCA SLAs commonly commit to?

Figures below are ranges seen across enterprise master service agreements and publicly posted vendor support policies. They are norms, not a standard; your own contracts govern. No individual company is being cited.

Enterprise MSA, written RCA after Sev1

Commonly 3–5 business days

Two-stage contracts

24–48h preliminary, longer final

Vendor SLAs seen publicly

Roughly 3–10 business days

Internal targets, where they exist

Often undefined or unmeasured

What does a typical MSA clause actually say?

The usual shape is: for each Severity 1 incident the provider will deliver a root cause analysis to the customer within N business days of resolution, describing the cause, the impact, the remediation and the measures taken to prevent recurrence. Some contracts split this into a preliminary RCA within one or two days and a final RCA within a longer window. Some attach the RCA to the same service credit schedule as availability; others make a late RCA a separate breach. Almost none define what a complete RCA contains, which is why the same clause produces a two-paragraph email at one vendor and a twelve-page evidenced report at another. If you are the customer, define the content. If you are the vendor, define it before your customer does.

What happens when the RCA window is missed?

A late RCA is rarely dramatic on the day. The cost arrives later, and it compounds.

Service credits, where the contract ties RCA delivery to the credit schedule; often larger than the availability credit for the same incident.

A logged SLA breach that surfaces at renewal, in a vendor-risk review or in an audit finding, long after the incident is forgotten.

Customer churn risk: the RCA is the last thing the customer remembers about the outage, and a late or thin one reads as indifference.

Repeat incidents, because corrective actions were never agreed while the evidence was fresh and the same cause returns the following month.

Escalation load on customer success and account teams, who end up chasing engineering for a document they cannot write themselves.

Regulatory or examiner questions in supervised industries, where an operational-resilience review asks to see the analysis behind material outages.

How do you instrument Time to RCA?

Three timestamps per Sev1 and Sev2: incident closure, RCA first delivered, RCA accepted by the recipient. Time to RCA is closure to acceptance; the gap between delivered and accepted tells you about quality. Record the committed window per contract or per internal policy alongside each incident, and report attainment as the percentage delivered inside the window per month, split by severity. Add abstention rate, the share of incidents where the cause could not be determined from available evidence, so that a fast RCA that says nothing does not count as a win. Publish the number internally every month whether or not it flatters you. Teams that see it move it.

How do you set an RCA SLA you can actually meet?

An RCA SLA is a promise about the slowest part of your incident process. Set it from measurement, not aspiration.

  • 1

    Measure your current Time to RCA over the last 90 days

    Pull every Sev1 and Sev2, find the date the RCA was actually sent and accepted, and compute the distribution. Most teams discover a median of one to three weeks and a long tail. That is your starting point, not your target.

  • 2

    Segment by telemetry coverage

    Incidents on well-instrumented services resolve to a cause quickly; incidents on services with no traces and a stale ownership record do not. Run a readiness assessment and commit tighter windows only where the coverage supports them.

  • 3

    Define what a complete RCA contains

    Timeline, impact statement, causal chain as trigger, propagation and failure, contributing factors held distinct, an evidence index, corrective actions with owners and dates, and a confidence statement. A window is meaningless without a definition of done.

  • 4

    Choose a two-stage window if your recipients need early comfort

    A preliminary RCA at 24 to 48 hours with the known facts and current hypothesis, then a final RCA at the full window. This satisfies the customer's own escalation path without forcing a premature cause.

  • 5

    Decide who is accountable and remove the writing from the on-call engineer

    The person who ran the incident should review and sign the RCA, not draft it at 11pm. Assign drafting to a problem management function or a managed service with its own SLA, so the deadline has a single owner.

  • 6

    Attach a remedy and publish attainment

    Internally, an RCA SLA with no consequence drifts. Externally, credits make it real. Either way, report attainment monthly, alongside acceptance rate and abstention rate, so the trend is visible before a customer points it out.

How Onepane relates

Time to RCA is the SLA Onepane signs.

Onepane commits to a Time to RCA in the contract, delivered inside your VPC on the observability you already own, with engineers signing off before delivery. Attainment, acceptance rate and abstention rate are published in the monthly SLA attainment report. The 90-day replay on your own historical Sev1s is how the window gets set.

The bottom line: Time to RCA is the metric a contractual RCA obligation actually tests. Measure it, define what a complete RCA contains, segment by telemetry coverage, take the writing off the on-call engineer, and publish attainment monthly. Then commit to a window.

How Onepane helps

Onepane delivers the RCA inside a contractual Time to RCA, credits when it is missed, and publishes attainment monthly; the 90-day replay sets the window.