Six steps from Sev1 closure to a signed root-cause report.
Managed root cause analysis works like this: we deploy inside your VPC, connect to the tools you already own, map services to owners, run an agentic investigation on every Sev1, have engineers sign off, and deliver evidence-linked artifacts into your ITSM inside a committed SLA window. You never replace monitoring, ship logs, or staff a platform.
What managed detection did for the SOC, applied to the NOC.
Managed detection and response providers didn't sell anyone a SIEM. They sat on the customer's existing security tools, sold triaged and investigated detections against an SLA, and displaced analyst headcount, not the tooling. That business passed a billion dollars in annual revenue.
RCA as a Service does the same thing to operations: we sit on the observability you already own, and the deliverable is the root cause instead of the detection.
How does a managed RCA get delivered?
Deploy inside your VPC
A Terraform or Helm reference deployment stands Onepane up in your own cloud account, AWS, Azure or GCP. Nothing you consider production data leaves. Model inference runs inside the boundary. Our expert pod's access is scoped, time-boxed, logged and revocable by you.
- , Reference architecture + data-flow diagram before you sign
- , Your keys, your encryption, your network controls
- , What egresses (service heartbeats, aggregate metadata) disclosed in writing
Connect the observability you already own
Read-only connectors into Datadog, Splunk, CloudWatch, New Relic, Dynatrace, Prometheus, your databases, your Kubernetes and cloud control planes, your CI/CD and IaC, and your ITSM. We investigate across all of it, including the mainframe and the change ticket your APM vendor can't see.
- , No agents to install on hosts, no log shipping
- , Coexists with a failed AIOps deployment, we run on top of it
- , Telemetry Coverage Report tells you which services we can reach a confident cause on
Build the service-and-ownership map
We assemble topology and ownership from what's actually running, deployments, DNS, traffic, IAM, tags, on-call schedules, ticket history, not from a CMDB you don't trust. In a ten-thousand-resource estate, "who owns this service" is the real bottleneck; we answer it, and you keep the map.
- , Durable asset, quietly becomes the CMDB you wish you had
- , Refreshed continuously as the estate changes
- , Snapshot captured at the moment of every incident
Agents investigate every Sev1
When a Sev1 closes, investigation starts automatically. Agents pull change history (deploys, config, IaC applies, cloud-provider events), walk topology from symptom to source, query the relevant telemetry, and state the causal chain as trigger → propagation → failure. Every claim is linked to the log line, trace, metric or change record that supports it.
- , Change-aware: most outages are caused by change, so we start there
- , Contributing factors held distinct from the root cause
- , Confidence scored; abstention when evidence is insufficient
Humans sign off
Our engineers review the agent's work before anything ships. If the evidence doesn't support a confident conclusion, the output says so explicitly, "insufficient evidence, here is what's missing", rather than guessing. Human sign-off stays mandatory in regulated accounts.
- , Decision support with an audit trail, not an autonomous actor
- , You can check our work, not just trust it
- , Percentage of zero-touch RCAs reported monthly, and it must trend up honestly
Artifacts land inside your process, under SLA
The Root Cause Report, Executive Summary, Customer-Facing RCA, Evidence Pack and Change Attribution Record are delivered inside the committed window measured from Sev1 closure. The Problem Record is written as structured fields into ServiceNow, Jira Service Management or your ITSM of record, not a PDF in an inbox.
- , Miss the window and we credit, that clause is what makes this a service
- , Monthly Problem Review, Known Error register, CAPA tracker and SLA Attainment Report
- , You own everything we produce
How is a managed root-cause service priced?
Per service under coverage and per accepted RCA, never per seat, per host or per GB. It comes out of the labour you already spend on triage and RCA authoring, or out of your managed-services contract, not a new tooling line.
| Layer | What it is | Why |
|---|---|---|
| Coverage fee | Per critical service per month, connected, mapped, ownership-resolved. | Forecastable for procurement; the shape of every managed-service contract you already run. |
| Outcome pack | A committed annual bundle of accepted RCAs, with discounted overage. | Gives the outcome narrative teeth; the annual bucket keeps the number forecastable. |
| SLA and credits | RCA delivered within the committed window after Sev1 closure, or we credit. | The clause that makes this a service rather than software. |
| Onboarding | VPC deployment, estate discovery, telemetry readiness, historical replay. | Real work with real cost, scoped once. |
We don't quote before the replay. The replay tells us how much investigation and authoring labour we would actually displace, and that is what the price should be based on.
How it works, the questions.
How does RCA as a Service work?
Onepane deploys into your VPC, connects read-only to the observability, change and ITSM tools you already own, builds a service-and-ownership map, and runs an agentic investigation on every Sev1. Engineers sign off, and the finished evidence-linked RCA artifacts are delivered into your ITSM inside a committed SLA window.
What does Onepane need access to?
Read-only access to your telemetry (metrics, logs, traces), change sources (CI/CD, IaC, config, cloud events), topology sources (Kubernetes, cloud control planes, DNS, IAM) and your ITSM. Access is scoped, time-boxed and logged, and you can revoke it. Nothing is written to production systems except the problem record you ask us to file.
How long does onboarding take?
The historical replay runs in about two weeks. Full deployment and coverage of your critical services follows; "days to first RCA in a new tenant" is a metric we track internally and will share.
Do you need our CMDB to be accurate?
No. We build the service-and-ownership map ourselves from what is actually running. That map is usually the piece that stalls in-house programmes, and you keep it.
How is Onepane priced?
A coverage fee per critical service under coverage, plus a committed annual bundle of accepted RCAs, with SLA credits and a one-off onboarding fee. Never per seat, per host or per GB. We don't price before the replay, because the replay tells us how much work we would actually be displacing.
What if our telemetry isn't good enough?
Probably true in places, and it caps our accuracy, which is why we run an RCA Readiness and Telemetry Coverage Report first and tell you exactly which services have blind spots. You get that report whether or not you buy.
See it on your own incidents. Not on our demo data.
Send us your last 90 days of Sev1 tickets. Two weeks, no cost.