Agents investigate, humans sign off: what RCA as a Service means in the agentic era
By Arun Mohan, Founder of Onepane · August 2026 · 8 min read
RCA as a Service is a managed root-cause service: agents investigate every Sev1 inside your own environment, on the observability you already own; human engineers sign off; and you receive a finished, evidence-linked root cause document inside a contractual SLA. The product is the deliverable and the commitment behind it, not the software that produced it. This post defines the category and says what we hold ourselves to inside it.
What is RCA as a Service?
Every enterprise has monitoring. Almost none of them have root cause. After a major incident, three engineers spend four hours reconstructing what happened across five tools, and someone spends another day writing the document that the customer, the regulator or the problem-review board is waiting on. That work is repeatable, evidence-driven and time-boxed, which makes it a service, and it is now a service that agents can do most of.
RCA as a Service, then, has four defining properties. Miss any one and you are looking at something else.
- The deliverable is a document with an evidence index. A root cause report where every claim links to the log line, metric or change record behind it, plus the derived artifacts: the sanitised customer-facing RCA, the evidence pack, the change attribution record, the problem-record payload for your ITSM.
- It is bound by an SLA. Time to RCA against a committed window, with credits when the window is missed. This is what makes it a service rather than software.
- It runs on what you already own, inside your perimeter. No replacement of your monitoring, no logs shipped to a vendor cloud, no production data sent to a third-party model provider. See deploy in your VPC.
- Humans are accountable for the output. Agents investigate; engineers review the evidence and sign off, or abstain. Someone other than your team is answerable when the RCA is wrong or late.
In an ITIL-mature enterprise, the same thing has a more familiar name: autonomous problem management. Problem Management is a named practice with a named owner, a mandatory deliverable and an existing services budget line. RCA as a Service is that practice, carried out by agents with human sign-off. We deliberately never abbreviate the phrase, because the abbreviation collides with application performance monitoring and confuses every technical buyer.
What did MDR do for the SOC, and why is the NOC next?
The closest precedent is managed detection and response in security. A decade ago, security teams owned a SIEM and a queue of alerts nobody had time to investigate. The MDR providers did not sell them another SIEM. They sat on the customer’s existing tooling, investigated the detections, and delivered triaged, evidence-backed findings against an SLA. What they displaced was analyst hours, not the tools. That category is now measured in billions of dollars of annual revenue and is the default way a mid-sized enterprise runs a SOC.
Operations teams are where security teams were. They own a monitoring estate and a queue of incidents, and the expensive part is the investigation and the write-up, not the alerting. RCA as a Service applies the MDR shape to the NOC: sit on the observability the customer already has, investigate the incidents, deliver the finished root cause against an SLA, and displace the labour rather than the tooling.
| MDR (security operations) | RCA as a Service (IT operations) | |
|---|---|---|
| Sits on | The customer’s SIEM, EDR, cloud logs | The customer’s metrics, logs, traces, change and topology data |
| Investigates | Detections | Sev1 and Sev2 incidents |
| Delivers | Triaged, evidence-backed findings | Evidence-linked root cause reports and derived artifacts |
| Bound by | Response-time SLA | Time-to-RCA SLA with credits |
| Displaces | Analyst hours | Investigation and RCA-authoring hours |
| Where it runs | Vendor cloud, typically | Inside the customer’s VPC |
The one deliberate departure from the analogy is the last row. Security telemetry has largely been accepted as something that can leave the perimeter. Production logs and traces have not, and in regulated accounts never will be. So the service has to come to the data.
Why is the deliverable the product, and not the software?
The agentic era has produced a wave of chat-first investigation tools, and many of them are good. They give an engineer a faster answer in a chat window during an incident. That is useful, and it is a different product from the one this post is about, for three reasons.
- The buyer is different. A tool for engineers is bought by a platform-engineering leader out of a tooling budget. A service with a deliverable is bought by an operations leader out of a labour or managed-services budget. Those budgets differ by an order of magnitude and are approved by different people.
- The obligation is different. A tool that is wrong or unavailable at 3am has a support ticket. A service that is wrong or late has a credit, and a named person accountable for the miss.
- The output is different. A chat answer is not something a customer’s procurement team, an auditor or a problem-review board will accept. A document with an evidence index is. Producing that document, and the sanitised, executive and ITSM variants of it from the same evidence, is most of the work that the chat answer leaves behind.
The consequence is that RCA as a Service is priced the way services are priced: per service under coverage, plus accepted RCAs. Not per seat, host or gigabyte, because those are the models that make an observability vendor’s incentives diverge from yours the moment the product starts reducing your consumption. We describe the model, without numbers, on how it works.
Why is abstention a first-class output?
The last generation of tools in this space oversold, and the memory is fresh. Buyers who lived through those programmes ask, correctly, “what if it is confidently wrong?” A confidently wrong root cause is worse than none: it closes the problem record, misdirects the corrective action, and gets sent to a customer who later discovers it was wrong.
So a service in this category has to treat “insufficient evidence” as a legitimate deliverable rather than a failure. The confidence and abstention statement that accompanies every report says what was concluded, at what confidence, what evidence was missing, and what could not be determined. When the answer is an abstention, it says specifically which telemetry or change record would have resolved it, which is itself actionable: it becomes the readiness backlog for that service.
This is also what makes human sign-off meaningful rather than ceremonial. The engineer reviewing the report is not asked “do you trust the agent?” They are asked “does this evidence support this claim?” and given the links to check. Where the answer is no, the report does not ship. That is the operating discipline behind “agents investigate, humans sign off”, and it is the reason regulated buyers who would never accept an autonomous actor in production will accept evidence-backed decision support with an audit trail.
What do we publish?
A category defined by transparency has to be transparent about itself. We publish, monthly and per account, the numbers that show whether the service is doing what the definition says, whether or not a given month flatters us:
- Accuracy against the customer’s own historical human-written RCAs, established first in the replay and tracked afterwards.
- Abstention rate, so a low error rate cannot be bought by refusing to answer.
- Time to RCA against the committed SLA window, with the credits it triggered. Benchmarks are on the RCA SLA page.
- Human-touch rate: the share of RCAs that shipped with and without an engineer intervening beyond sign-off, trending over time.
The definition of the category is not a slogan. It is a set of obligations, and the only fair way to evaluate anyone claiming it, ourselves included, is against your own incidents. Send us your last ninety days of Sev1 tickets and we will show you what the RCA would have said, how long it would have taken, and where we would have abstained. Run the replay.
Send us your last 90 days of Sev1 tickets. In two weeks, at no cost, we show you what we would have found, how fast, and what the RCA document would have looked like.