You can build the investigation. You can't build the SLA.
Build vs buy for root cause analysis comes down to one distinction: a good platform team can build an AI incident-investigation agent in a quarter, but it cannot build the committed window, the accountability when the answer is wrong at 3am, or the maintenance eighteen months later when the author has moved teams. Onepane sells those, delivered in your VPC.
"We'll just build this ourselves." You probably can.
This is the most common alternative we meet, and the one we take most seriously, because it is true: a good platform team really can build a passable investigation agent in a quarter. Model access is cheap, the observability APIs are documented, and the first demo on a clean incident looks great. We do not argue that you cannot.
What we ask instead: when your version is wrong at 3am, who is accountable? Who maintains the integrations when the observability vendor changes its API and the engineer who built it has moved teams? Who writes the customer-facing version, the evidence pack, the CAPA tracker, or is that still by hand?
We are not selling you an investigation engine. We are selling the fact that the RCA arrives, on time, and someone other than your team is answerable for it.
How does an in-house build compare with a managed root-cause service?
| Dimension | Build in-house | Onepane, managed RCA |
|---|---|---|
| Time to value | A passable investigation agent in a quarter, if the platform team is not pulled onto the roadmap. Coverage of change sources, topology and ITSM takes longer. | The 90-day replay in about two weeks; production coverage of critical services after onboarding. Days to first RCA in a new tenant is a metric we track and share. |
| Accountability when it is wrong | Yours. At 3am, when the internal tool names the wrong cause, the on-call engineer owns the consequence and nobody outside the team is answerable. | Ours. Engineers sign off before anything ships; abstention is a first-class output; the SLA and credits put someone other than your team on the hook. |
| Maintenance at month 18 | Observability vendors change APIs, the estate changes, the engineer who built it moves teams. Internal tooling projects taken on alongside a roadmap commonly stall around 70 percent. | Connectors, models and the service-and-ownership map are maintained as part of the service, in your VPC, with automated upgrades. |
| Coverage of the estate | Usually the stack the platform team knows: cloud, Kubernetes, the primary observability vendor. The database, the mainframe, the change ticket and the acquired stack come later, or never. | Everything the causal chain touches, from day one: multiple observability vendors, databases, cloud control planes, CI/CD, IaC, ITSM. |
| Evidence artifacts | A summary in Slack or a wiki page. The Customer-Facing RCA, Evidence Pack, CAPA Tracker and Problem Record still get written by hand. | The full artifact set, generated from one evidence-linked investigation and delivered into your ITSM inside the SLA. |
| Deployment | In your environment, which is a genuine advantage of building. | In your VPC as well, the one property of building that a SaaS tool cannot match, and one we share. |
| What it costs you | Senior platform-engineering time, indefinitely, plus the opportunity cost of what they were meant to ship. | A coverage fee per service and a bundle of accepted RCAs, funded from the investigation and authoring labour it displaces. |
What can't a platform team build alongside its roadmap?
The SLA
An internal tool has no committed window and no credits. When the RCA is late, that is a conversation, not a contract. The recipient of the RCA, the customer, the regulator, the review board, does not care that the tool was built in-house.
The accountability
When the internal agent is confidently wrong at 3am, who is answerable? The engineer who built it, who has other work; or the on-call, who trusted it. A managed service puts someone outside your team on the hook, with human sign-off before anything ships.
The maintenance
Eighteen months from now the observability vendor has changed its API, the estate has grown by an acquisition, and the engineer who wrote the agent has moved teams. Who maintains the integrations? Ask what happened to the last internal tooling project your platform team took on alongside its roadmap.
The expert pod
Behind the service is a team whose whole job is root cause across many estates, the pattern library, the abstention discipline, the review. A platform team building on the side does not get that, and should not have to.
The artifacts an internal build rarely gets to: Customer-Facing RCA, Evidence Pack, Confidence and Abstention Statement, CAPA Tracker, Problem Record Payload. All generated today, delivered inside your VPC.
Run the replay. Then decide.
Send us your last 90 days of Sev1 tickets. We show what we would have found, how fast, and what the document would have looked like, scored against the RCA your team actually wrote. If your platform team can produce that output in a quarter, you should build it, and you will have a much better spec for having seen ours.
If they cannot, or if they can but should be shipping the roadmap instead, you will know that too, with numbers from your own incidents rather than a vendor's demo. See how the service runs.
Build vs buy, the questions.
Should we build or buy root cause analysis?
If your platform team can genuinely dedicate a quarter and then own the tool indefinitely, and nobody outside engineering is waiting on the RCA document, building is reasonable. If the RCA is owed to a customer, a regulator or a review board on a clock, buy the outcome: what you cannot build is the SLA, the accountability when it is wrong, and the maintenance eighteen months later. Run the replay first either way, it is a better spec than a whiteboard.
Can't a good platform team build an AI incident investigation agent?
Yes. A capable team can build a passable investigation agent in a quarter, and we do not argue otherwise. The investigation is the part that is buildable. The service around it, committed windows, credits, human sign-off, abstention discipline, connector maintenance, the artifact set, is the part that is not, and it is the part the recipient of the RCA actually depends on.
What if we build it and it works?
Then you should keep it, and you will have a much better spec for having seen the replay. Our honest closer is: run the replay; if your team can produce that output in a quarter, build it. Most teams find the investigation is the easy 70 percent and the artifact, ownership and maintenance are the hard 30.
Does buying mean giving up control of our data or environment?
No. Onepane deploys in your VPC, the same place your internal build would run. Telemetry stays in your account under your keys; our engineers' access is scoped, time-boxed, logged and revocable; source escrow and data portability are agreed in the contract. You keep the service-and-ownership map and every artifact.
A better spec than a whiteboard.Your own last 90 days, replayed.
Two weeks, no cost. If your team can match the output, build it. Either way you get the readiness report.