Kubernetes RCA
Overview
Reading events off a broken Deployment tells you it is broken. It does not tell you that the Kafka consumer group is stuck, or that the Istio sidecar never became ready — that takes knowing how the specific thing fails, which is knowledge most on-call engineers hold for a handful of workloads and nobody holds for all of them. Name a resource — a kind and a name — and it investigates that resource: reconstructing what changed around it, reading the state of its dependencies, and working through a domain-specific method for whatever it is. It is built from Klaudia’s root-cause knowledge, which is why it can diagnose a Kafka consumer group or a stuck Istio sidecar rather than only reading events. It returns a structured root-cause analysis for that one resource. For: root-causing one named resource, in the one cluster the agent runs in. Not for: a fleet. One instance does not answer for more than its own cluster — use Klaudia Investigator for that.Tools supported
Read-onlykubectl against the cluster it runs in. There is no connection to bind, because the
ServiceAccount’s own permissions are simultaneously its only data source and its only security
boundary.
Scenario examples
This agent takes a resource, not a question. A run’s input names what to investigate:
A stuck StatefulSet
Prerequisites
Catalog ID:
kubernetes-rca — what the deploy API takes.
The self-hosted install renders a values file that binds the agent’s ServiceAccount to Kubernetes’
own aggregated view role, plus a supplemental role for the cluster-scoped reads view omits.
Leaning on the built-in role means custom resources are covered as they are installed, rather than
by a list that goes stale.
The one grant that is not read-only
pods/exec is granted cluster-wide, because several domain methods diagnose from inside the pod —
that state exists nowhere else, and at install time nobody knows which namespace an incident will
land in.
It is a separate value from the read access, so you can narrow or remove it:
What may actually run through that grant is decided by the agent itself, not only by RBAC. It admits
a short list of read-only diagnostics and refuses everything else, interactive sessions included.
No database client is on that list, so a Postgres incident that turns on a grant or a row-security
policy is out of reach — and the agent names that as a blind spot rather than approximating.
The
view role excludes Secrets, so Helm release history is unavailable and the change-timeline
method falls back to its other sources. Two grants are deliberately off by default and need an
edit to the values file: reads on custom resource groups that do not aggregate to view, and
cluster-wide Secret reads.As part of a workflow
A specialist bound to one cluster, so a workflow step that targets it has already decided which cluster the problem is in. Where a fleet-wide investigation is the starting point, an orchestrator reaches for a multi-cluster investigator first and delegates here once the cluster is known. See Orchestrator for how a step delivers work to it.Limitations
- It changes nothing. Its read access is the whole of its reach, and the one non-read grant is confined to a fixed list of diagnostics by the agent’s own classifier.
- It sees exactly one cluster and no external context. No Komodor backend, no metrics vendor and no cloud provider API — a cause in any of those is outside what it can observe, and it says so rather than inferring.
- Some evidence is off by default. The
viewrole excludes Secrets, so Helm release history is unavailable unless you edit the values file. - No database client is admitted through exec, so a Postgres incident that turns on a grant or a row-security policy is out of reach — named as a blind spot rather than approximated.
Recommended model
claude-sonnet-5 — the shipped default. The work is judgement over evidence someone else gathered,
where a weaker model produces confident-sounding conclusions that do not hold.
Cluster Investigator
Overview
A pod restarting because of a GPU fault looks like an ordinary crash loop until something reads both the cluster and the hardware. The same is true of a latency regression that is visible in traces but explained by a node: the answer sits between two consoles, and nobody correlates them under pressure. Describe what is wrong and it investigates across all three signal sources at once: reading cluster state through read-onlykubectl, pulling the metrics, logs and traces around the incident from
Grafana, and checking GPU health where the workload depends on it. It reconstructs the change
timeline and returns a structured root-cause analysis.
For: one cluster where the cause could be the workload, the telemetry or the GPU — especially ML
serving.
Not for: a fleet. Run one per cluster, and name each after the cluster it watches.
Tools supported
Read-only, apart from the in-pod diagnostics noted under Prerequisites. It carries the same domain
knowledge as Kubernetes RCA — twenty-three workload-specific methods — plus five
GPU-specific ones.
Scenario examples
In chat, or as a run’s prompt: Restart loop with GPU in playPrerequisites
Catalog ID:
cluster-investigator — what the deploy API takes.
The Grafana connection is optional because the agent can instead reach your own in-cluster MCP
server, or run its own bundled Grafana tooling. Add a connection only to use the built-in path.
Without any of the three, it still investigates the cluster — it simply has no telemetry to
correlate against.
view role, a
supplemental role for the cluster-scoped reads it omits, and pods/exec for the workload methods
that diagnose from inside the pod. Exec is a separate value, so you can narrow it to named
namespaces or remove it and let those methods fall back to resource state.
As part of a workflow
A specialist bound to one cluster, so a step targeting it has already established which cluster the problem is in. Because it covers three signal sources, it often replaces two steps that would otherwise run in parallel and need synthesizing afterwards. See Orchestrator for how a step delivers work to it.Limitations
- It changes nothing in the cluster or in Grafana. Its non-read grant is confined to a fixed list of read-only in-pod diagnostics, and everything else is refused.
- It sees exactly one cluster. Deploying one instance and expecting it to answer for a fleet is the mistake worth avoiding — run one per cluster, and name each after the cluster it watches.
- Without a Grafana path it still investigates, but it has no telemetry to correlate the cluster state against.
- A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
Recommended model
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where
the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit
it better than a larger model.
Klaudia Investigator
Overview
The in-cluster investigators each answer for one cluster, which is the wrong shape for a question that starts “which of our clusters…”. Answering that by deploying an agent everywhere and collating the results is work nobody wants to do during an incident. Ask a Kubernetes question or describe an incident. The agent puts it to Klaudia, who investigates across the clusters your Komodor account already covers, and relays the answer back — as prose for a question, or as a structured root-cause analysis when that is what she produces. For: questions that span clusters, or an incident whose cluster is not yet known. Not for: answering an approval or a choice on your behalf, and not a substitute for an investigator where you already know the cluster.Tools supported
It reads nothing directly. Its entire surface is a small, fixed set of tools for holding a conversation with Klaudia — asking a question, collecting the result, and nothing else. The investigation happens on Komodor’s side, and its reach is whatever your Komodor account’s agent already sees. Because that surface is fixed by the agent itself, the wizard’s tool checklist has no effect on it: what it can do is decided by the implementation, not by what you tick.Scenario examples
In chat, or as a run’s prompt: One workload, one clusterPrerequisites
Catalog ID:
klaudia-investigator — what the deploy API takes.
There is no Komodor connection to create. The Komodor MCP server authenticates itself on the gateway
server object, configured once in gateway administration, so the Integrations step has nothing to
collect for this agent — only the MCP group needs binding.
As part of a workflow
The multi-cluster counterpart to the in-cluster investigators: where a step needs a Kubernetes answer but the cluster is not known in advance, this is the specialist to route to. An orchestrator can also use it as a first pass, then delegate to a cluster-bound investigator once the affected cluster is identified. See Orchestrator for how a step delivers work to it.Limitations
- It cannot answer an approval or a choice on your behalf. That is enforced in the agent rather than by configuration, so a gate Klaudia raises always comes back to a human.
- It performs no investigation of its own. Everything it reports comes from Klaudia, so what it can see is exactly what your Komodor account covers — no more, and no less.
- The wizard’s tool checklist does not apply. Its surface is fixed by the implementation.
Recommended model
claude-sonnet-5 — the shipped default. The work is judgement over evidence someone else gathered,
where a weaker model produces confident-sounding conclusions that do not hold.
Next steps
Agent catalog
Every deployable agent, and what each one is for.
APM
The Datadog and Grafana investigators.
Integration groups
The gateway endpoint Klaudia Investigator needs.
Kubernetes Cost
The module that owns Kubernetes spend rather than incidents.