Skip to main content
Three catalog agents investigate Kubernetes, and which to deploy depends on how much you want it to see: Kubernetes RCA for one resource in one cluster, Cluster Investigator for one cluster correlated against Grafana and GPU health, and Klaudia Investigator for a question spanning every cluster your Komodor account covers. The first two run inside a cluster and see only that cluster. Deploying one instance and expecting it to answer for a fleet is the mistake worth avoiding before you install rather than after.

Kubernetes RCA

Overview

Reading events off a broken Deployment tells you it is broken. It does not tell you that the Kafka consumer group is stuck, or that the Istio sidecar never became ready — that takes knowing how the specific thing fails, which is knowledge most on-call engineers hold for a handful of workloads and nobody holds for all of them. Name a resource — a kind and a name — and it investigates that resource: reconstructing what changed around it, reading the state of its dependencies, and working through a domain-specific method for whatever it is. It is built from Klaudia’s root-cause knowledge, which is why it can diagnose a Kafka consumer group or a stuck Istio sidecar rather than only reading events. It returns a structured root-cause analysis for that one resource. For: root-causing one named resource, in the one cluster the agent runs in. Not for: a fleet. One instance does not answer for more than its own cluster — use Klaudia Investigator for that.
This agent investigates only the cluster it is deployed in. One instance does not answer for a fleet, and that is the mistake worth avoiding before you install rather than after. For multi-cluster work, use Klaudia Investigator instead.The wizard states this on its first step and holds Next until you confirm it, and it recommends putting the cluster name in the agent’s name for exactly this reason.

Tools supported

Read-only kubectl against the cluster it runs in. There is no connection to bind, because the ServiceAccount’s own permissions are simultaneously its only data source and its only security boundary.

Scenario examples

This agent takes a resource, not a question. A run’s input names what to investigate: A stuck StatefulSet
A deployment that will not roll out
A job that failed overnight

Prerequisites

Catalog ID: kubernetes-rca — what the deploy API takes. The self-hosted install renders a values file that binds the agent’s ServiceAccount to Kubernetes’ own aggregated view role, plus a supplemental role for the cluster-scoped reads view omits. Leaning on the built-in role means custom resources are covered as they are installed, rather than by a list that goes stale.

The one grant that is not read-only

pods/exec is granted cluster-wide, because several domain methods diagnose from inside the pod — that state exists nowhere else, and at install time nobody knows which namespace an incident will land in. It is a separate value from the read access, so you can narrow or remove it: What may actually run through that grant is decided by the agent itself, not only by RBAC. It admits a short list of read-only diagnostics and refuses everything else, interactive sessions included. No database client is on that list, so a Postgres incident that turns on a grant or a row-security policy is out of reach — and the agent names that as a blind spot rather than approximating.
The view role excludes Secrets, so Helm release history is unavailable and the change-timeline method falls back to its other sources. Two grants are deliberately off by default and need an edit to the values file: reads on custom resource groups that do not aggregate to view, and cluster-wide Secret reads.

As part of a workflow

A specialist bound to one cluster, so a workflow step that targets it has already decided which cluster the problem is in. Where a fleet-wide investigation is the starting point, an orchestrator reaches for a multi-cluster investigator first and delegates here once the cluster is known. See Orchestrator for how a step delivers work to it.

Limitations

  • It changes nothing. Its read access is the whole of its reach, and the one non-read grant is confined to a fixed list of diagnostics by the agent’s own classifier.
  • It sees exactly one cluster and no external context. No Komodor backend, no metrics vendor and no cloud provider API — a cause in any of those is outside what it can observe, and it says so rather than inferring.
  • Some evidence is off by default. The view role excludes Secrets, so Helm release history is unavailable unless you edit the values file.
  • No database client is admitted through exec, so a Postgres incident that turns on a grant or a row-security policy is out of reach — named as a blind spot rather than approximated.
claude-sonnet-5 — the shipped default. The work is judgement over evidence someone else gathered, where a weaker model produces confident-sounding conclusions that do not hold.

Cluster Investigator

Overview

A pod restarting because of a GPU fault looks like an ordinary crash loop until something reads both the cluster and the hardware. The same is true of a latency regression that is visible in traces but explained by a node: the answer sits between two consoles, and nobody correlates them under pressure. Describe what is wrong and it investigates across all three signal sources at once: reading cluster state through read-only kubectl, pulling the metrics, logs and traces around the incident from Grafana, and checking GPU health where the workload depends on it. It reconstructs the change timeline and returns a structured root-cause analysis. For: one cluster where the cause could be the workload, the telemetry or the GPU — especially ML serving. Not for: a fleet. Run one per cluster, and name each after the cluster it watches.

Tools supported

Read-only, apart from the in-pod diagnostics noted under Prerequisites. It carries the same domain knowledge as Kubernetes RCA — twenty-three workload-specific methods — plus five GPU-specific ones.

Scenario examples

In chat, or as a run’s prompt: Restart loop with GPU in play
Slow, not failing
Node-level suspicion

Prerequisites

Catalog ID: cluster-investigator — what the deploy API takes.
The Grafana connection is optional because the agent can instead reach your own in-cluster MCP server, or run its own bundled Grafana tooling. Add a connection only to use the built-in path. Without any of the three, it still investigates the cluster — it simply has no telemetry to correlate against.
The install grants the same cluster access as Kubernetes RCA: Kubernetes’ aggregated view role, a supplemental role for the cluster-scoped reads it omits, and pods/exec for the workload methods that diagnose from inside the pod. Exec is a separate value, so you can narrow it to named namespaces or remove it and let those methods fall back to resource state.

As part of a workflow

A specialist bound to one cluster, so a step targeting it has already established which cluster the problem is in. Because it covers three signal sources, it often replaces two steps that would otherwise run in parallel and need synthesizing afterwards. See Orchestrator for how a step delivers work to it.

Limitations

  • It changes nothing in the cluster or in Grafana. Its non-read grant is confined to a fixed list of read-only in-pod diagnostics, and everything else is refused.
  • It sees exactly one cluster. Deploying one instance and expecting it to answer for a fleet is the mistake worth avoiding — run one per cluster, and name each after the cluster it watches.
  • Without a Grafana path it still investigates, but it has no telemetry to correlate the cluster state against.
  • A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit it better than a larger model.

Klaudia Investigator

Overview

The in-cluster investigators each answer for one cluster, which is the wrong shape for a question that starts “which of our clusters…”. Answering that by deploying an agent everywhere and collating the results is work nobody wants to do during an incident. Ask a Kubernetes question or describe an incident. The agent puts it to Klaudia, who investigates across the clusters your Komodor account already covers, and relays the answer back — as prose for a question, or as a structured root-cause analysis when that is what she produces. For: questions that span clusters, or an incident whose cluster is not yet known. Not for: answering an approval or a choice on your behalf, and not a substitute for an investigator where you already know the cluster.

Tools supported

It reads nothing directly. Its entire surface is a small, fixed set of tools for holding a conversation with Klaudia — asking a question, collecting the result, and nothing else. The investigation happens on Komodor’s side, and its reach is whatever your Komodor account’s agent already sees. Because that surface is fixed by the agent itself, the wizard’s tool checklist has no effect on it: what it can do is decided by the implementation, not by what you tick.

Scenario examples

In chat, or as a run’s prompt: One workload, one cluster
Across the fleet
A question, not an incident

Prerequisites

Catalog ID: klaudia-investigator — what the deploy API takes.
There is no Komodor connection to create. The Komodor MCP server authenticates itself on the gateway server object, configured once in gateway administration, so the Integrations step has nothing to collect for this agent — only the MCP group needs binding.

As part of a workflow

The multi-cluster counterpart to the in-cluster investigators: where a step needs a Kubernetes answer but the cluster is not known in advance, this is the specialist to route to. An orchestrator can also use it as a first pass, then delegate to a cluster-bound investigator once the affected cluster is identified. See Orchestrator for how a step delivers work to it.

Limitations

  • It cannot answer an approval or a choice on your behalf. That is enforced in the agent rather than by configuration, so a gate Klaudia raises always comes back to a human.
  • It performs no investigation of its own. Everything it reports comes from Klaudia, so what it can see is exactly what your Komodor account covers — no more, and no less.
  • The wizard’s tool checklist does not apply. Its surface is fixed by the implementation.
claude-sonnet-5 — the shipped default. The work is judgement over evidence someone else gathered, where a weaker model produces confident-sounding conclusions that do not hold.

Next steps

Agent catalog

Every deployable agent, and what each one is for.

APM

The Datadog and Grafana investigators.

Integration groups

The gateway endpoint Klaudia Investigator needs.

Kubernetes Cost

The module that owns Kubernetes spend rather than incidents.