Skip to main content
Two catalog agents investigate a regression through application performance telemetry: Datadog Investigator and Grafana Investigator. Both run the same aggregate → pivot → trace flow — narrow on the failing slice, read the flame graph of one bad trace within it, correlate against deploys — and return structured findings with a deeplink on every claim. Both are strictly read-only. Neither can mute a monitor, silence an alert, close an incident or edit a dashboard.

Datadog Investigator

Overview

An on-call engineer with a latency regression and a Datadog account has the evidence but not the time: localising it means narrowing on the failing slice, opening one bad trace, reading its flame graph, then checking what deployed. That is twenty minutes of retrieval before the thinking starts, repeated every incident. This agent does that retrieval. Hand it an incident description and the questions you want answered, and it returns structured findings with a Datadog permalink on every claim — so the engineer starts from evidence rather than assembling it. For: SREs and on-call engineers who already run Datadog as their primary telemetry. Not for: infrastructure-level causes — that is a cloud provider investigator’s job.

Tools supported

Read-only. There is no mutating tool in its surface, so it cannot mute, acknowledge, close or edit anything in Datadog.

Scenario examples

In chat, or as a run’s prompt: Latency regression
Error-rate spike
Post-remediation verification
When a workflow drives it instead of a person, it takes a structured envelope:

Prerequisites

Catalog ID: datadog-investigator — what the deploy API takes.
An account can hold several Datadog connections. Provider presence alone does not identify the organization or environment you meant, so select the specific connection this agent should use rather than assuming the only active one is the right one.

As part of a workflow

A specialist, not a lead. An orchestrator delegates the Datadog slice of an incident to it and synthesizes its findings with other specialists’. Because it reads monitors as well as telemetry, it also works as the verification step after a remediation — asking whether the signal that fired has actually recovered. See Orchestrator for how a step delivers work to it, and Incidents for the module that drives it.

Limitations

  • It writes nothing to Datadog. No muting, acknowledging, closing or editing.
  • It is bounded by what Datadog holds. If the span was sampled away or the retention window has passed, it reports that it could not verify the step rather than inferring.
  • It sees only Datadog. A cause in the cloud provider, the cluster, or the code is outside its reach — pair it with the relevant investigator.
  • A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit it better than a larger model; a faster model degrades the correlation step, which is where the value is.

Grafana Investigator

Overview

The same problem in a different stack. A regression visible in Prometheus means pivoting to Loki for the error text, to Tempo for the trace, and to Pyroscope when the cost is CPU or memory rather than a downstream call — four tools and four query languages before the diagnosis starts. Hand it an incident description and the questions you want answered. It runs the aggregate → pivot → trace flow across those data sources and returns structured findings with Grafana deeplinks, so each claim can be checked. For: SREs and on-call engineers whose telemetry lives in the Grafana stack, on Grafana Cloud or their own Grafana. Not for: infrastructure-level causes — that is a cloud provider investigator’s job.

Tools supported

Read-only across all of them. It cannot silence an alert, close an incident, edit a dashboard or add an annotation.

Scenario examples

In chat, or as a run’s prompt: Error-rate spike
Saturation, not errors
Did the alert tell the truth
When a workflow drives it instead of a person, it takes the same structured envelope the other investigators take — incident_description, questions, affected_services, investigation_start, context and budget_seconds.

Prerequisites

Catalog ID: grafana-investigator — what the deploy API takes.

Three ways to reach Grafana

The Grafana connection is listed as optional because the second path replaces it: a self-hosted agent pointed at your own in-cluster server needs no connection here. Add one only to use the agent’s built-in Grafana tooling.
The wizard asks where the agent runs and then disables the paths that target does not support, rather than hiding them — so a choice that is unavailable is visibly unavailable.

As part of a workflow

A specialist, not a lead. An orchestrator delegates the Grafana slice of an incident to it and synthesizes its findings with the other specialists’. Because it reads alerts as well as telemetry, it also works as the verification step after a remediation — asking whether the signal that fired has actually recovered. See Orchestrator for how a step delivers work to it.

Limitations

  • It writes nothing to Grafana. No silencing an alert, closing an incident, editing a dashboard or adding an annotation.
  • It is bounded by retention, and by what its path can reach. A self-hosted agent pointed at your own in-cluster server sees exactly what that server exposes, and it reports an unreachable data source rather than working around it.
  • It sees only Grafana. A cause in the cloud provider, the cluster, or the code is outside its reach — pair it with the relevant investigator.
  • A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit it better than a larger model; a faster model degrades the correlation step, which is where the value is.

Next steps

Agent catalog

Every deployable agent, and what each one is for.

Cloud providers

The AWS, Azure and Google Cloud investigators.

Integrations list

Connecting Datadog and Grafana.

Integration groups

The group-scoped gateway endpoint Datadog needs.