Datadog Investigator
Overview
An on-call engineer with a latency regression and a Datadog account has the evidence but not the time: localising it means narrowing on the failing slice, opening one bad trace, reading its flame graph, then checking what deployed. That is twenty minutes of retrieval before the thinking starts, repeated every incident. This agent does that retrieval. Hand it an incident description and the questions you want answered, and it returns structured findings with a Datadog permalink on every claim — so the engineer starts from evidence rather than assembling it. For: SREs and on-call engineers who already run Datadog as their primary telemetry. Not for: infrastructure-level causes — that is a cloud provider investigator’s job.Tools supported
Read-only. There is no mutating tool in its surface, so it cannot mute, acknowledge, close or edit
anything in Datadog.
Scenario examples
In chat, or as a run’s prompt: Latency regressionPrerequisites
Catalog ID:
datadog-investigator — what the deploy API takes.
An account can hold several Datadog connections. Provider presence alone does not identify the
organization or environment you meant, so select the specific connection this agent should use
rather than assuming the only active one is the right one.
As part of a workflow
A specialist, not a lead. An orchestrator delegates the Datadog slice of an incident to it and synthesizes its findings with other specialists’. Because it reads monitors as well as telemetry, it also works as the verification step after a remediation — asking whether the signal that fired has actually recovered. See Orchestrator for how a step delivers work to it, and Incidents for the module that drives it.Limitations
- It writes nothing to Datadog. No muting, acknowledging, closing or editing.
- It is bounded by what Datadog holds. If the span was sampled away or the retention window has passed, it reports that it could not verify the step rather than inferring.
- It sees only Datadog. A cause in the cloud provider, the cluster, or the code is outside its reach — pair it with the relevant investigator.
- A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
Recommended model
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where
the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit
it better than a larger model; a faster model degrades the correlation step, which is where the
value is.
Grafana Investigator
Overview
The same problem in a different stack. A regression visible in Prometheus means pivoting to Loki for the error text, to Tempo for the trace, and to Pyroscope when the cost is CPU or memory rather than a downstream call — four tools and four query languages before the diagnosis starts. Hand it an incident description and the questions you want answered. It runs the aggregate → pivot → trace flow across those data sources and returns structured findings with Grafana deeplinks, so each claim can be checked. For: SREs and on-call engineers whose telemetry lives in the Grafana stack, on Grafana Cloud or their own Grafana. Not for: infrastructure-level causes — that is a cloud provider investigator’s job.Tools supported
Read-only across all of them. It cannot silence an alert, close an incident, edit a dashboard or
add an annotation.
Scenario examples
In chat, or as a run’s prompt: Error-rate spikeincident_description, questions, affected_services,
investigation_start, context and budget_seconds.
Prerequisites
Catalog ID:
grafana-investigator — what the deploy API takes.
Three ways to reach Grafana
The Grafana connection is listed as optional because the second path replaces it: a self-hosted
agent pointed at your own in-cluster server needs no connection here. Add one only to use the
agent’s built-in Grafana tooling.
As part of a workflow
A specialist, not a lead. An orchestrator delegates the Grafana slice of an incident to it and synthesizes its findings with the other specialists’. Because it reads alerts as well as telemetry, it also works as the verification step after a remediation — asking whether the signal that fired has actually recovered. See Orchestrator for how a step delivers work to it.Limitations
- It writes nothing to Grafana. No silencing an alert, closing an incident, editing a dashboard or adding an annotation.
- It is bounded by retention, and by what its path can reach. A self-hosted agent pointed at your own in-cluster server sees exactly what that server exposes, and it reports an unreachable data source rather than working around it.
- It sees only Grafana. A cause in the cloud provider, the cluster, or the code is outside its reach — pair it with the relevant investigator.
- A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
Recommended model
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where
the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit
it better than a larger model; a faster model degrades the correlation step, which is where the
value is.
Next steps
Agent catalog
Every deployable agent, and what each one is for.
Cloud providers
The AWS, Azure and Google Cloud investigators.
Integrations list
Connecting Datadog and Grafana.
Integration groups
The group-scoped gateway endpoint Datadog needs.