AWS Infrastructure Investigator
Overview
An incident with no deploy behind it sends an engineer through CloudTrail by hand, looking for the change that explains it — and the change is rarely in the service that broke. It is in a policy, a security group, or a dependency two hops away. Hand it an incident description and the questions you want answered. It traces root causes through configuration changes, access-control layers and resource dependencies, then returns findings that cite the CloudTrail event, metric or policy behind each claim. Findings come back structured, so an orchestrator can synthesize them with other specialists’ work rather than re-reading prose. Run it on its own and you get the same report in the run’s output. For: SREs and on-call engineers investigating an incident inside one AWS account, especially when nothing deployed. Not for: acting on what it finds — that is Remediation Executor’s job — or for a cause that lives in a different account.Tools supported
Read-only. The investigator has no mutating tool, in any deployment.
Scenario examples
In chat, or as a run’s prompt: A change you cannot seePrerequisites
Catalog ID:
aws-investigator — what the deploy API takes.
The AWS connection offers two authentication modes, and which ones you see depends on where the
agent runs:
The AWS connection is listed as optional because IRSA replaces it. If you deploy self-hosted with a
ServiceAccount that already carries the role, the agent needs no connection at all.
As part of a workflow
A specialist, not a lead. In an incident workflow it takes a scoped sub-task from an orchestrator, investigates the AWS slice of the problem, and hands its findings back to be synthesized with the other specialists’ work. Add a second investigator for another system when an incident spans both — they run independently and their findings are combined. See Orchestrator for how a step delivers work to it, and Incidents for the module that drives it.Limitations
- It changes nothing in AWS. There is no mutating tool in its surface, in either deployment mode, so it cannot stop an instance, edit a policy or roll back a change — it can only tell you which one it thinks is responsible.
- It sees only what its credentials reach. An investigation scoped to one account cannot observe a cause that lives in another, and it reports the gap rather than inferring across it.
- A clean deploy is not proof of access. A well-formed role ARN with insufficient permissions fails on the first run, not at deploy time.
- A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
Recommended model
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where
the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit
it better than a larger model; a faster model degrades the correlation step, which is where the
value is.
Azure Investigator
Overview
Azure spreads the evidence for one incident across Resource Graph, Monitor, Log Analytics and the cluster’s own metadata, each with its own query surface. Assembling that by hand is the slow part of an Azure incident. Hand it an incident description and the questions you want answered. It locates the affected resources, reads their metrics and logs, and correlates what it finds into structured findings — each one carrying a portal deeplink, so the evidence can be opened rather than taken on trust. For: SREs and on-call engineers running workloads in an Azure subscription, including AKS. Not for: acting on what it finds, or for a cause outside the subscription the connection reaches.Tools supported
Read-only. The five capabilities above were measured against a plain Reader role, not assumed from
documentation.
Scenario examples
In chat, or as a run’s prompt: Readiness failureincident_description, questions, affected_services,
investigation_start, context and budget_seconds.
Prerequisites
Catalog ID:
azure-investigator — what the deploy API takes.
Unlike the Datadog and Grafana investigators, this agent has no MCP Gateway path. Azure publishes no
hosted MCP endpoint, so there is no remote server for the gateway to be bound to and no group to
pick — the agent’s Azure tooling runs alongside it, configured from the connection you select.
As part of a workflow
A specialist, not a lead. An orchestrator delegates the Azure slice of an incident to it and synthesizes its findings with other specialists’ work. Where a workload spans Azure and Kubernetes, pairing it with a cluster investigator covers both halves. See Orchestrator for how a step delivers work to it.Limitations
- It changes nothing in Azure. It has no mutating tool, so it cannot restart a resource, edit a configuration or adjust a scaling rule.
- It sees one subscription’s worth of evidence, bounded by what the connection’s role permits. Where a capability is not granted, it reports the gap instead of inferring around it.
- There is no MCP Gateway path. Azure publishes no hosted MCP endpoint, so the gateway is not an option for reaching it.
- A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
Recommended model
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where
the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit
it better than a larger model; a faster model degrades the correlation step, which is where the
value is.
Google Cloud Investigator
Overview
Free-text log search is the wrong instrument for a Google Cloud incident: it finds occurrences, not the first one. The useful question is when the failure actually started and what changed immediately before it, which means narrowing structurally rather than by keyword. Hand it an incident description and the questions you want answered. It narrows the log stream by resource, log name and severity, pins the earliest failing entry, and correlates it against the audit log to find the change that preceded it. Findings cite the exact filter and time window behind each claim, so a conclusion can be reproduced by running the same query yourself. For: establishing when a failure began in one Google Cloud project, and what changed just before it. Not for: metrics, traces, resource configuration or billing — none of those are in its reach.Tools supported
Read-only, and one source. There is nothing else in its surface.
Scenario examples
In chat, or as a run’s prompt: Earliest failureincident_description, questions, affected_services,
investigation_start, context and budget_seconds.
Prerequisites
Catalog ID:
gcp-investigator — what the deploy API takes.
The connection is required rather than optional, because the gateway is this agent’s only path.
There is no alternative tooling bundled alongside it, so without a connection it has no tools at
all.
As part of a workflow
A specialist with a narrow surface, which makes it a good first step rather than a sole one: it establishes when the failure started and what changed around it, and an orchestrator takes that timestamp to the specialists that hold the rest of the picture. See Orchestrator for how a step delivers work to it.Limitations
- It writes nothing to Google Cloud, and it reads nothing outside Cloud Logging. Metrics, traces, resource configuration and billing are all outside its reach — where a cause lives in one of those, it establishes the timeline and says plainly that the rest is not something it can see.
- It is scoped to a single project — the one its service-account key belongs to.
- Without a connection it has no tools at all, because the group-scoped gateway is its only path.
- A partial report is possible. Budget exhaustion returns what it has, so read the findings rather than assuming completeness.
Recommended model
claude-sonnet-4-6 — the shipped default. Investigation is long-context, multi-step tool use where
the work is retrieval and correlation rather than novel reasoning, so Sonnet’s speed and cost suit
it better than a larger model; a faster model degrades the correlation step, which is where the
value is.
Next steps
Agent catalog
Every deployable agent, and what each one is for.
APM
The Datadog and Grafana investigators.
Integrations list
Connecting AWS, Azure and Google Cloud.
Remediation Executor
Turning a finding into an action, with approval.