The four views
Agent state, precisely
An agent’s state is computed from presence, not declared. The control plane pings each open worker connection; the worker’s reply is the only thing that updates presence. If replies stop arriving for 90 seconds, the agent reads as offline.
Several qualifiers can sit beside the status, because they are independent of presence:
- Disabled — triggers will not fire for this agent. Manual and chat invocations still work.
- Archiving / Archived — retiring, and retired. An agent reads Archiving until its last replica is genuinely gone, then Archived.
- New version deploying, Restarting, Reverting — a rollout on an agent that is already live, so its status stays Online while the badge shows what is happening.
- Deploy failed — a rollout onto an already-live agent failed. The previous version keeps serving.
“Offline” and “unreachable” are different facts. Offline means presence lapsed. An agent can be
heartbeat-online while its message channel is not reachable — the Overview calls that out
separately, because the fix is different.
The Overview tab
Choosing a period
The window selector offers Today, Last 7 days, and Last 30 days, and defaults to 30 days. It is calendar-day inclusive: Today means today so far, not a rolling 24 hours. The selector governs this tab’s period figures only.The four headline figures
Trends and leaderboards
Three charts plot the window day by day: Runs & errors over time, Cost over time, and Latency over time. Splitting runs from errors on separate rows is deliberate — an outage day stands out without hiding the throughput underneath it. Three panels rank what needs attention:- Worst error rate — agents ordered by failure rate, with the run count behind each rank so a rate comes with its volume. A minimum-run floor applies to the denominator, so a single bad run on a barely used agent cannot top the board. On a short window the floor can legitimately exclude everything, and the panel says so rather than implying a healthy fleet.
- Slowest agents — ordered by mean duration, with no floor. A slow agent is worth seeing on its first few runs.
- Model mix — which models the fleet actually used, sized by run volume and labelled with spend.
Live state
Three cards describe right now rather than the window, and refresh on their own.Fleet composition
Fleet composition
Four exhaustive buckets — online, offline <24h, offline >24h, and draft —
each linking into the filtered agent list. Splitting recent from long-standing offline matters:
a worker that stopped an hour ago is an incident, and one that stopped last month is a cleanup
task.A Just went offline (last hour) call-out names the agents that changed state most recently,
which is almost always what you came to find.
Capacity & backlog
Capacity & backlog
Live replicas, how many are busy versus free, how many runs are queued, and the age
of the oldest queued run. Oldest-queued is the number to watch: a growing backlog with idle
replicas is a routing problem, and a growing backlog with no free replicas is a capacity one.An Online, 0 live replicas call-out flags an agent that reads online with nothing serving —
which happens when a candidate arm keeps the agent present while its production arm is dead.
Reachability & version mix
Reachability & version mix
Which online agents cannot currently be reached over their channel, headed
Heartbeat-online, channel-unreachable, plus the SDK version mix across live workers.
The version mix is how you find the worker that never got upgraded.
Live figures and period figures refresh on different cadences — presence is polled every 10 seconds
and the window metrics every 60 — so a period number beside a presence row can be up to a minute
older than it.
The Agents tab
The list
Click any sortable header to rank by it. Because the activity cell holds four numbers, its header
offers a Sort by choice of Runs, Cost, Error rate, or Avg latency — so the
ranking always names what it ranked by. Error-rate sorting suppresses agents with too few runs in
the window, for the same reason the leaderboard has a floor.
Narrowing the fleet
- Search matches an agent’s name, id, worker id, description, tools, and labels.
- Status facets: All active agents, Online, Offline.
- Shadow: Shadow active filters to agents currently running a candidate arm — see Model Evaluation.
- Labels: one facet group per label key your agents declare, so
team,env, ortierbecome filters without anyone configuring them. - Clear filters returns you to the whole fleet. Use it when you arrive from a composition card’s link — those can select combinations the facet buttons cannot express, which leaves the buttons reading inactive.
Search scopes the Agents tab only. The Overview deliberately ignores it, so that a query typed on
one tab cannot silently narrow numbers on another with no visible control to clear it.
An individual agent
Opening an agent gives you the same picture for one agent, plus what only it can tell you:- Run history and Run activity — its runs, and its own outcome, cost, latency, error, and change trends.
- Instances — every live replica, including which arm each belongs to.
- Versions — its deployed versions, with invocations split into production and shadow.
- Evals — its judge verdicts and score trend, when it has been evaluated. See Evaluations.
- Skills & tools, Metadata, Permissions — what it knows, what it is, and what it may do.
Reading the fleet in practice
1
Start on Overview at 30 days
Look at Outcomes & reliability and Changes together. A dip that coincides with version churn
points at a rollout.
2
Check live state before period state
Fleet composition and the just-went-offline call-out tell you whether you are looking at a
current outage or at history.
3
Name the agents
Use Worst error rate and Slowest agents to get from “the fleet is degraded” to two agent names.
4
Switch to Agents and sort by error rate
Confirm the ranking against volume, then open the worst agent.
5
Read one failing run
The run’s evidence says what actually happened. If the same failure repeats,
Insights has already grouped it.
Next steps
Runs & evidence
What a single run records, and how to read it.
Agent spend attribution
Where the Cost figure comes from and how it is attributed.
Manage a deployed agent
Enable, disable, archive, roll forward, and roll back.