Skip to main content
Building and running agents gets work done; Optimize is how you find out whether it was done well. The Komodor Agentic Operation Platform (KAOP) measures a running fleet through four independent lenses, each answering a question the others cannot. The console groups most of these under Improve — Insights, Quality Lab, and Cost. Fleet health is the exception: it sits under Operate → Fleet. Quality Lab carries both quality surfaces — Evaluations on its Runs tab, Model Evaluation on Agent Score Results.

Four lenses, four questions

An agent can be healthy and wrong, correct and ruinously expensive, or cheap and invisible outside KAOP. Those are four separate failure modes, so they get four separate readings.
Reliability and quality are the pair worth internalizing. A failing run announces itself — it has a status, a category, and a place to look. A wrong run does not: it succeeds, costs what it always costs, and leaves no operational trace at all. Reliability tooling will never find it. That gap is why quality gets its own instruments rather than being folded into fleet monitoring.

Reliability

Reliability starts from presence and run outcomes, and both are computed rather than declared. Fleet health gives you the fleet’s live state — who is online, who has capacity, what is queued — alongside period figures for volume, cost, latency, error rate, and how many agent versions were deployed. Insights takes the failures and does the grouping for you. Failing runs are categorized from structural proof — the run’s status, the shape of its span tree, the agent’s presence — never by guessing at error text, and identical failures collapse into one ranked issue with its verbatim reason and an agent-produced recommendation.

Quality

Quality is graded, not observed. Evaluations puts LLM judges on a sampled share of runs and records a score per rubric with the reasoning that produced it — so a low number comes with the criteria that were checked and what moved it off full marks. Model Evaluation answers the comparative question. Deploy a candidate beside production, let both run every real request, and have a blinded judge say which output was better — which is far more sensitive to a small real difference than diffing two absolute scores.

Cost

Agent spend attribution builds every figure from what each run reported, and keeps the derivation visible: cost is reported by the worker, estimated from model rates and token counts, or unattributed. Spend pivots by label, agent, model, provider, MCP server, or tool, and budgets cap what an agent may spend per window — with the option to stop it when the cap is reached.

Portability

Everything above lives in KAOP. Observability & OTel export is what makes it useful outside it, emitting agent activity as standard OpenTelemetry traces, logs, and metrics to your own backend — so an agent’s tool call sits in the same trace as the service it touched.

The order to read them in

1

Is the fleet up?

Fleet health’s live cards. A capacity or presence problem makes every other reading unreliable.
2

What is failing, and is it new?

Insights. Grouped, ranked, and already categorized — with the run evidence one click away, and a recommendation produced by an agent that investigated your account.
3

Of what succeeded, what was any good?

Evaluations. This is the reading nothing else substitutes for.
4

What did all of it cost?

Agent spend. Read cost per successful run before total spend — volume flatters a total.
5

Then change something, and prove it

Model Evaluation. Run the candidate beside production on live traffic and let a blinded judge settle it before anything ships.

The pages

Fleet health

Presence, capacity, backlog, reliability, and version churn across the whole fleet.

Insights

Failures grouped and ranked, each with its verbatim reason and an agent-produced next step.

Evaluations

LLM judges grading runs against explicit rubrics, with the reasoning behind every score.

Model Evaluation

A candidate beside production on real traffic, settled by a blinded judge.

Agent spend attribution

What the fleet costs per run, agent, and model — and budgets that cap it.

Observability & OTel export

Agent traces, logs, and metrics in your own OpenTelemetry backend.

Next steps

Runs & evidence

The record every reading on these pages is computed from.

Build — how it works

Where a fix lands once you know what to change.

Troubleshooting

The fixes for the failure categories Insights reports.

Limits & quotas

The caps that shape what a fleet can do.