Skip to main content
A run that succeeded can still be wrong. Evaluations is how the Komodor Agentic Operation Platform (KAOP) grades what agents actually produced — an LLM judge reads a run and returns a score per rubric with its reasoning attached. Find it under Improve → Quality Lab, which opens on Runs.

The tabs

Grading is itself agent work. A judge is an agent, and its grading pass is a real run against your control plane — with an input, tool calls, and an output. That is why a verdict can carry reasoning rather than a bare number: the reasoning is something an agent produced and the platform stored, not a label the platform inferred.

Setting up grading

Grading starts when you create an eval rule — the configuration that decides which agents’ runs get graded, how often, and by whom. Rules live on the Evaluators tab alongside the judges themselves: New eval rule opens the form, and each existing rule has Edit.
1

Name the rule and choose what to grade

Single run grades each matching run on its own — the default. Run pair (shadow vs. production) grades a completed comparison pair together and belongs to Model Evaluation.
2

Pick the agents

A rule names the agents whose runs it grades. Nothing outside that list is touched.
3

Set the sample rate

10%, 25%, 50%, or 100%, defaulting to 10%. The rate is per rule: the whole panel of judges evaluates the same fraction of matching runs, so their verdicts are comparable. It is the one knob that trades cost for confidence — see Choosing a rate below.
4

Choose the panel of judges

One or several. A judge is an ordinary agent that advertises the grading capability, listed on the Evaluators tab; each one you select produces its own independent verdict.
5

Add rule instructions

Free text telling the judges what “good” means for these agents specifically — for example, that a weekly report must lead with the total, or that a table missing a share-of-total column is a defect. Optional, and the most under-used field on the form.
6

Activate

The rule starts matching new runs immediately, samples them at the rule’s rate, and dispatches one grading task per judge. Edits apply to new matching runs from then on.
A rule can be Active or Paused, and pausing it stops grading without losing anything already recorded.
What each judge measures lives in the judge, not in the rule. A rule says whose runs to grade and how often; the judge decides which rubrics it reports and how it computes them, and the rule’s instructions give it your nuance. That is why two judges on the same rule can return different rubrics — and why adding a judge that measures something new needs no change to the rule.
By default a rule grades runs that succeeded — a failed run is not sampled. Grading also keeps a shadow candidate separate from the agent it is being compared against: the candidate’s runs advance their own sampling counter, and its verdicts stay out of the production agent’s average. A live experiment therefore cannot quietly move the numbers it is being judged against.

Choosing a rate

The rate is the one knob that trades cost for confidence, and both directions have a real failure mode.
Grading costs model calls. A panel of three judges at 100% on a busy agent is three extra runs for every production run, and it shows up in Agent spend like any other agent work. Sample the everyday rules and pay for depth only where a decision hangs on it.

Reading a verdict

A verdict is one judge’s score on one rubric for one run, with its reasoning — where a rubric is the named dimension the score is reported against: quality, precision, recall, and whatever else a judge measures. The score runs from 0 to 1, shown to two decimals with its rubric name attached — safety 0.92. The chip is coloured by how strong the score is, but never labelled with a word for it: calling a number “good” or “poor” is a firmer claim than one score supports. A run graded by several judges has one entry on Runs and opens into one review per judge. Open it and each review expands into three parts: Below those sits the judge’s analysis prose, and a link straight to the run itself so you can read what the agent actually did.
The points under What moved the score are in the judge’s own scale, not the 0-to-1 score — a judge marking out of 50 reports them out of 50. Two judges’ points are not comparable with each other, and KAOP never adds them up.

Two rubrics stay two numbers

If a judge measures precision and recall, you see two scores — never their mean. A blend cannot distinguish a reviewer that invents findings from one that misses them, which is precisely the distinction two rubrics exist to preserve. Where a single headline is genuinely needed, KAOP shows the lowest rubric with its name attached, which is the honest summary of “is this run’s evaluation good”.

Human grades are kept separate

A person can grade a run directly, on a 1-to-5 star scale, and those grades are shown in their own amber styling rather than the score colouring. They are never blended into a judge’s average and never move an agent’s rubric scores, which are built from LLM verdicts only. They are not discarded, though: human grades roll up into their own Quality figure on the agent’s page — a star rating out of five, over both runs graded individually and the agent graded directly. So the two live side by side and neither moves the other.
A human grade is one person’s spot check; an LLM verdict at a known sample rate is a measurement. Reporting them as one number would let machine grading quietly inflate or mask what your team actually thinks — which is why they are kept apart rather than averaged.
Pick a window of 7d, 30d, or 90d — 30 by default — and optionally narrow to one rubric, one agent, or one judge. Every filter narrows the whole Overview tab together, so the tiles always describe the same slice the chart is drawing.
  • Tiles — Evaluations in the window, Agents evaluated, Active evaluators, and the Weakest rubric named alongside its average. Never a cross-rubric mean.
  • Score over time — a daily series per rubric, with volume. A day with no verdict is omitted rather than plotted as zero: a zero would draw a cliff on every quiet weekend that looks identical to a quality collapse.
  • Movers — the agents whose rubric average moved most against the preceding window, split into Dropped and Improved. A mover needs a baseline on both sides, so an agent first evaluated this window is not reported as the period’s biggest improvement.
  • Evaluators — each judge with its volume, its average, and when it first evaluated anything, so a newly added judge is visible as new and its early verdicts can be sanity-checked against the incumbent’s.
Counts on Overview are of LLM verdicts only. That is deliberate — the tab exists to trend a measurement, and a human spot check is not one.

Per-agent quality

An agent that has been evaluated carries its own score-over-time series per rubric, plus its best and worst runs per rubric, each linking to its review. Those two lists are what turn “3% of this agent’s output is weak” into “read this specific review”. They are computed over the agent’s whole history rather than folded from whatever page of results happens to be loaded, so the worst run is genuinely the worst run. The same view appears as an Evals tab on the agent’s own page in Fleet health, shown once the agent has LLM verdicts to plot.

Next steps

Model Evaluation

Grade a candidate against production on the same input, with a blinded judge.

Runs & evidence

The run a verdict grades, and the trail behind it.

Skills

Where a weak rubric usually gets fixed.