The tabs
Grading is itself agent work. A judge is an agent, and its grading pass is a real run against
your control plane — with an input, tool calls, and an output. That is why a verdict can carry
reasoning rather than a bare number: the reasoning is something an agent produced and the platform
stored, not a label the platform inferred.
Setting up grading
Grading starts when you create an eval rule — the configuration that decides which agents’ runs get graded, how often, and by whom. Rules live on the Evaluators tab alongside the judges themselves: New eval rule opens the form, and each existing rule has Edit.1
Name the rule and choose what to grade
Single run grades each matching run on its own — the default. Run pair (shadow vs.
production) grades a completed comparison pair together and belongs to
Model Evaluation.
2
Pick the agents
A rule names the agents whose runs it grades. Nothing outside that list is touched.
3
Set the sample rate
10%, 25%, 50%, or 100%, defaulting to 10%. The rate is per rule: the whole panel
of judges evaluates the same fraction of matching runs, so their verdicts are comparable. It is
the one knob that trades cost for confidence — see Choosing a rate below.
4
Choose the panel of judges
One or several. A judge is an ordinary agent that advertises the grading capability, listed
on the Evaluators tab; each one you select produces its own independent verdict.
5
Add rule instructions
Free text telling the judges what “good” means for these agents specifically — for example,
that a weekly report must lead with the total, or that a table missing a share-of-total column
is a defect. Optional, and the most under-used field on the form.
6
Activate
The rule starts matching new runs immediately, samples them at the rule’s rate, and dispatches
one grading task per judge. Edits apply to new matching runs from then on.
What each judge measures lives in the judge, not in the rule. A rule says whose runs to grade
and how often; the judge decides which rubrics it reports and how it computes them, and the rule’s
instructions give it your nuance. That is why two judges on the same rule can return different
rubrics — and why adding a judge that measures something new needs no change to the rule.
Choosing a rate
The rate is the one knob that trades cost for confidence, and both directions have a real failure mode.Reading a verdict
A verdict is one judge’s score on one rubric for one run, with its reasoning — where a rubric is the named dimension the score is reported against:quality, precision, recall, and whatever
else a judge measures.
The score runs from 0 to 1, shown to two decimals with its rubric name attached — safety 0.92.
The chip is coloured by how strong the score is, but never labelled with a word for it: calling a
number “good” or “poor” is a firmer claim than one score supports.
A run graded by several judges has one entry on Runs and opens into one review per judge. Open
it and each review expands into three parts:
Below those sits the judge’s analysis prose, and a link straight to the run itself so you can read
what the agent actually did.
Two rubrics stay two numbers
If a judge measuresprecision and recall, you see two scores — never their mean. A blend cannot
distinguish a reviewer that invents findings from one that misses them, which is precisely the
distinction two rubrics exist to preserve. Where a single headline is genuinely needed, KAOP shows
the lowest rubric with its name attached, which is the honest summary of “is this run’s
evaluation good”.
Human grades are kept separate
A person can grade a run directly, on a 1-to-5 star scale, and those grades are shown in their own amber styling rather than the score colouring. They are never blended into a judge’s average and never move an agent’s rubric scores, which are built from LLM verdicts only. They are not discarded, though: human grades roll up into their own Quality figure on the agent’s page — a star rating out of five, over both runs graded individually and the agent graded directly. So the two live side by side and neither moves the other.A human grade is one person’s spot check; an LLM verdict at a known sample rate is a measurement.
Reporting them as one number would let machine grading quietly inflate or mask what your team
actually thinks — which is why they are kept apart rather than averaged.
Trending quality over time
Pick a window of 7d, 30d, or 90d — 30 by default — and optionally narrow to one rubric, one agent, or one judge. Every filter narrows the whole Overview tab together, so the tiles always describe the same slice the chart is drawing.- Tiles — Evaluations in the window, Agents evaluated, Active evaluators, and the Weakest rubric named alongside its average. Never a cross-rubric mean.
- Score over time — a daily series per rubric, with volume. A day with no verdict is omitted rather than plotted as zero: a zero would draw a cliff on every quiet weekend that looks identical to a quality collapse.
- Movers — the agents whose rubric average moved most against the preceding window, split into Dropped and Improved. A mover needs a baseline on both sides, so an agent first evaluated this window is not reported as the period’s biggest improvement.
- Evaluators — each judge with its volume, its average, and when it first evaluated anything, so a newly added judge is visible as new and its early verdicts can be sanity-checked against the incumbent’s.
Per-agent quality
An agent that has been evaluated carries its own score-over-time series per rubric, plus its best and worst runs per rubric, each linking to its review. Those two lists are what turn “3% of this agent’s output is weak” into “read this specific review”. They are computed over the agent’s whole history rather than folded from whatever page of results happens to be loaded, so the worst run is genuinely the worst run. The same view appears as an Evals tab on the agent’s own page in Fleet health, shown once the agent has LLM verdicts to plot.Next steps
Model Evaluation
Grade a candidate against production on the same input, with a blinded judge.
Runs & evidence
The run a verdict grades, and the trail behind it.
Skills
Where a weak rubric usually gets fixed.