The shape of it
Every queued request runs on both arms. Callers only ever receive the production answer. The two executions are first-class runs, linked as a comparison pair, and that pair is what a judge grades.Deploying a candidate
A candidate is a second deployment of the same agent, and KAOP walks you through setting one up.1
Open the agent and edit it
From Operate → Fleet, open the agent you want to test and choose Edit. Change whatever
the candidate should do differently — its model, its instructions, its tools.
2
On the Review step, choose Deploy as a shadow
The last step of the wizard offers two outcomes. Update the main agent saves the change to
the agent you already run. Deploy as a shadow leaves that agent untouched and sets the
change up as a candidate instead.
3
Deploy the candidate with the environment you are given
KAOP cannot reach your infrastructure, so it hands you what only it can provide: a worker token
minted for the shadow alone, and the declaration the SDK needs. Run your candidate image as a
second deployment alongside the one you already run, with that environment.The token is shown once. It belongs to the shadow alone — the main agent’s token is
untouched, so bringing the candidate up or taking it down never rotates the credential your
production deployment is using.
AGENTOPS_VARIANT_ROLE declaration means production, so nothing changes on the serving side.
Because both arms share one agent id, Fleet health
shows one row per agent. A live candidate makes itself visible there:
- The Instances cell splits into
N production · M shadow. - A Shadow active facet filters the whole fleet to agents currently running one.
- The agent’s Remove shadow action ends the experiment.
Use the token the wizard minted, not a copy of production’s. The guard above keys on the
credential: a replica on the shadow’s token that declares no arm is refused, but one on
production’s token is a production pod whatever else it declares — it answers real callers, and
nothing catches it. A candidate also has to run the Python SDK, since that is what reads
AGENTOPS_VARIANT_ROLE.What the candidate cannot do
The whole design rests on the candidate being unable to affect anything. Specifically:
The single deliberate exception is grading. Grading a candidate is the entire point, so a shadow run
does publish for grading — with the carve-out that a shadow judge run never publishes, so a
candidate judge cannot overwrite the production judge’s verdict.
Attaching a judge
Fan-out gives you pairs; grading them is an eval rule away. Create a rule in Evaluations and set What to grade to Run pair (shadow vs. production). Rules configured that way carry a Run pair badge in the rules list. A pair rule fires on a completed pair rather than on a single terminal run, and both arms are selected together or not at all — sampling them independently would produce half-pairs, which cannot answer a comparison question.Why the judge is blinded
The judge is handed the two outputs as Candidate 1 and Candidate 2, in randomized order, with nothing in its context naming which arm is which. It sees each side’s input, output, status, error, and tool trace, and nothing that hints at role.Unblinded pairwise comparison is not offered, on purpose. A judge told which output came from
production will favour it — and the resulting number looks exactly like a trustworthy one. Blinding
is the only thing that makes a small preference meaningful, so it is not a setting you can turn off.
Why a pairwise verdict beats two separate scores
Both instruments exist and neither replaces the other.
An absolute score is only as stable as the judge’s internal scale, and that scale wobbles between
independent calls: a
0.72 beside a 0.68 cannot tell you whether the candidate is worse or the
judge was stingier the second time. A pairwise verdict deletes the scale — the judge only says which
of two things in front of it is better, which is an easier judgment and far more sensitive to a
small real difference.
The converse matters just as much: a pairwise winner from two bad answers is still a bad answer.
Keep your per-arm rubric rules running alongside the pair rule. They tell you the absolute bar is
met; the pair tells you which version to ship.
Reading the result
Open Agent Score Results for the list of agents that have been compared, then an agent for its shadow-comparison history. The same report also appears under Operate → Fleet → agent → Versions, against the version that produced it — the better entry point when you are asking “was this change any good” rather than “how are my candidates doing”.A/B Test Results
Each row is one judge’s verdicts under one eval rule against one agent’s candidate:
Expanding a row lists every graded pair with, for each rubric, the result, both arms’ scores, and
the tie band that produced the result. A tie therefore reads as “within the band” rather than as
a coin flip — the two scores are right there.
Two charts summarize the same data: Win distribution (how often each arm scored higher) and
Score distribution (both arms’ scores over time).
Cost, tokens, and latency per arm
Beside the verdicts sits each arm’s rolled-up usage — cost, token breakdown, and mean latency — with the delta expressed as production minus shadow, so a green figure means the candidate used less. Every figure carries its own sample size: how many of that arm’s runs actually contributed. This matters more than it sounds. A run can carry no usage record at all, and pricing is not resolvable for every model, so a missing cost is reported as unknown rather than as zero — “cost from 3 of 17 runs” instead of a blended figure that reads as “this arm is free”. Token counts are kept disjoint for the same reason: fresh input tokens sit beside cached and cache-write tokens rather than being summed into them, because blending fresh with cached input misprices a cache-heavy agent by roughly an order of magnitude.Deciding
The comparison gives you three independent readings — preference, absolute quality, and marginal cost — and the decision is yours to make and execute.1
Wait for volume
A handful of pairs is an anecdote. Watch the comparison count, not the win rate, until the
counts are worth reading.
2
Read preference per rubric
A candidate that wins on one rubric and loses on another has told you something specific. Do not
average it away.
3
Check the absolute bar
Confirm the candidate’s own per-arm rubric scores clear the level you require. A preference win
between two weak answers is not a release.
4
Price it
Read the per-arm cost and latency totals, with their sample sizes. A candidate that wins
narrowly and costs twice as much is a different decision from one that wins narrowly and costs
a tenth.
5
Promote deliberately
Edit the agent again, make the change the candidate was carrying, and take Update the main
agent on the Review step this time. For an agent KAOP deployed, that saves the config and
gives you an updated
values.yaml and a helm upgrade command to run against its existing
release. For an agent you deploy yourself, it saves the config and what reaches the running
agent is down to how you ship it.6
Then tear down
Remove shadow on the agent ends the experiment, and you delete the candidate deployment
yourself — KAOP cannot reach your infrastructure. Removing the shadow is also what lets an
ordinary undeclared replica of that agent register normally again.
No result ships itself. Nothing here auto-promotes, auto-blocks a release, or writes into your
serving path — the platform produces the evidence and a human makes the change. Every promotion is
therefore recorded as an ordinary agent change and shows up in the Changes figure on
Fleet health.
Next steps
Evaluations
Eval rules, judges, rubrics, and how a verdict is read.
Fleet health
Spotting a live candidate, and removing it.
Runs & evidence
Both arms of a pair, side by side, with the pair’s own id.