Skip to main content
See what’s new in the Komodor Agentic Operation Platform (KAOP). Updates are listed by release date, with the newest first. They are the same updates you see under What’s new in the product.
Improved

Improved: A dataset’s page leads with its latest replay

A dataset’s header now shows four things: the replay agent, the cases (and how many are switched on for replay), the latest replay with its verdict, and when it was created. The default replay agent and Minimum coverage moved into Settings in the dataset’s menu, next to Copy dataset ID. The Live/Candidate filter on Run history is gone — each replay already names the agent version it ran against.Try it in KAOP
Improved

Improved: Add a run to a dataset and fix its output in one step

Add to dataset on a run now shows the run’s output, and you can correct it before saving. What you add is ready to replay straight away; there is no separate approval step. On a dataset’s page each case has a Replay switch to include or exclude it, Edit output in its row menu, and clicking a case opens the run it came from. Cases added by an agent or a capture rule wait with Replay off until someone turns them on. The Recorded cases tab and the review screen are gone.Try it in KAOP
Improved

Improved: A replay compares two answers, not three

A dataset replay now shows the Recorded answer beside the Replay answer, next to the evaluator’s verdict. The middle “Original recorded” column is gone: approving a recording makes its recorded output the answer a replay is judged against, so the two columns always showed the same thing.Try it in KAOP
Fixed

Fixed: A tag you type is saved even if you don’t press Enter

Tag fields used to keep only the tags you had turned into chips with Enter. Text still sitting in the box when you clicked Save was dropped, so Edit tags on a dataset could report “Dataset tags updated” while saving nothing. Now any text left in the box becomes a tag as soon as you click away, on every tag field: dataset and record tags, capture rules, scenarios and integrations.Try it in KAOP

Fixed: Pick tools in the candidate wizard again

The Tool selection picker on a candidate’s Change step said “No tools in the catalog yet” even when your MCP servers advertised tools, because the catalog request was rejected. It now loads your tools, and if the catalog ever cannot be loaded the step says so and offers a retry instead of showing an empty list.Try it in KAOP
NewImprovedFixed

New: Tag datasets and filter the Datasets list by tag

A dataset can now carry its own key=value tags (team=payments, env=prod). Add them in New dataset, or change them any time with Edit tags on the dataset page.The Datasets list shows each dataset’s tags, and its tag chips filter the list: pick several and only datasets carrying every one of them stay. Dataset tags are for finding datasets — which records a replay runs is still chosen by the records’ own tags.Try it in KAOP

Improved: Set tool permissions before you deploy

The Tool permissions card on the Add-Agent wizard’s Permissions step used to tell you to come back after the deploy: a permission attaches to an agent, and an agent deployed from the catalog only gets its id at the very end. So the one screen that asks which tools an agent may run on its own, and which ones need a person to say yes first, was the one screen you could not answer there.You can answer it now. Author the policy while you are building the agent, and it is applied with the deploy — the agent’s first run is already governed by it. The last step says Done rather than Save, because nothing is written until the deploy goes through.The tool picker works too. It offers the tools this agent will actually be able to reach, read from the integration scope you chose one step earlier — not the whole account’s catalog, which would let you write a rule over a tool the agent has no route to: armed, matching nothing, and silent about why.Try it in KAOP

Improved: Scope a catalog agent while you deploy it

Deploying an agent from the catalog used to skip the Permissions step: the roles it holds and the secrets it may fetch were something you set afterwards, from the agent’s Permissions tab in Fleet, once it was already running. The step is on that rail now, and what you choose there rides the deploy itself — there is no window in which the agent exists unscoped.A handful of workers do not have the step, and would not use it. The workflow orchestrator is scoped by the roster its Workers step defines, and each specialist it delegates to carries its own roles. The Golden Scenarios workers, the base evaluator and the memorizer are invoked by a pipeline or by the platform rather than by you, so their scope is not yours to set. Every other catalog agent gets the step, and the tool-permission card on it, exactly as a hand-built one does.Try it in KAOP

Improved: Replaying a dataset has its own permission

Replaying a dataset now needs the golden_scenario.run permission as well as permission to invoke the agent it replays against. Developers, admins, owners and automation hold it by default; members and operators, who can read datasets, no longer see Replay now or Replay by tags enabled. A custom role that replayed datasets before needs golden_scenario.run added.Try it in KAOP

Fixed: Replays show their evaluator, filter by candidate or live, and stay out of bulk capture

A dataset replay now records and shows the evaluator’s model and prompt version on its page, next to whether it ran on the live agent or a candidate, and a dataset’s Replays list filters by All, Live or Candidate. Capture to dataset from History no longer picks up dataset replays as new cases.Try it in KAOP
NewImprovedFixed

New: Standing capture rules automatically add runs to a dataset

Quality Lab → Datasets now has a Capture rules tab. A rule watches an agent’s succeeded runs and, on a match, captures them into a dataset automatically — tagged, sampled at whatever rate you choose, and always left awaiting approval so nothing ships without a person reviewing it. Pause, edit or delete a rule at any time.Try it in KAOP

New: Review awaiting-approval cases one at a time

Quality Lab → Datasets now has a review screen that walks through awaiting-approval cases one at a time, scoped to a single dataset or the whole case library. Each case shows its input, tool calls and results, and expected output, with Approve (with an optional note) and Skip actions and a running “N remaining” count.Try it in KAOP

New: Rerun history shows candidate vs live, and which evaluator judged it

A dataset rerun’s history now says whether it ran on a Candidate or the agent’s live version, and records the judge’s model plus its prompt version. Rerun history can be filtered by agent_variant=live or candidate.Try it in KAOP

Improved: One verdict per rerun — pass, failed or inconclusive

Every dataset rerun now shows a single Verdict: pass, failed, or inconclusive. A replay gap (a tool call with no recorded answer) always means inconclusive, even if the judge still reached a same/different opinion — that opinion stays visible, it just no longer counts as the verdict. A contaminated run (a live call happened) keeps its own distinct label rather than being folded in.Try it in KAOP

New: See which org and repositories a GitHub connection reaches

Open View details on a GitHub App connection in Integrations to see the organization it is installed on, whether it covers all repositories or a selected set, and the repositories themselves — searchable, with a link to manage the installation on GitHub. The dialog also shows who connected the integration. The list is read live from GitHub and cached for ten minutes; Refresh fetches it again on demand.Try it in KAOP

Improved: Funnel drill-down rows open the incident

Clicking a terminal stage of the alerts funnel lists the incidents that ended there. Each of those rows is now a link to the incident’s own page, so the drawer is a way in rather than a dead end. The link closes the drawer as it navigates; open it in a new tab or from the keyboard and the drawer stays where it was. Alert rows are unchanged — an alert has no page of its own.Try it in KAOP

New: Edit a recorded case’s expected outcome, with full history

A recorded case’s expected outcome or recorded tool responses can now be replaced without losing the original. Editing creates a new version that starts awaiting approval — the original recording and every replacement stay visible in the case’s decision history, with who made each change, when, and why. A run that already used the old version keeps using it; only future runs see the edit.Try it in KAOP

New: Delete a dataset

Quality Lab → Datasets can now delete a dataset, from the trash icon on its row or Delete dataset on the dataset page. Its records are kept — they stay under Recorded cases and in any other dataset — and every finished replay keeps its results. Delete is unavailable while a replay of the dataset is queued or running, or when the dataset is declared by a cluster.Try it in KAOP

Fixed: Dataset reruns no longer skew an agent’s run history and metrics

A dataset replay run used to show up in its agent’s own run history and count toward its success rate, run counts, activity, and cost panels — making a batch of regression reruns look like a spike of real traffic. Replays are now excluded from an agent’s history and metrics by default, while staying fully reachable from the dataset’s own rerun history and by direct link.Try it in KAOP

Improved: Dataset replays work for kubectl agents, with a minimum coverage per dataset

Replaying a dataset of a Bash/kubectl agent now answers a kubectl command from the recording even when the agent spells it differently, and the agent’s skills load as in a normal run. Each dataset has a Minimum coverage (default 80%): a replay with a few unanswered calls is judged by the evaluator as long as enough of its calls came from the recording. The replay page shows the coverage and, on each call, how the recording answered it.Try it in KAOP

Improved: See a dataset’s health at a glance — including which agents fed it

A dataset’s counts (total, approved, awaiting approval), when a case was last added, and when it last ran are now computed by the server and available on every surface — UI, API and MCP. The dataset list and detail pages also show which agents supplied its recordings.Try it in KAOPThe Documentation link on a catalog agent’s card, and on a deployed agent’s detail panel, used to point at a private Komodor repository — so it opened a GitHub 404 for everyone outside Komodor. Each of the fourteen catalog agents that has a published page now links straight to its own section of the Komodor docs, and the few with no page yet simply show no link rather than a broken one.Try it in KAOP

New: Capture a filtered list of runs to a dataset in one action

The run History page can now capture many completed runs into recorded cases at once, from whatever filter you already have applied (agent, status, text, time range). Already-captured runs are skipped, only succeeded runs are eligible, and you can tag every captured case and drop them straight into a dataset in the same action.Try it in KAOP

Improved: The alerts funnel prices each stage

On a workflow’s Overview, the Investigation and Remediation headers now read as two lines: the count, then Total $12.40 · P95 latency 9m. The number is what the agent runs that stage’s steps dispatched spent, over every run behind the window’s incidents — investigating is priced under Investigation and remediating under Remediation, so the two add up to the runs’ whole spend. A stage nothing priced shows no dollar figure rather than $0. The header is now the one place a P95 is drawn; the line under a bar keeps its share and hint, and the node’s own P95 moves to its tooltip.Ingest and Correlation keep the count alone: nothing before an incident opens runs an agent, and both happen inside the webhook request, so neither has a spend or a latency worth a line.Try it in KAOP
NewFixed

New: Use your own agent as an evaluator

Turn on Evaluator when you add an agent and it appears in Quality Lab as an evaluator you can put on an eval rule. A rule either grades against its evaluator’s default criteria, shown on the criteria step, or against criteria you define; criterion weights are now sliders that always add up to 100. The platform checks every evaluator’s answer against the expected shape and asks it to fix a malformed one before the score is recorded.Try it in KAOP

Fixed: Replays are judged on the agent’s answer

Your Base Evaluator now reads the rerun’s own answer when it judges a dataset replay. Before, it was shown the replay’s bookkeeping instead, so a replay that reached the approved conclusion could still be marked different and flagged as a regression.Try it in KAOP
NewImproved

Improved: Tag dataset records and take them out of a dataset

Each record in a dataset now shows its tags, which you can edit in place, and has a Remove from dataset action that takes it out of that dataset without deleting it. On a run that has not succeeded, Add to dataset now stays visible but disabled, and says why the run cannot be captured.Try it in KAOP

New: See every recorded case in one place, and clear the ones awaiting approval

Quality Lab → Datasets now has a Recorded cases tab: every case captured from a run, with its approval status, tags and the datasets it sits in. Filter to Awaiting approval to see how many are waiting and approve them right from the list. A replay now shows Expected, Before (the original run’s output) and After side by side, and recordings no longer clutter the Golden Scenarios library.Try it in KAOP

New: Define weighted criteria on an eval rule

An eval rule can now carry custom criteria, each with a name, description, and weight, plus pass and warn thresholds. Every criterion becomes its own score you can trend; beside them you get a weighted overall and a Pass / Warn / Fail badge. Need to steer how a criterion is judged? Add it to the rule’s own instructions. Only evaluators that accept custom criteria appear in the picker when criteria are on.Try it in KAOP

Improved: Datasets hold any agent’s runs, and you pick the agent to replay against

A dataset is no longer tied to one agent. Add to dataset on a finished run now offers every dataset. A run already in one dataset can be added to another, and it stays the same record, so approving or re-tagging it shows in both. On a dataset page, Add case brings in a record captured from any agent. Remove from dataset takes a record out of that dataset only: it is not deleted, and it stays in every other dataset that holds it.Replay now asks which agent to replay against, with the dataset’s default agent picked for you, and each replay shows the agent it ran against. A dataset no longer needs an agent when you create it.Try it in KAOP

Improved: Only people approve dataset records, and capturing and approving are their own permissions

Approving a dataset record, or removing its approval, is now a person’s decision only: the assistant and other agents can capture runs and organize datasets, but they can no longer approve an outcome. Capturing a run, approving an outcome and managing datasets are separate permissions, so a role can be given one without the others. Replaying a dataset needs permission to invoke its agent, not a dataset permission.Try it in KAOP

New: key=value tags on cases, and a faster way to build a dataset

A case’s tags are now key=value pairs (env=prod, not a bare word) — edit them from any dataset that holds the case, via the new Edit tags action on each record row.A dataset’s header gains Add by tag, which pulls in every live case carrying a pair you give it, and a new Tag coverage section shows how many cases and approved cases each tag has, with an Untagged row alongside — click a tag to filter the records below to it.Replay by tags narrows a whole-dataset replay to cases that carry (or don’t carry) given pairs, with a live preview of how many approved cases will run before you start it.Try it in KAOP
NewImproved

New: Ask Slack what it can do, and who answers you here

Type the AgentOps slash command on its own, or with help, and it lists what it can do rather than guessing what you meant.Type it with here and it tells you, privately, which agent answers you in that conversation and what is stopping you if nothing does: an allowlist you are not on, an agent that is switched off, a Slack account you have not connected yet, or no routing rule for that conversation at all. Until now each of those looked identical from Slack, which was silence.Both work before you have connected your Slack account, which is when people are most lost, and both answer only you.Try it in KAOP

New: See whether a replay reached the approved outcome

Every completed dataset replay is now judged same or different against the outcome you approved by your Base Evaluator, which checks the conclusion rather than the wording, and a different outcome is flagged as a regression. The replay page shows the evaluator’s reasoning beside the verdict, and the approved expected outcome, the original recorded output and the rerun’s output side by side, next to both tool transcripts. The Datasets page shows whether your evaluator is live.Try it in KAOP

Improved: See whether a replay got more expensive or slower

Every dataset replay now shows its cost and duration next to the original recorded run, with the change in dollars, time and percent. Any increase is flagged, so a version that answers the same way but costs more or takes longer stands out, without changing the replay’s outcome.Try it in KAOP

Improved: Skills can carry many files, and every agent can use them

A skill is no longer just its SKILL.md. Add references, scripts and assets to it, and your agents get all of them, in the layout their framework expects.Claude Code and ADK agents load the skill with their own built-in skill support and open the file they need when a task calls for it, including files the skill’s instructions never link to. A Claude Code agent can also run the skill’s scripts.Try it in KAOP

Improved: Edit one file of a skill from the co-pilot or the API

The chat co-pilot can now read a skill’s files and add, replace or remove one at a time, and the API has the same with PUT and DELETE /api/v1/skills/{skill_id}/files/{path}. Each edit publishes a new version that keeps the body and every other file, so nothing is resent.The agent now sees every file too: the SKILL.md it reads ends with a list of the skill’s files by path, so a file the body never links is still found.Try it in KAOP

Improved: Delete a dataset record without losing its replays

Each record in a dataset now has a Delete action. The record leaves every dataset and can no longer be replayed, but its source run and every finished replay stay, marked Deleted, and the dataset’s replay totals do not change. A record, or a dataset, with a replay still queued or running cannot be deleted until that replay finishes.Try it in KAOP

Improved: Dataset records show who captured them, and what the run cost

Each dataset record now shows who captured it, plus the original run’s cost and duration. Only a run that succeeded can be added to a dataset, so failed and cancelled runs no longer offer Add to dataset.Try it in KAOP

Improved: Approving a dataset record keeps its recorded output, and approval can be removed

Approving a dataset record now always makes the run’s original output the expected outcome, and records who approved it, when, and why. An approved record has a Remove approval action that sends it back to awaiting approval: future replays skip it, and replays that already ran keep their results. A record can no longer be edited, archived or published without review.Try it in KAOP
NewImprovedFixed

Improved: Pick the tools before a server connects, and find them in a long list

Adding an MCP server used to publish every tool it advertised, and the Configure step let you save without looking. A server now starts publishing nothing and will not let you continue until you say what an agent may call — the disabled button tells you why on hover.Long tool lists are navigable now. MCP tools fold into Read-only and Writes groups, OpenAPI operations into GET, POST, PATCH and DELETE, each collapsed with a running 12/40 selected count and its own select-all. An upstream advertising hundreds of tools fits on one screen instead of scrolling past the destructive ones.A server that answers but advertises no tools at all is now refused with that reason, rather than saving as something an agent can reach and never use.Try it in KAOP

Fixed: Autoscaling a self-hosted agent now adds a replica per queued run

Turning on KEDA autoscaling for a self-hosted agent used to start you at a queue depth of five per replica, so six runs had to pile up before a second replica appeared. An agent takes one run at a time, so those five were waiting on capacity that could never exist.The field now starts at one, matching the Helm chart, and one queued run brings up one replica. Agents already deployed keep whatever value you set; change it on the agent’s scaling settings if you want the new default.Try it in KAOP

Fixed: Insights now counts the same 30 days everywhere on the page

The Run status card and the Run trend chart on Insights were counting a different set of runs from the 30-day tiles above them — the newest 100 runs of all time, rather than the last 30 days. On a workspace with more than 100 runs the card could not show the window at all, and on an older workspace it counted runs from outside it, so the card’s own numbers did not add up to the total beside them.Both now read the same 30-day window as the rest of the page, and the card’s outcomes add up to the run total it reports. Cancelled and queued runs, which were being dropped entirely, are included.Try it in KAOP

Fixed: Correcting your LLM gateway’s URL now reaches the agents already using it

Editing a “Managed by me” gateway’s base URL changed the provider but not the agents bound to it. Each agent holds its own copy of the address, taken when it was bound, so workers kept dialling the old endpoint while the Providers page, the agent’s own page and the generated values.yaml all showed the new one. The only way out was to move the agent to another provider and back.The edit is now carried to every agent already on that gateway, matched on the address each one actually holds — so an account running several gateways only moves the agents on the gateway that changed. The next run picks it up; no redeploy needed.Try it in KAOP

New: See exactly what a run and a deployed agent executed with

A run’s Configuration tab now shows the resolved config version, worker build, and every credential bound to it by reference — kind, source, version, and which keys were delivered or denied. A deployed agent’s Used secrets card carries the same detail, and its Versions tab adds a Deploy history card showing who created, edited, redeployed, or archived it, and when, straight from the audit log.Try it in KAOP

Fixed: Editing a catalog agent now produces a values.yaml helm accepts

Editing an agent deployed from the catalog produced a values.yaml with an empty agent.instructions, and the chart requires that field, so the generated helm upgrade failed with agent.instructions is required before it applied anything. The wizard had stopped short of the file for these agents; once it started producing one, the file itself was unusable.A catalog worker’s prompt ships inside its image and is never written into your values, so the file now carries the same standing-in line every catalog release already has, with a comment saying so. Nothing about the agent changes and no prompt is exposed.Try it in KAOP

Improved: Author a whole skill folder in the editor, not just SKILL.md

A skill is a folder — SKILL.md plus the references/, assets/ and scripts/ files its body links to — and the editor now writes all of it. Each file is a tab beside the body, a script can be marked executable, and the agent gets the folder back on disk exactly as you laid it out.The SKILL.md tab also has a Preview toggle, so a body reads as rendered markdown instead of raw source.Editing a body no longer touches the folder: publishing content on its own keeps every file the skill already had, whether the publish came from this editor, the chat co-pilot, the API or a GitOps Skill resource, which can now declare the whole folder too.Try it in KAOP

Improved: The alerts funnel reads as one picture

On a workflow’s Overview, the Ingest bar now spans the whole chart. Everything that arrives is ingested — folding a duplicate or dropping an alert on a filter is a decision taken afterwards — so that bar is the 100% the rest of the funnel is read against, however much of the volume the noise exits take.A provider’s “resolved” notice is no longer counted as an alert or drawn as a Recovery signal stage. It closes an incident whose firing alert was already counted at Ingest, so counting it again made the ingested total and every share of it too high. The incident it closed is charted under Auto remediation instead: it resolved without anyone acting on it, whatever the run had applied.Each stage’s number sits centred over its own stage and stays inside it, instead of running into the column beside it, and Remediation now carries how long a plan waited on a person alongside the count.The Show values table is gone: its numbers were already on the graph. The two readings the chart cannot draw — a stage nothing reached this window, and a caveat on a number that is known to be biased — are printed under the legend instead.Try it in KAOP

Fixed: A workflow that finished is no longer charted as Failed

Failed on the alerts funnel now means one thing: a run that broke. It is read from the workflow’s own step records — the run failed outright, or a step in it failed or timed out — instead of from the incident’s status, which carries the same word for a run that broke and for one that did not.Three populations move out of Failed as a result. A run that walked every step of its graph and stopped because nothing was left to do is charted where its steps got to, not as a breakage. So is a run an operator cancelled. And an incident whose workflow never ran at all — every incident on an account with the engine switched off — now lands under Triaged directly, which says why the engine never ran, rather than under a failure that never happened.A run still in flight stays under Still running even if one of its steps threw and is being retried.Try it in KAOP
New

New: Approve a paused action, and stop it asking next time

When an agent pauses on a governed action and you approve it, you can now approve and stop asking — one click that releases this call and edits the agent’s action policy so this tool stops pausing in future, the way Claude Code’s “Yes, and don’t ask again” works. It allows the tool for any arguments — the agent stops asking about that tool from then on.It is honest about what it can do. Because a stricter rule always wins, simply allowing a call does nothing while a broader rule still asks about it — so the control edits the rule that asks, and when it can’t (a pattern it can’t cleanly narrow, or a rule that refuses the call outright) it tells you so rather than writing a rule that changes nothing. Editing the standing policy needs the action_permission.manage permission, so an approver who may answer but not change the rules still gets the plain Approve, and every rule change is recorded in the audit trail. The affordance shows up wherever you answer — the run, the Actions inbox, and in chat — behind the actions_approvals flag.Try it in KAOP
New

New: Watch an agent work in Slack, and find one without leaving

Ask an agent something in Slack and the thread shows what it is doing while you wait, step by step, with a link to the full run. A run parked on an approval says it is waiting on a person, with the deadline in your own timezone, so a working agent no longer looks the same as a stuck one.If you do not know which agent to ask, /agentops agents lists the ones you are allowed to run, and /agentops agents kubernetes searches them by name. Each result carries the handle to type, and asking costs you nothing: the list is private to you and starts no run.Agents also read less of your channel than before. Each channel carries a history window, a week by default, and an agent cannot read further back than that however it asks.Try it in KAOP

New: Turn any finished run into a replayable dataset record

Open a finished run and choose Add to dataset: the input, every tool call and the output are recorded. Approve the record under Quality Lab → Datasets, then Replay now to run the agent against the recording. Each replay says whether it was clean, inconclusive (a call had no recorded data) or contaminated (a live call slipped through), and shows Expected, Before and After side by side.Try it in KAOP
ImprovedFixed

Fixed: An agent’s Versions tab lists real changes, not every image rebuild

An agent got a new version every time its worker image was rebuilt, because the build commit was stamped onto each of the agent’s skills and counted as part of the agent’s identity. The version history filled up with entries whose diff was empty, and the genuine changes were buried among them — one code reviewer showed 60 versions covering 7 actual edits, the most recent of which was five weeks older than the newest entry claimed. Rebuilding an unchanged agent no longer creates a version, and a version that really does carry no change to the definition now says so instead of opening an empty diff.Try it in KAOP

Fixed: The right-sizing policy picker explains an account with no policies yet

Picking a right-sizing policy for a workload used to open onto a blank list on an account that hadn’t created one yet, with no explanation and no way out. The picker now disables itself and shows “No policies yet” with a “Create your first policy” link straight to the policy drawer, so you’re never staring at an empty dropdown wondering what’s missing.Try it in KAOP

Improved: One webhook endpoint can feed several agents, and every one of them fires

An endpoint used to serve exactly one destination, and dispatch used to stop at the first one it found — so an endpoint pointed at an agent silently starved the incident workflow listening on the same URL: the request returned 200, the agent ran, and no incident was ever opened. Endpoints now attach to as many agents and workflows as you need, and a single inbound request reaches all of them, incident workflows included. You pick the endpoint from the agent’s own trigger step rather than routing it when you create it, so the endpoint wizard ends at Test & mapping; the Endpoints table’s new “Used by” column names everything listening on each one. Unhooking an agent detaches it and leaves the endpoint’s URL and token working for everything else.Try it in KAOP

Improved: The co-pilot now lives in the top-right header

The Co-pilot button has moved from the floating bottom-right corner into the global header, as the first control in the top-right cluster — before Search, Demo and Notifications. It no longer sits on top of a wizard’s own buttons, so nothing behind it is lost.Opening it still slides the co-pilot in beside the page and resizes the content to make room, and it stays aware of the page and step you’re on. On a narrow window the button collapses to an icon with a tooltip.Try it in KAOP

Improved: Awaiting approval now follows the run, not just the agent

A run paused on a governed action used to read “Awaiting approval” only in the Actions inbox or on the run’s own page — everywhere else it looked like it was quietly working. The badge now follows the run wherever it’s listed: Overview’s recent runs, the History ledger, an agent’s own run history, and an incident’s orchestrator and specialist runs.When a specialist or orchestrator is paused on a governed tool call, the incident’s Workflow runs row now shows “Needs your answer” and the incident increments both the “Needs your answer” and “Needs your attention” facet counts. The incident header badge and the Incident feed tab count update the same way, and the specialist row in the agent team panel marks which agent is waiting. The “Needs your answer” help text now names this as a third kind of pause alongside a remediation choice and a guardrail hold.Try it in KAOP
ImprovedFixed

Fixed: The workflows list counts every incident, not just the open ones

The Incidents module’s workflows table showed each workflow’s open incident count under a column headed plainly “Incidents”, so a workflow whose incidents had all been resolved reported zero — on the same row as the resolution time and the spend those incidents produced, and next to an Overview that reported them. The column now counts every incident triggered in the selected window, matching the Overview’s Volume, and names the open subset beneath it when there is one.Try it in KAOP

Fixed: A workflow’s field-mapping override is saved, and the Review step says so

The Field Mapping section of the workflow wizard’s Trigger step confirmed an override in three places — the “1 field overridden” badge, the line naming what decides the field, and the filter’s own field list — while the value reached no request and no column. Finishing the wizard produced a workflow that still read every field the way its endpoint does: wrong titles, and a dedup key that silently splits or merges incidents. Re-opening it showed the section back at “Inherits the endpoint’s mapping” with the box empty.An override now rides the same save every other field on the step does, on create and on edit alike, and a field nobody touched still follows the endpoint rather than being restated. The Review step carries its own Field mapping card listing each overridden path, and the Dedupe card names a dedup key this workflow overrode instead of reporting one it inherited.Try it in KAOP

Improved: Tooltips now stand out from the page they float over

A tooltip’s surface used to be the same colour as the card underneath it — literally the same design token — so it was told apart only by a thin border and a shadow. In dark mode that meant dark text panels on dark cards; in light mode, white on white.Tooltips now render on their own surface, dark in light mode and light in dark mode, so the content reads as an overlay at a glance. The help-icon tooltips on the cost cards already looked this way; every other tooltip in the product now matches them.Try it in KAOP

Fixed: The orchestrator is not its own specialist

An incident’s Agent team card named the same agent twice — once as ORCHESTRATOR and once as SPECIALIST — so a workflow with no specialist bound at all read as a two-agent team. The cause was one level down: every step of a workflow dispatches to the same step-orchestrator agent, and only the first of those runs is recorded as the orchestrator, so each later step was filed under a specialist role while being the orchestrator itself.The card now lists only agents other than the orchestrator, and the Investigation summary counts the remainder in runs rather than in agents — 3 (2 orchestrator runs + 1 specialist run) instead of 3 (1 orchestrator + 2 specialists), with a side that has no runs left out entirely. Every step run is still reachable from the step breakdown, and still counted.Try it in KAOP

Fixed: Steps say when a specialist never ran

A workflow step’s status reported on its orchestrator run alone. So a step whose bound specialist could not be dispatched — the agent offline, its run failing with no worker to serve run — completed green, with no failure reason, while the investigation behind it had never reached the cluster. The failure existed only on a sub-run one level down, and the incident’s own specialist list named the dead run without a status, so it read as a specialist that took part.The step row now carries a 1 specialist failed marker without being expanded, and expanding it names the agent and the error its run recorded. The Agent team card annotates a specialist that failed or was cancelled with its outcome instead of listing it like any other. The step itself is still reported honestly as completed — its own run did succeed — but an operator being asked to authorize a change built on that investigation is now told what it could not see.Try it in KAOP

Fixed: The investigation time limit now says when your account cannot change it

Changing the orchestrated step time budget away from 900 in the incident workflow wizard failed the whole activation, with “This workflow has no workflow definition yet … so there is nothing to set orchestrated_step_max_duration_seconds on.” The limit is stored on the workflow definition, and an account whose workflow engine is disabled — or that has no step-orchestrator agent deployed — never gets one stamped, so the value had nowhere to land. The field was offered anyway, and every save in between was rejected without saying so.The field is now read-only on those accounts and states which of the two reasons applies, the same way the workflow template picker already did. This covers both the current wizard’s Investigation team step and the new wizard’s Investigation time limit. Leaving it at the default always worked and still does.A rejected auto-save also no longer passes in silence: the wizard keeps a banner up saying your last change was not saved and what the API said, at the step where it happened rather than three steps later at Activate. Navigation is still never blocked.Try it in KAOP

Fixed: Set up notifications without leaving the wizard

Picking where a run or an incident workflow reports back used to be a dead end on an account with no notification sink yet: the step said “Nothing here yet” and offered nothing to click. It now explains what a sink is for and opens the Add-notification-sink wizard right there — Slack or a signed webhook — so you can finish the setup in one pass.

Fixed: The incident timeline reads in your own timezone

An incident’s Timeline card printed each entry’s timestamp exactly as the control plane stores it — 2026-09-16T10:38:29.777697+00:00. Every reading of “when did this alert land, and how long did triage take” started with converting UTC in your head.Those timestamps now render on your own clock, as Sep 16, 1:38:29 PM. Seconds are kept, because a live incident routinely logs two entries a second apart, and hovering one still reveals how long ago it was.Try it in KAOP

Fixed: The alerts funnel stops calling a completed workflow “Failed

A workflow that runs to the end and applies nothing marks its incident Needs attention, so a person reads the findings and decides what to do. The funnel treated that status as a failure and charted the incident under Failed — telling an operator their workflow had broken on a run that had done exactly what its template asked of it, investigated the alert and reported that the step meant to apply a change could not act.The chart now reads the run’s own status alongside the incident’s. A run that completed is placed at the stage its step runs earned, which for this population is No remediation, reached through its RCA verdict like any other finished run; the terminal’s hint counts them, so “1 closed by hand · 2 completed without an applied change” says which incidents are behind the bar. Failed keeps the incidents it was written for: one that errored, and one that needs attention because its run stopped with nothing left to retry. Those incidents are still waiting on a person — the Needs human tile on the overview counts them, and the bar drills through to every one.Try it in KAOP

Fixed: The alerts funnel charts incidents whose workflow graph never ran

The funnel’s stages are read from a workflow run’s step runs, and an incident triaged by the orchestrator directly has none. Those incidents were counted at Investigated and placed nowhere to the right of it — so an account whose workflow graph does not run saw an empty right half and a headline of “0 of N remediated”, however much work its agents had actually done. That is not only old data: an account still on the legacy dispatch produces these today.Every stage those incidents can answer for themselves is now drawn. Status places a failed or in-flight one, the incident’s own root cause places the RCA verdict, and a close recorded outside any run lands under No remediation with the hint saying which. What none of them settles ends at Triaged directly, reached straight from Investigated rather than through an RCA verdict, with a caveat saying no remediation stage can be proven either way and a hint naming why the graph never ran. Every bar is clickable through to the incidents behind it.Try it in KAOP

Fixed: The endpoint tester says what actually happened to a request

Testing an endpoint told you a request had arrived, then ended with “Captured — no target is wired yet, so nothing ran.” On an endpoint feeding an incident workflow that was never true: the workflow attaches to the endpoint and opens an incident rather than starting a run, so the panel’s one question — is there a run? — always answered no, even for an alert that had just created an incident.Each captured request now reports what every incident workflow on the endpoint did with it: routed, filtered out, parse failed, throttled, ignored or errored, with the workflow’s name, the reason, and a link to the incident it opened. Filtered out is the one worth having — it names the condition that rejected the alert, so a filter that is quietly dropping everything no longer looks exactly like a quiet week. The same detail now appears in the endpoint’s History drawer.“Nothing ran” is still there for the endpoint it was written for: one with no target and nothing consuming it.Try it in KAOP

Fixed: Moving a self-hosted agent onto your own LLM gateway now takes effect

Editing a self-hosted agent onto a “Managed by me” provider recorded the choice but could leave Komodor’s gateway credentials still bound to it. The agent kept inferring through Komodor — billing to Komodor’s budget and not to your gateway — while the wizard, the agent’s detail page and the generated values.yaml all reported the new provider. Saving again did not help: the selection already matched, so nothing was re-applied.The save now reconciles against what the agent actually holds rather than against what was last recorded, so an agent left in that state is repaired by the next edit. An agent deployed from the catalog can be moved between providers too, which it previously could not be at all — its Edit wizard refused to produce a values.yaml and its provider was fixed at whatever it was deployed with.Deploying a catalog worker onto your own gateway works too. It previously accepted the choice and bound Komodor’s credential anyway, with no field to enter your gateway key in.The Permissions step no longer lists the gateway credentials AgentOps manages for you (LLM_PROXY_TOKEN and friends). They are decided by the Model step, and a checkbox that disagreed with it could only break the agent.Try it in KAOP

Fixed: Edit keeps the in-cluster MCP server you picked at create

A self-hosted catalog agent deployed with “Use your own in-cluster MCP server” lost that choice the moment you re-opened Edit: the toggle came back off, the URL and Kubernetes Secret fields were empty, and the regenerated values.yaml quietly dropped the agent.mcpServers override — so applying it sent the worker back to the MCP server baked into its image. The create wizard now records the URL and the Secret name/key with the agent, Edit re-opens on them, and the Review step lists In-cluster MCP server among the changes waiting for your helm upgrade instead of claiming there are none. A catalog agent that was configured here also no longer opens under the “recovered from the running agent” warning.Try it in KAOPOn the incident workflow wizard’s Investigate & remediate step, “Deploy from catalog” listed the catalog one full-height card at a time in a narrow modal — fourteen specialists, about two of them on screen, and no way to type a name.The dialog is wider now, lays the workers out in two columns, and clamps each card’s description so a screenful is six workers rather than two. A search box filters by name, category or what the worker does, within the role scope rather than out of it, and the footer keeps naming the worker you picked even once a search has scrolled it out of view.Try it in KAOP

Fixed: Cancel no longer leaves a draft workflow behind

The incident workflow wizard saves a draft the moment you continue past the first step, so the rest of the wizard has something to write to. Cancel then navigated away without mentioning it, and the draft stayed — a catch-all workflow nobody chose to create, sitting in the same list you read to judge what the account is running.Cancel now says what it is holding and asks. Discard draft deletes it; Keep draft leaves it in the workflows list to finish later. Cancelling before the first Next still creates nothing, and cancelling out of an edit never offers to delete the workflow you opened.Try it in KAOP

Fixed: Budgets now stop runs mid-flight, and blocked runs explain themselves

A block new runs budget is meant to halt an agent the moment its window spend crosses the cap — even for a run already in progress. That mid-run enforcement had stopped firing, so a long run could sail past the cap; it now terminates in-flight runs again.When a run is blocked or terminated by a budget, the run page now says so plainly — “Stopped by budget, not a run failure” — with the cap, the spend so far, and an Increase budget button that opens the edit dialog for that exact window, with a note on why you landed there.Try it in KAOP

Improved: An agent’s tool permissions are one screen, starting with what it does by default

Writing an action permission meant filling in a dialog on the Actions page — which knows about no particular agent — and the one question that decides the most had no field at all: what does this agent do about a tool call nothing covers.Now it is the first step. Every agent gets a default — let everything run, ask a person first, or refuse everything — and rules become overrides on top of it, grouped by what happens rather than listed one by one. A third step says where an ask is delivered: a run started from Slack is still answered in its own thread, and everything else can now land in a channel you choose instead of waiting unannounced in the Actions inbox. You can try a named call against the whole thing before saving it, arguments included.How long a question waits is now set in one place — on the agent, in that third step, anywhere from 5 minutes to 12 hours, defaulting to an hour. The per-rule wait and the account-wide setting are both gone: three places could answer “how long does this wait” and no single screen could, so a number you set in one could be quietly overruled by another.Choosing ask now says what it costs, wherever you choose it: the agent stops on the call and holds its only slot until somebody answers, and runs queued behind it start failing after about an hour.The Actions page no longer offers New permission — a rule is one line of an agent’s posture, so it is written on the agent, under Fleet → the agent → Permissions. Rules already there stay editable.Try it in KAOP

Improved: AgentOps is now KAOP

The product you use to build, run, and operate your agents is now called KAOP. You’ll see the new name across the app, in emails, and throughout the docs. Nothing about how it works changes — your agents, runs, integrations, and API keys are exactly as you left them, and existing links keep working.
NewImprovedFixed

Fixed: Workflows say when no step graph will run

The incident workflow engine is not enabled on every account. Where it is off, incidents are triaged by the orchestrator directly — which is intended — but the product said nothing about it. The wizard drew the full step graph, including a Select remediation card badged Awaits operator; the Remediation step said only that a different template could not be saved; the workflow’s Metadata panel described a remediation mode where “an operator selects before anything is applied”; and an incident’s Steps card showed an empty list that read as a graph yet to start.All four now say the same true thing: no workflow graph runs for this account, why, and the consequence — no step runs, and no operator approves anything before it is applied. The step cards are marked Will not run rather than drawn as a graph that executes, and the template you pick is described as recorded for when the engine is turned on.Try it in KAOP

Fixed: The workflow wizard’s Source step opens on the essentials

Step 2 of the incident workflow wizard no longer greets you with two expanded advanced sections. Field Mapping and Trigger Filtering now start collapsed, so the step opens on the endpoint picker. Each one still unfolds on its own when the workflow already carries something to show — a stored filter or a mapping override — and stays where you put it once you touch it.Try it in KAOP

Improved: The workflow wizard no longer assigns specialists that are not running

The specialist grid let you assign an agent that is not running, marked only by a line saying it would not run until it reconnected — so a workflow could be created with a specialist that silently never did anything. Those tiles now sort below the agents that are up and refuse the click, naming what has to happen first: reconnect, finish deploying, or become healthy. A workflow you are editing keeps the specialist it is already bound to, so you can still clear or change it.The review step’s last two bands — the one-line summary of what the workflow runs, and the flow of what happens to an alert — now carry titles like the cards above them.Try it in KAOP

Fixed: Role holder counts only count active holders

The Holders column on Settings → Roles counted every grant ever issued for a role, including ones held by disabled service accounts or members. On accounts with many disabled principals this could overstate a role’s real holders several times over — most visibly for admin.The count now reflects who can actually use the role today: it excludes disabled principals and any grant that has expired.Try it in KAOP

Fixed: Two right-sizing policies that would not save now save

Editing a right-sizing policy could fail in two ways that had nothing to do with what you were trying to change.Picking “Only when pod is created/restarted” and then switching back to “Apply Immediately” left a hidden restart setting behind, and saving the policy was rejected with ALLOW_RESTART_NOT_APPLICABLE — an option that only applies to the on-creation protocol being sent alongside the immediate one. The restart setting is now sent only when it applies.Separately, a policy with priority 0 — which the default policy has — reported “Priority is required” and blocked the step, even though 0 is a real priority and the API accepts it. Priority now accepts 0, and an empty field says it is required while an out-of-range one says so instead.Try it in KAOP

New: Create and find right-sizing policies from the Right-Sizing page

The Right-Sizing by Workload header now carries the two policy actions it was missing. “Add Policy” opens the full policy wizard in place — the same create flow Settings → Right-Sizing Policies uses, without leaving the table you were reading. “View Policies” takes you to that settings tab when you want the whole list.Add Policy requires the k8s_cost.policy.manage capability and is hidden without it. View Policies is always available, since the policy list itself is readable by anyone who can open the page.Try it in KAOP

Fixed: A resolved incident’s banner says what actually happened

The resolved status banner on an incident’s run drilldown always read “Closed — either manually or automatically on Datadog recovery.” — even for a person’s manual remediation, or a run that stopped at Investigate and applied nothing.The banner now reads the incident’s own resolution record: who or what closed it, and what was (or wasn’t) applied.Try it in KAOP

Fixed: The Remediation agent is no longer offered as an investigation specialist

The investigation specialists picker offered the catalog’s Remediation agent and counted it among the agents that can investigate — but the draft then refused to save, with a validation error naming an internal value you had never typed. The only way forward was to work out which of your picks was at fault and untick it.That agent is not an investigator: it reads a finished investigation and proposes a fix, and it is chosen on the remediation step instead. The picker now says so where it is listed, rather than offering a selection the save would reject. An agent of your own that happens to be named “Remediation” is still yours to assign as an investigator, and now saves.Try it in KAOP

Fixed: Pausing a workflow says so, and the status updates straight away

Pausing or activating a workflow from the Incidents overview table left the row showing the status it had before the click — the badge still read Active and the button still offered Pause — until the page was reloaded. The action had already succeeded; only the table was out of date, which made a successful click look like a no-op and invited a second one.The overview now refreshes with the rest of the surface when a workflow’s status changes, and every Activate/Pause control across the Incidents screens confirms the result with a short message.Try it in KAOP

Fixed: A newly connected cluster says it is collecting data, not that its agent is out of date

A cluster connected through Add cluster appeared under Unavailable clusters in the K8s Cost cluster picker with the tooltip “Agent update required” — on an agent installed minutes earlier, often the newest one in the account. Nothing was wrong with it, and there was no update to install.Cost data is written hourly, so a cluster in its first hours has no rows yet. That was being read as the same condition as an agent too old to report any, and the two shared one message. A live cluster still waiting for its first rows now says “Collecting data - Cluster will be available in less than 24 hours”, which is what is actually happening; it becomes selectable once it has a day of history. A cluster that has been connected for longer and still has no data, or whose agent has stopped reporting, still says the agent needs attention — there, the message is true.Try it in KAOP

New: Ask a specific agent by name in Slack

Slack messages used to reach whichever agent a routing rule pointed at, so sending one question to the right specialist meant editing standing configuration that changes who answers for everyone in the channel. You can now name the agent you want by typing its handle to the bot, and that agent answers instead. It applies to that one message only and changes no configuration, so the next message is routed the way it always was.Get the name wrong and the bot says so and stops, rather than quietly passing your question to a different agent. Naming an agent you are not allowed to run reads the same way, so a reply never reveals that an agent you cannot see exists.Answers can now also carry the answering agent’s own name and icon rather than the app’s, which makes a thread with several agents in it readable. A specialist the assistant hands work to posts its own message in your thread instead of arriving folded into the assistant’s summary. Workspaces connected before this shipped keep the single app identity until an admin reinstalls the Slack app and grants the new permission.Try it in KAOP

Fixed: Memory spaces say when nothing is reviewing them

A memory a run saves is a draft until the space’s memory worker reviews it, and a draft is invisible to search. A space with no memory worker therefore keeps collecting drafts that no agent can ever retrieve — memory looks switched on and returns nothing. Manage spaces said so only in a line of grey helper text, styled exactly like the two neutral hints beside it, and only while the picker sat on Find by role.That state now draws a warning that names the consequence: every memory written here stays an unreviewed draft, so agents keep saving and never retrieve anything. It appears whenever the space has no effective worker — including when someone picked an agent that no longer resolves — and it says what to do: pick an agent, or give one the role=memory-worker label. Nothing is lost in the meantime; drafts already waiting are reviewed as soon as a worker exists.Try it in KAOP

Fixed: Added by” now shows a name, not an internal id

An integration created with a service account’s API key showed its raw internal id (prin_...) in the Integrations table’s “Added by” column instead of a readable name — while one created by a person correctly showed their email. That column now resolves a service account to its display name, and falls back to a generic “Automated” label rather than an opaque id when the account can’t be resolved.Try it in KAOP

Fixed: The Insights Failures breakdown now adds up to its total

The Failures card on the Insights page shows a headline total with a per-category breakdown underneath — but the breakdown only ever showed the top 3 categories, so whenever a 4th category had failures, the numbers shown no longer summed to the headline.The breakdown still leads with the top categories by count, but now folds anything beyond that into a +N more remainder, so what you see always reconciles with the total. If the line is too long to fit the card, hover it to see the full breakdown.Try it in KAOP

Fixed: An incident workflow can no longer go live with nobody to investigate

Creating an incident workflow with no investigation specialist used to be allowed with only a warning. Every incident it opened was a no-op — the orchestrator had nobody to delegate to, so no evidence was gathered and no root cause was established, yet the incident still spent tokens and closed as Resolved.Assigning at least one specialist is now required: the Investigate & remediate step holds until you pick one, and activation is refused over the API as well.Try it in KAOP

Fixed: New incident workflow stops toasting “Connect Slack first.

Clicking Next on the Name step of a new incident workflow no longer raises a red “Request failed / Connect Slack first.” toast when your account has no Slack integration. Slack not being connected is an expected state here — the Endpoint step already says so inline and lets you paste a channel ID — so it is no longer reported as a failed request. The same applies to the Slack channel pickers in notification sinks.A Slack integration that genuinely fails still surfaces an error as before.Try it in KAOP

Improved: Know when a policy’s clusters cannot right-size immediately

“Apply Immediately” resizes running pods in place, which Kubernetes only supports from 1.33 with the in-place pod vertical scaling feature gate enabled and a recent enough Komodor agent. A policy scoped to clusters that do not meet that bar still saves, but its changes quietly wait for the next pod restart instead.The When to Apply step now says so. Pick “Apply Immediately” while any cluster in the policy’s scope cannot honour it, and the step flags that those workloads will wait for a pod restart instead — so what you are reading matches what the policy will actually do.The warning is advisory — it never blocks saving the policy.Try it in KAOP

Fixed: Deploy from catalog offers only what the picker will accept

On the incident workflow wizard’s Investigate & remediate step, “Deploy from catalog” opened the same unfiltered worker catalog whichever picker you clicked it from. From the orchestrator picker you could deploy a plain specialist — which that picker then refuses to list — so the dead end only showed up after the deploy had finished.The dialog now scopes itself to the picker that opened it: orchestrators from the orchestrator picker, everything an orchestrator can delegate to from the specialist pool. A banner names the scope and counts what it hid, and “Show all agents” drops it if you want the full catalog back.Try it in KAOP

Fixed: The config operator’s install commands work on a clean cluster

The Install dialog for a configuration cluster showed kubectl create secret before helm install --create-namespace, so copying the commands in order failed on a cluster that had never seen the agentops namespace: the Secret step ran first and had nowhere to land.The dialog now leads with kubectl create namespace agentops, and its title names what it installs — “Install the config operator into …” — rather than just the cluster.Try it in KAOP

Improved: You can un-pick an integration or group in the agent wizard

The Integrations step lets you scope an agent to a single integration or to a group of them. Once you had picked one, there was no way to un-pick it — every option in the list selects something, and none of them means “none”.Both pickers now carry a clear button. It empties the field back to its placeholder so you can choose again, and it takes the tool selection with it so no tick from the previous server is carried over. It does not unscope the agent: the kind of scope you chose stays chosen, waiting for a new pick.Try it in KAOP

Fixed: Back returns you to the workflow feed you came from

Opening an incident from one workflow’s feed and clicking Back could drop you on the general incident feed instead of the workflow you were working in — losing the scope, and the filters and search you had set on it.Back now returns to the exact feed the incident was opened from, and names it, so a shared link, a reload, or an incident opened in a new tab all land back where you started.Try it in KAOP

Fixed: Connecting AWS from the agent wizard offers both ways in

AWS can be connected two ways: a static key pair, or a cross-account IAM role you create by launching a stack in your own console, which stores no long-lived keys anywhere. The Integrations page has always asked which you want. The agent wizard never did — it rendered the key form and nothing else, so the keyless route was unreachable without leaving the wizard.“Connect new” on a provider that offers a choice now asks the same question the Integrations page asks, on that provider, and brings you back to the step you left with the new connection selected. Providers with a single way in are unchanged.Try it in KAOP

Fixed: Activate is disabled on a workflow that is not ready yet

An incident workflow draft used to offer an enabled Activate button even when it could not possibly activate — the wizard saves the draft as soon as you name it, so abandoning it before picking an investigation team left a workflow whose only feedback was an “Activate failed” toast after the click.Activate is now disabled on a workflow the control plane would refuse, and hovering it says which step is missing — an orchestrator agent, or an endpoint / Slack trigger to route alerts in. This applies everywhere the action appears: the workflows list, the workflow page, and the “Workflow not activated” empty state on its feed.Try it in KAOP

Fixed: Connecting an integration from the agent wizard asks for the right fields

The wizard’s inline “Connect new” form now renders each provider’s fields the way the provider describes them, matching the Integrations page. A field with a fixed set of values — Datadog’s site, an AWS region — is a picker rather than a box you had to spell the value into, defaults are filled in, and the guidance the catalog writes for a field appears beside it.Fields that only apply to the authentication method you picked are hidden, and no longer submitted. Saving is no longer blocked on one of those hidden fields, which could leave the button inert with nothing on screen left to fill in.Try it in KAOP
NewImprovedFixed

Improved: A webhook sink URL says why it was refused

A notification sink’s destination URL has always been checked for private, loopback and link-local addresses before it is saved — but the form only asked that it start with https://, so anything else came back as a save error after the fact.The field now applies the same rules the server does, as you type: a non-HTTPS scheme, an internal hostname such as api.default.svc.cluster.local, and a private or loopback address each say what is wrong in the field itself. The server still has the last word, because it also checks the address the hostname resolves to at delivery time.Try it in KAOP

Fixed: Switch between knowledge bases

The Knowledge page always showed the most recently created knowledge base and had no way to reach any other. An account that had more than one — for example a base declared through cluster IaC alongside the account’s own base — could permanently lose access to every older base and its documents the moment a newer one appeared.The page now defaults to the account’s original base rather than whichever is newest, and a selector appears next to Add page whenever there is more than one to choose from. An account with a single knowledge base sees no change.Try it in KAOPThe “N skills · N agents · N tags” badges on the Skills page always showed the full catalog, even while a search or kind filter had narrowed the list beside them — so searching for something that matched nothing still showed the total count next to “No skills match this search.”The badges now count the filtered results, so they always agree with the list.Try it in KAOP

Improved: Search for an agent in the incident workflow wizard

Choosing the orchestrator and the investigation specialists for an incident workflow meant reading the whole fleet — a specialist grid with every eligible agent in it, and an orchestrator dropdown with no filter. With no way to narrow either one, finding an agent by name came down to the browser’s own Ctrl+F.Both now filter as you type, and match on the agent id as well as the display name, so an agent you know only as datadog-alert-triage is one query away. A query that matches nothing says so rather than leaving you with an empty panel, and the “N assigned” count and the fleet-coverage line keep counting the fleet, not whatever the search left on screen.Try it in KAOPThe incident run page names the workflow that produced the run, but the name was plain text — to read the workflow’s overview you had to open Workflows and find it by name yourself.The name is now a link straight to that workflow’s overview, so a run drills up to the workflow that ran it in one click.Try it in KAOP

New: Register a private Azure AI Foundry as your own gateway

The “Managed by me” path — your own LLM gateway that your self-hosted agents dial directly, with no inference traffic through Komodor — now offers a gateway kind. Alongside the OpenAI-shaped kinds (LiteLLM, Hosted vLLM, Ollama), you can register an Azure AI Foundry resource: point it at https://<resource>.services.ai.azure.com and Komodor lists the deployments over Foundry’s own route.This is the private Foundry path, and it is deliberately distinct from “Bring your own Azure AI Foundry” under Add a new key. The Managed-by-me one stays inside your network: your workers reach the resource directly and Komodor is never in the inference path — only discovery, verify and spend travel the control-plane channel (through an Outpost if the resource is private). The Add-a-new-key one is public and routes inference through Komodor’s shared gateway. The kind picker names the difference so you register the right one.Try it in KAOP

Fixed: K8s Cost Overview’s Pod Placement cards no longer flash 0% while loading

The “Pod placement savings” card and the “Savings from Pod Placement” figure on the “Potential Savings” card briefly showed 0% while their data was still loading, before jumping to the real value. They now show a loading state until the real figure is ready.Try it in KAOP

Improved: Connections and tools are one step, and picking the tools picks the connection

Setting up a catalog agent asked for its integration twice — once as a connection, once as an MCP server — and the two could disagree with nothing on screen saying which one the agent would use. They are now a single Integrations step.For an agent built around one provider, picking its server or its integration group is the whole answer: the connection comes from the pick and is shown back to you, so there is no second field to keep in sync. The pickers list only what that agent can actually use, and anything else stays visible with the reason it cannot be chosen. Where an integration is optional, every scope stays available and the agent runs without it.Activate no longer lets a half-finished scope through. A scope whose kind you chose but whose server you did not used to deploy as no scope at all; it now says so and takes you back to the step.Try it in KAOP

Improved: Registering your own LLM gateway is easier to get right

Adding a “Managed by me” gateway now guides you past the two places it used to trip. When the gateway is reached through an Outpost, the wizard checks the host against that Outpost’s allowlist before you submit — so a host the relay would refuse is caught up front, with a note to add it and re-install the Outpost, instead of failing only once it tries to list models.Picking which models to expose is no longer a wall of pre-selected chips: the discovered models come as a searchable, checkable list with Select all and Deselect all, so paring hundreds down to the few your agents use is a couple of clicks. And once a gateway is created you can step Back to review it, or Edit to change the endpoint or key.Try it in KAOP

Fixed: GitHub can be connected from the agent wizard again

Creating a GitHub agent walks you through an Integrations step that asks you to connect GitHub. That step only ever offered one way to do it — type a name, press Save connection — and GitHub does not work that way: it is connected by installing its app, so every attempt came back with “GitHub uses an app installation; start the app installation flow instead” and no way to start that flow without abandoning the wizard.The step now offers the flow the provider actually uses. GitHub opens its app installation in a new tab; a provider that signs in with OAuth opens its authorization instead. Finish in that tab and the step picks the new connection up on its own and selects it, so the wizard carries on where you left it. Where a provider has not been configured on your deployment, the step says so up front rather than failing at the last press.Try it in KAOP

New: Bring your own Azure AI Foundry

Azure AI Foundry now sits alongside Azure OpenAI in the provider catalogue, so an account can register its own Foundry endpoint and point agents at the models deployed there.It is a separate provider from Azure OpenAI because Foundry is a different API, and the difference matters for Claude: a Claude deployment on Foundry speaks Anthropic’s Messages API but authenticates the Azure way. Picking Azure AI Foundry routes it correctly, while a non-Claude model registered under the same key takes Foundry’s OpenAI-compatible path — so one key covers a tenant’s whole Foundry catalogue.Add the key and the endpoint (https://<resource>.services.ai.azure.com) and the wizard lists the deployments that key can reach. An endpoint behind your own API gateway may publish no catalogue at all; the model id can always be typed by hand. The API version field is optional and only widens that listing.Try it in KAOP
ImprovedFixed

Improved: Every Slack message the bot accepts now gets an answer

A /agentops question is answered privately to whoever asked it, rather than posted where they cannot see it, and one typed in a direct message now reaches your DM rules. An incident raised by mentioning the bot reports back into the thread it came from, both when it concludes and while it is waiting on someone to pick a remediation. Handing work to an incident pipeline or a workflow leaves a reply and a ✅ or ❌ instead of a message stuck wearing 👀. A DM conversation keeps its context across follow-ups, and a job stranded with no worker is reported in minutes rather than after an hour of silence.

Fixed: Renaming a key you added a while ago works too

Renaming a key without re-entering its secret only worked for keys added recently. An older one asked for the key again and refused to save without it. Any key can now be renamed on its own.Try it in KAOP

Improved: Every key on Providers shows what it is costing

Each key row on Providers now carries the spend, calls and tokens of all the models underneath it, so you can read what a key is costing without opening it. A key whose models the gateway has not metered in the current window stays blank rather than claiming zero.Try it in KAOP

Improved: The alerts funnel now follows your workflow’s steps

“Where the alerts went” on a workflow’s Overview tab now reads how far each incident actually got: Investigated, RCA found or not found, then Remediation plan, Auto remediation or No remediation, with Still running and Failed as their own ends. Every stage shows its share and a P95 duration, and the Remediation plan stage shows how long it waited on a person.The left half is unchanged in substance: alerts ingested, folded as duplicates, filtered out, recovery signals, grouped into incidents. The chart changes unit once, at “Grouped into incidents”, and says so on the node — 7,280 alerts became 286 incidents, not 7,280 of anything else.Placement comes from the workflow’s own record — a step run for a human decision means a person was asked; an applied action in the run’s ledger means the agents changed something — never from how the incident was eventually closed. A manual close or a provider recovery is now a line of detail under the stage the run reached, not a stage of its own. The five tiles above the chart are all read from the funnel itself: Incidents created is now Incidents investigated and counts the distinct incidents behind the window’s grouped alerts, RCA success is Root cause found, and Remediation coverage is Remediated, a count split by plan and automatic. Clicking an incident-unit stage lists the incidents behind it.Try it in KAOP
NewImprovedFixed

Improved: Slack pings you when the answer lands, not before

Asking an agent in Slack no longer gets you an instant ”…” placeholder that is quietly edited minutes later — Slack never notifies on an edit, so the answer used to arrive in silence. The bot now posts a single message, the moment it has something to say, and reports progress on your own message instead: 👀 received, ⏳ running, ✅ answered, ❌ failed.

Improved: Your keys are the rows on Providers, and they can be renamed

Providers now groups the table by the key that serves each model, so one key backing four models is one row you can open rather than four rows repeating its name. Each key row carries the name you gave it, its identifier underneath, and how many models it serves; a search narrows the table to the keys that match and opens them. The actions that act on a key (add models, rotate, delete) now live on the key’s own row, and the ones that act on a model stay on the model.You can also rename a key without re-pasting it. Adding a provider no longer asks you to invent an identifier either: give the key a name if you want one, and we generate the identifier. Across the screen the same thing is now called a key everywhere, rather than a credential in one place and a key in the next.Try it in KAOP

Fixed: A model you just added becomes selectable in the wizard on its own

Registering a model leaves it unverified for a moment, and the agent wizard used to need a browser reload before it would offer it. It now keeps checking while something is waiting on verification, and refreshes the list when you come back to the tab, so a model that becomes ready while the wizard is open can be picked there.A model that is not selectable yet also says why in the dropdown, with the same explanation the Providers page gives, rather than appearing as a greyed-out name with no reason attached.Try it in KAOP

New: Connect a Kubernetes cluster to K8s Cost without leaving the page

Add cluster connects a new cluster in three steps: name it, run the Helm command it gives you, and watch the agent report in. The command carries the installation key already filled in, with a copy button, so nothing has to be pasted together by hand. It sits beside the cluster selector on Settings → Configured Clusters, and on every K8s Cost tab that has nothing to show yet. Asking the co-pilot to connect a cluster opens the same three steps.The name is checked as you type, against the same rules Komodor’s own installer enforces, so a name it would reject is caught before anything is created. A name that is already connected is flagged too — as a warning rather than a block, since reconnecting after a key rotation or a reinstall is a real thing to want.The last step is the useful one. Creating a cluster mints a key but tells you nothing about whether anyone actually ran the command, so the window keeps checking and confirms the moment your cluster first reports in — including while you are away in a terminal. If it is taking longer than usual it says so rather than spinning silently.The agent installs in cost mode: metrics and live cost data only. It does not read pod logs or exec into containers.Try it in KAOP
NewImprovedFixed

Fixed: Incidents created now agrees with the tiles beside it

A workflow’s Overview tab could state 0 under Incidents created while RCA success right beside it read “221 of 269 incidents”. The tile was counting alerts — the alerts funnel’s opened node, which counts an alert row and so counts a re-fire twice — while its two neighbours divided by incidents. It now reads the same incident count they do, so the three tiles share a denominator and the row can no longer contradict itself.Expect this number to change on workflows where alerts re-fire: it was previously inflated by every re-fire, and it is now the count of incidents actually created in the window.Run statistics in the same tab is now five tiles rather than five bare rows — Duration, Runs, Tokens, Cost and Errors — each with a line saying what its number means (for example, Cost is “Run cost in window”, not all-time spend). A non-zero Errors count stays red. The section also carries one name: it previously sat under a “Run statistics” heading with a card inside titled “Investigation summary”.Try it in KAOP

Fixed: Uploading a page again replaces it

Re-uploading a corrected file used to add a second page at the same path, leaving the old text in the knowledge base and retrievable — so an agent could cite the version you had just replaced. Now the page is replaced in place: it keeps its identity and history, and only the new text is retrievable. Before the upload runs, the dialog names any pages that are about to be replaced, so a batch of files can’t quietly overwrite something you couldn’t see.Try it in KAOP

Improved: Supply your “Managed by me” gateway key in the helm command instead

When you pick a “Managed by me” model in the add-agent wizard, the Model step now offers a checkbox — “I’ll provide the API key myself in the helm command”. Tick it and the Gateway API key field is disabled and no longer required, and the generated helm install command omits the key entirely, leaving you to add your own --set-string llm.anthropicApiKey.value=….It’s for teams who would rather not paste an inference key into the browser at all: the wizard never sees it, and it reaches your cluster only through the command you run yourself.Try it in KAOP

New: Scope which knowledge pages each agent can read

Any agent that could search the knowledge base could read all of it. Now a page carries tags, an agent carries a kb-tags label, and the two are matched at search time: an agent labelled kb-tags: payments retrieves the payments pages and not the security ones. Tag a page from its Details panel in Knowledge; set the label from the agent’s Metadata tab, where a picker offers the tags already in use and tells you how many pages the agent would reach.Two rules are worth knowing before you rely on it. An agent with no kb-tags label still reads everything, so nothing changes for the agents you have until you scope one deliberately. And a page with no tags stays readable by every agent — a page nobody has classified is not hidden, so the Details panel says so on the page itself rather than leaving an empty row. Tags are read on each search, so retagging a page takes effect on the agent’s next question with no re-indexing.Try it in KAOP

Improved: Scope an agent to an integration in one step, not two

The agent wizard’s MCP tools step used to ask twice. First a Connection question — no MCP servers, or connected integrations — and only then, nested underneath it, the scope that actually mattered. The first question added a click without adding a decision.It now opens straight on the scope: Single integration or Integration group. Choosing neither is how you say an agent needs no tools, and the step says so plainly rather than leaving you to infer it.Two things the dropdown now tells you before you commit. Each option is labelled with what it is — an Integration authenticated by a connection you set up, or a Custom integration, a raw MCP server registered by hand. And an integration whose tools the gateway has not discovered yet is listed but cannot be selected, because binding it would grant the agent nothing.Outpost routing moved into a collapsed Advanced section at the bottom. Its current value stays on the header, so a routed agent still says where its traffic goes without being opened — and an agent already bound to an Outpost opens with the section expanded. On a Komodor-cloud agent the control is now visible and disabled with the reason, instead of missing.Try it in KAOP

Improved: Schedules now read in plain English

Fleet → Schedules leads each row with what the schedule actually does — “Every 15 minutes”, “Mon–Fri at 07:00”, “Day 1 of every month at 09:00” — with the cron expression and its timezone underneath for anyone who reads cron directly. The schedule drawer says the same thing in its header.An expression that cannot be stated exactly still shows itself rather than a rough paraphrase: a day-of-month and a weekday together, for instance, are OR’d by cron, and a description that is only nearly right is one you have no way to check.Try it in KAOP

Fixed: A looped run’s steps now read in the order they ran

On a run whose workflow looped back — Verify did not pass, so the graph returned to Investigate — the Workflow steps card numbered its rows 1 2 3 4 5 1 2 3 6 and showed Investigate twice with nothing but a small badge to tell the two apart. The list was in execution order, but the number was the step’s position in the workflow, so a repeat read as a sorting glitch rather than as what it was: the run genuinely doing that step a second time.Rows are now numbered by the order they actually ran, 1..N, and each row states its place in the workflow beside its name — step 2/6 · attempt 2. Where the run went back, the card says so, on its own line: Looped back to Investigate.No run data changed, and nothing was collapsed — every step execution is still its own row. Only what the card claims about them is different.Try it in KAOP

New: Route a single-server agent through an Outpost, not just a group

A self-hosted agent scoped to one MCP server can now be routed through an Outpost in its own cluster, the same way a group-scoped agent already could. Pick one on the agent wizard’s MCP step and its tool calls go straight to that Outpost’s local listener instead of AgentOps.As with group scope, that trade is real: tool calls made this way never reach AgentOps, so no guardrail, policy, or audit applies to them, and the listener authenticates no caller — a NetworkPolicy in your own cluster is what decides who may use it.Try it in KAOP

New: Reach a private “Managed by me” gateway through an Outpost

A “Managed by me” gateway that lives inside your own network used to have to be publicly reachable for Komodor to list its models, verify it or read its spend. Now the Add-provider form and a gateway’s Edit dialog carry an Outpost picker: bind the gateway to one of your Outposts and Komodor reaches it for that metadata over the Outpost’s tunnel instead of a direct dial. Leave it on None for a gateway that is reachable directly.This changes only Komodor’s metadata path — your self-hosted workers still dial the gateway directly on their own network with their own key, exactly as before, so nothing about inference or where your credential lives changes. The picker is available anywhere the customer-gateway tier is, and the chat co-pilot can prep the field too.Try it in KAOP

Fixed: An incident no longer reports Agent offline for an agent that was never going to run it

An incident workflow names an orchestrator agent, and on workflows the platform runs step by step that agent is not what executes them — each step runs on the platform’s own step orchestrator. The incident was still checking the named agent’s heartbeat, so a workflow that was running perfectly well showed an Agent offline badge, and its timeline an entry saying the investigation was queued until that agent reconnected. Neither was true: nothing was waiting on it.Those workflows no longer read the named agent’s presence, and an incident that picked up the badge before its workflow moved onto the platform engine drops it on its next alert. Workflows that do dispatch to the agent you named — accounts not yet on the platform engine — are unchanged and still tell you when it is offline.Try it in KAOP

Improved: “New workflow” now sits in the same place on every Incidents screen

The Incidents Overview offered New workflow as a small action tucked into the right-hand end of the tab row, while the Incident feed offered it as the filled primary button beside the page title. Same action, two different buttons in two different places — so the one you had just used was never where you next looked for it.Every Incidents screen now puts it in one place: the primary button at the top right, beside the page title. The tab row carries tabs only.The Workflows list stopped offering it twice. That page rendered both copies at once — the title-row primary and the tab-row action — which read as two different buttons rather than one action. It now shows the single primary. Creating a workflow is still one click from ⌘K, and the workflow and run drilldowns keep their own actions instead of a create button they never needed.Try it in KAOP

Fixed: Runs keep the model they actually ran on, not the “Synthetic” placeholder

When an agent’s last turn failed, the CLI reported that turn under a placeholder name rather than a real model, and the run’s whole token bill was reattributed to it. <synthetic> then appeared in the Fleet Analytics model mix as if it were a model your fleet had chosen to run, taking one of the eight slots. Runs recorded from now on keep the model they actually spent their tokens on, and the failed turn’s tokens are still counted against it; a run whose only reported model was the placeholder now lands in (unattributed) instead, so growth there is the fix working. Runs recorded before this fix may still show the placeholder until they age out of the window.Try it in KAOP

Improved: Your knowledge base is already there

Knowledge used to open on a setup gate — “Knowledge base not set up”, and a Set up knowledge base button you had to press once before the page did anything. Every account now has its knowledge base from the start, named after the account, so the page opens straight into the wiki and there is nothing to set up.Nothing changes for an account that already had one: it keeps the knowledge base it has, under the name it was created with. The Set up knowledge base button is gone either way.Try it in KAOP

New: See every cluster reporting into K8s Cost, and how recently each one checked in

K8s Cost → Settings has a new Configured Clusters tab listing the Komodor agents reporting into your account — the agent version each one runs, when it was connected, and how long ago it last sent a heartbeat. Sort on any column, hide the ones you don’t need, narrow to specific clusters, and page through at 10, 25 or 50 rows.Agents that have stopped reporting are hidden by default, since a dormant cluster is rarely what you came to look at. Turn on Show inactive agents to bring them in; every row carries a status dot, so a cluster that has gone quiet is visible at a glance rather than something you infer from its heartbeat. A cluster that has never reported its version shows a dash, not a blank — it means the agent has not said, not that it is running nothing.Restarting an agent from this table is not available yet.Try it in KAOP

Improved: Click a configured cluster to see everything its agent reports

Rows in K8s Cost → Settings → Configured Clusters now open a details panel. Alongside what the table already showed, it carries the agent’s id, the Helm chart version it was installed from, and the Kubernetes version it reports running on — each with a copy button, so an agent id goes straight into a ticket or a kubectl command without retyping a uuid.The timestamps copy the full instant rather than the shortened form on screen, since “Aug 12, 2026” pasted into a query loses the time. Anything the agent has never reported shows a dash and offers no copy button, rather than looking like an empty value.Below the details, a JSON section shows the agent’s full reported configuration — the Helm values it is running with, useful when a cluster is behaving differently from its neighbours and you want to see what actually differs. Credential-shaped values are replaced with a marker before they leave the control plane.Try it in KAOP

Fixed: An incident’s counters now see a failed specialist

An incident whose specialist agent failed reported a clean run: Errors 0, Errors % at 0, a Runs count that left the specialist out, and a Specialist runs list naming only the step orchestrator. The runs an orchestrator delegates to were never linked to the incident at all, so nothing that counts them could see them — including any alerting built on those numbers.They are linked now, through the delegation lineage the platform records rather than anything the agent reports about itself. On an affected incident expect Runs to rise, Errors to show the failure in red, and the failed specialist to appear in the run list by name.Errors % on the incident feed and the incident page now reads 100% when any run of the investigation errored, not only when the incident’s own status is Errored — a hard specialist failure on an otherwise-progressing incident is no longer invisible there.One line was removed: the Runs tile on a workflow’s Overview tab no longer shows an “N orchestrators + M specialists” breakdown. That split subtracted the workflow’s configured specialist count from a run total covering every incident in the selected window, so the number it printed was not a count of anything. The Runs total itself, and the other four tiles, are unchanged.Try it in KAOP

New: Connect Grafana through an MCP server you already run in-cluster

The Grafana integration now asks which shape it is. Cloud, the default, is what it always was: a service-account token and your Grafana URL, for Grafana Cloud or a self-hosted Grafana. In-cluster MCP server records an MCP server you already run inside your cluster — its cluster-internal URL and, optionally, the name and key of a Kubernetes Secret that holds its bearer token. The token itself never leaves your cluster; only the reference is stored, and the control plane is never in the MCP data path.Because a cluster-internal address cannot be reached from the control plane, an in-cluster connection reports as unverified rather than healthy — the worker proves it at runtime. Existing Grafana connections are unchanged and stay in cloud mode.Try it in KAOP

Improved: Fleet’s Triggers tab is now Schedules, and you can filter and edit from the list

The tab is named for what it actually holds — recurring schedules. Inbound webhook endpoints never lived here; they are under Integrations → Endpoints. Alongside the rename it picks up the facet rail the Agents tab already had: filter by Enabled / Disabled and by any label on the agent a schedule fires, plus a search box that matches the agent, the cron expression and the schedule ID. The old ?tab=triggers links still work.Each row now shows the timezone under its cron expression, so a 9am schedule can no longer be misread as 9am UTC, and carries edit and delete buttons — with a confirmation step before anything is removed. A schedule declared in an agent’s agent-spec.yaml says so on the row and explains, on hover, why its controls are refusing instead of leaving them dead. The detail drawer drops the diagram that restated the cron expression and shows the input payload each run is invoked with in its place, keeps its edit and delete together, and marks run rows as the links they always were.Try it in KAOP

Fixed: Dropdowns show the detail beside each option again

Several dropdowns carry a second piece of information beside the option’s name — the servers an integration group contains, what a policy targets, what a saved selection resolves to. On any dropdown attached to a full-width field, that detail was being cut off: the list rendered narrower than the field above it and everything past its right edge disappeared.The most visible casualty was the agent wizard’s Integration group picker, which is where you decide what an agent may call. It listed group names alone, so answering “which MCP servers does this group actually grant?” meant selecting a group to find out, one at a time. It now reads ops-tools — kubernetes, datadog in the list itself, and a group with no servers says so before you pick it.A dropdown is now never narrower than the field it belongs to, so the same detail is back in the integration, policy, history and incident pickers that were losing it too.

Fixed: Disconnecting an integration no longer reports a failure that did not happen

Disconnecting an integration — and deleting a credential, an MCP server, a memory space, a notification sink or a golden suite — showed a success message and an error message at the same time. The page was re-reading the row it had just deleted, getting the expected “not found” back, and reporting it as a failure. The delete itself always worked; only the alarming second message was wrong, and it is gone.Try it in KAOP

Improved: Every cron field now reads back what it will actually do

The plain-English reading that Fleet → Schedules got has spread to every place a cron is written or shown. As you type into a schedule field — in the agent wizard, the Schedules dialog, the workflow builder or a maintenance window — the hint under it stops describing the format and starts describing your expression: “Every 15 minutes”, “Mon–Fri at 07:00”, “Day 1 of every month at 09:00”. So a mistake is visible before you save it, not after it fires.The confirmation you get after deploying an agent says the same thing, as do the workflow canvas’s trigger node and the maintenance wizard’s review step. An expression that cannot be stated exactly still shows itself instead of a rough paraphrase.

New: Agents deployed from the catalog can now be edited

An agent you deployed from the catalog used to be frozen at creation. The two things most likely to be wrong on the first attempt — which integration connection it uses, and which model it runs on — could only be changed by deleting the agent and starting again, which lost its history and its triggers.Those agents now open in the same edit wizard a self-built agent uses, for exactly the fields the deploy asked you for: the integration connection per provider, the model, the MCP group or server and its tools, the display name, labels, and how many copies run. A connection change takes effect on the agent’s next run; the rest arrives with its redeploy.Everything the catalog image defines — its instructions, skills and tool surface — stays fixed, and now says so where you would look for it, with the reason, instead of the field quietly disappearing. Komodor-cloud and self-hosted catalog agents are both editable, for the same set of fields.Try it in KAOP

Improved: An agent’s metadata now shows which LLM provider it uses

The Metadata tab on an agent’s page now has a Provider field beside Model, so the model name no longer hides which gateway actually serves it. It always shows one of Komodor managed gateway, Bring your own, or Managed by me — for every agent, including hosted and catalog ones, falling back to the managed gateway when an agent declared no provider of its own.Try it in KAOP
NewImprovedFixed

Improved: Workflow wizard’s ”+ Create new endpoint” now opens the real endpoint popover

The New incident workflow wizard’s Endpoint step used to create an endpoint through its own stripped-down inline form. ”+ Create new endpoint” now opens the same endpoint popover the Endpoints page and the chat co-pilot use — every source, auth, and mapping option is available, not just a name and a sender. On success the new endpoint is selected in the wizard automatically, and nothing else you had already entered is touched.Try it in KAOP

Improved: A workflow row opens from anywhere on the row, not just its name

On the Incidents Overview and Workflows tables, only the workflow’s name was clickable — the volume, resolution, error-rate and spend cells you were actually reading were dead space, so drilling into a workflow meant travelling back to the first column. The whole row is now the click target, and it shows a pointer on hover so you can see that before you try it.The row’s own controls still belong to themselves: the Activate/Pause button, the Actions kebab and the copy-the-agent-id affordance do their own job and nothing else — clicking one does not also open the workflow. The name is still a real link, so it keeps its keyboard focus, Enter to open, and open-in-new-tab.Try it in KAOP

New: A workflow now tells you when it was last updated

The Record group on a workflow’s Steps & config tab now shows Last updated beside Created and Last incident — so “did this workflow change before or after the alert I’m looking at?” is answerable without guessing.It reports the last write to the workflow for any reason, and says so: activating or pausing a workflow moves it too, not only a config edit. A workflow nothing has touched since it was created reads Never updated, which is deliberately different from a workflow whose timestamp we could not read.Try it in KAOP

Improved: Sort the incident workflow tables by any column

The Incidents Overview and Workflows tables arrived in one fixed order, so finding the workflow that spent the most, errored the most, or has not fired in weeks meant reading every row. Each column header is now a sort control: click it to sort, click again to reverse. Metrics, money and dates open on the end you are usually looking for — the biggest number, the most recent date — and text columns open A-Z. The active column and its direction are marked in the header.Sorting is a view, not a setting: the tables still open in the order the platform returns them, and a workflow with no data yet stays at the bottom in both directions rather than pretending to be the cheapest or the fastest.Try it in KAOP

Improved: Shadow A/B results now sit on the version they judged

A shadow experiment compares two versions of one agent, and each side is its own generation. But production keeps shipping while a candidate stays pinned, so the same candidate ends up judged against several production versions — and until now every one of those verdicts was folded into a single row. On one staging agent that blended three separate experiments into one, and the two baselines it was actually measured against disagreed on whether the candidate used more tokens or fewer.The Versions tab on an agent now marks a generation that ran as a shadow candidate, carries its verdict on the row, and expands into one result per production baseline — each with its own win distribution, score spread, and cost/token/latency deltas. Baselines are summary lines you open one at a time, so a candidate measured against ten production versions stays readable. A generation that served only as the baseline links across to the candidate instead. Quality Lab’s A/B Test Results table splits the same way.Verdict counts still fold across baselines — “Production wins 6/7” — but scores and costs never do, because averaging verdicts from different baselines states something no evaluator said.One number changes. Per-generation invocation counts were attributed by timestamp, which broke whenever two generations were live at once — exactly what a shadow is. They now follow the worker that actually served each run, so a candidate that served no production traffic correctly reads zero instead of borrowing its baseline’s runs. Expect existing counts on agents with shadow history to shift.Try it in KAOP

New: Shadow comparisons now show why an arm won, not just that it did

A shadow experiment could tell you Production or Shadow came out ahead and by how much, and nothing else — the evaluator’s reasoning was recorded on every verdict and never rendered.Clicking any comparison in an experiment’s sample list now opens its full report: both arms’ weighted totals with the winner, a one-line summary of what separated them, each rubric’s contribution to the total, and the evaluator’s point-by-point comparison behind every score. Below that, what each arm did better — and what both missed, the gap neither arm covered, which is the one thing no score can tell you.Rubrics are now weighted rather than counted, so an arm that wins the dimension that matters most is no longer outvoted by two narrow losses on lighter ones.Comparisons graded before this show their scores and say plainly which parts an older evaluator did not record.Try it in KAOP

Improved: A workflow run page now offers one way out, and it says where that goes

Opening a workflow run used to hand you the module’s whole tab bar and a New workflow button — so a page you opened to read one incident also offered to switch sections and start something new. The run page now carries a single back link at the top left and nothing else.The link names where you came from. Arrive from the Incident feed and it reads Incident feed; arrive from one workflow’s Workflow runs tab and it reads Workflow runs. Either way it returns you to that exact list with your filters still applied. A run opened from a bookmark or a pasted URL has no origin to name, so it points at the unfiltered Incident feed and says so.Try it in KAOP

Improved: The run page headline now spans the page, with findings beside the summary

The incident run page put every section inside one two-column grid, so the headline — title, severity, status, workflow and trigger time — was squeezed into two-thirds of the width with an empty rail beside it, and Findings was read through a narrow column.The headline card now spans the full page width, and directly under it Findings and the Investigation summary sit side by side in one row. Findings gets the wide column, so an agent-written root cause takes far fewer screens to read. The panel that asks you for a decision and the errored / needs-review banners that carry Retry stay full width above that row, and everything below it — the steps panel, deduped alerts, the collapsed alert payload, the Agent team card, Mark resolved and the timeline — keeps the order it already had.On a narrow window the row stacks, and a run that never started shows its findings full width with no empty summary card beside them.Try it in KAOP

Improved: A run waiting on your decision now keeps it docked at the bottom of the page

When a workflow run stops on something a person has to answer, the decision used to sit inline in the middle of the page — so on a long run it scrolled out of sight exactly when you went looking for the findings you needed in order to answer it. It is now docked to the bottom of the page instead: it stays put however far you scroll, and you can collapse it to a single line while you read. Everything above it stays fully readable and scrollable — the dock sits at the bottom of the content, it does not float over it. Both the workflow run page and the incident detail page use it, and nothing about the decision itself changed: the same frozen options, the same steps and undo instructions, the same “selecting nothing applies nothing”.Try it in KAOP

Improved: Field mapping now shows what reads each field, and tests a path before you save it

The four field-mapping boxes on an endpoint — and their per-workflow overrides — used to sit empty, which read as four blanks waiting to be filled when an empty box actually means “let the sender decide”. Each one now shows what does decide the field and is locked until you press Override, with the button that hands it back named for what takes over: Use default on the endpoint, Follow endpoint on a workflow.What a path may look like is now stated rather than implied: dot-separated keys that can nest, like details.incident.title, with your sender’s own example in each box. Paths that could never resolve — a pasted or list, brackets, quotes — are caught before the step will advance, instead of being saved and quietly doing nothing on every delivery.And the mapping can finally be tried before it is trusted. The endpoint’s Test & mapping step carries a payload you can edit, seeded from the last request it received, and shows exactly what your paths find in it — including when a dedup key resolved to nothing and the sender’s default answered instead. Nothing is sent and no incident is opened.Try it in KAOP

Improved: Incident Slack messages carry three times more findings

An incident’s Slack notification used to cut its findings at 2,800 characters, which trimmed roughly half of them mid-report. The budget is now 8,400, so about nine in ten incidents arrive whole — and the ones still long enough to be cut are trimmed on a whitespace boundary and marked, as before. Slack allows 40,000 characters in a message, so there is room left over.Try it in KAOP

Improved: Triaging and Triaged now read Investigating and Investigated

“Triage” was jargon that appeared in exactly one place — the status badge. Everything around it already said investigation: the help text, the detail banner, the timeline entry, even the agent doing the work. The badge now matches. Triaging reads Investigating, and Triaged reads Investigated, on the incident tile, the Status filter rail, the detail page and the notification event list.Investigated does not mean “waiting for you.” It covers every stage after the investigation finishes, so an incident reads Investigated while the workflow is proposing a remediation, applying one, verifying it, or writing the postmortem — as well as when it is genuinely sitting there. The status tooltip used to claim it was “awaiting operator sign-off to resolve”, which was simply untrue for an incident mid-remediation. It now says what the status actually covers.If you want the incidents that really are blocked on a person, that is the Status rail’s needs-a-person filter and the count on the Incident feed tab — not this status.Filters and API values are unchanged: status=triaging and status=triaged still work, so saved filters and scripts keep working.Try it in KAOP

Improved: The incident status formerly shown as Errored now reads Error

The status an incident lands in when its workflow stops on a failure used to render as Errored on the incident tile, the Status filter rail, the detail banner and the notification-event list. It now reads Error, matching the noun-shaped vocabulary the rest of the status set uses.Nothing about the state itself changed: the same incidents are in it, Retry still resumes the workflow from the step that did not finish, and the value the API accepts and returns (status=errored) is unchanged — so a saved filter, a webhook subscription or a script keyed on it keeps working.Try it in KAOP

Improved: The wait marker now says who owes the answer

An incident parked on a person used to carry a Waiting for response marker. That wording said neither who owed the response nor that the row was waiting on you rather than on the alert provider, and it sat next to Needs review — a different state, wanting a different action — with nothing to tell the two apart.The marker now reads Needs your answer, on the incident feed, every per-workflow feed and the incident’s own page. The three things a person can be asked for are now one name each: Error (press Retry), Needs review (the workflow finished with nowhere left to go — read the findings and close it out), and Needs your answer (the workflow asked you a question and is paused on it).The marker’s tooltip also stops claiming the question is always a remediation choice: since a guardrail hold parks a run the same way, it can equally be an approval the workflow is waiting on. Nothing about the underlying state changed, and the marker still layers on the status — a parked incident still reads Triaged while it waits.Try it in KAOP

Improved: The incident feed reads as a tile list, keeping every column it had

The Incident feed is no longer a table. Each incident is now a divided tile — provider mark, severity, title, and a muted line carrying the module, the workflow, when it started and its orchestrator run.Nothing was dropped. Duration, Errors and Cost keep fixed, right-aligned slots that still line up down the list, each labelled so a number reads on its own, and Status keeps an aligned position of its own beside them — together with the badges for an incident waiting on a person, one whose agent is offline, and one that needs attention. The per-workflow Workflow runs tab gets the same row.Try it in KAOP

New: Spot incidents that are stuck, right from the incident list

An incident that cannot move without a person — errored, out of retryable steps, or parked on a remediation decision — now carries an amber marker on the row itself, distinct from a running one, and a new Stuck filter in the Status rail narrows the list to exactly those rows in one click.Stuck is a different question from the feed tab’s own badge: the tab counts the whole backlog that will eventually need a person, including incidents still being investigated; this filter is only what is blocked right now. The two numbers are allowed to differ.The marker clears the moment you act (resolve, retry, or answer the parked decision) — no reload needed, on the module feed and every per-workflow feed.Try it in KAOP

New: The Incident feed tab shows how many incidents need a person

The Incident feed tab in the Incidents module now carries a count of the incidents waiting on a human — open, triaging, errored, needs-review, or parked on a decision — reading “3 incidents need you” on hover and to a screen reader. It is visible from every page of the module, so you can see the backlog without leaving the workflow you are reading.It is a needs-a-person count, not an unread count, and the difference is worth knowing: the number drops when any operator resolves an incident, not when you look at it. Each incident is counted once even when it is both errored and awaiting a decision. The badge is absent while the count is loading, if the read fails, or when nothing needs a person — so a number on the tab always means there is genuinely something waiting.Try it in KAOP

Fixed: The Incident feed count now says what it counts — and counts the right rows

The number on the Incident feed tab used to read “3 incidents need you”, and it was counting more than that claimed. Incidents still being investigated — Open and Triaging — were included, even though the workflow is running and nobody is waiting on a person. The badge was a backlog figure wearing a needs-a-person label.It now counts exactly what is blocked on someone: incidents in Error, in Needs review, or carrying a Needs your answer marker. It reads “3 incidents need your attention”, so the number says who is being asked and what for.Expect the number to get smaller. That is the point — the old one was too high. If you want the rows behind it, the Status rail’s filter (renamed from Stuck to Needs your attention) now selects exactly the same set, so the badge and the filter can no longer disagree about what needs a person.The badge is still absent while the count is loading, failed, or genuinely zero, so a rendered badge always means a real, positive number.Try it in KAOP

Fixed: Fleet status badges no longer spill into the next column

An agent carrying two status badges — Online plus Disabled, or Deploy Failed plus Disabled — laid them side by side and overflowed the Status column, rendering on top of Last run. Each badge takes its own line now, so the pair stays inside its column. Rows with a single badge are unchanged.Try it in KAOP

New: Edit a knowledge page in place

A wiki page can now be corrected without deleting and re-uploading it. Open a page, hit Edit, and change its title or Markdown body — the page keeps its identity, so the citations agents have already made and its retrieval count both survive the edit. Saving a body change re-indexes the page so agents retrieve the corrected text; renaming a page skips re-indexing entirely.Try it in KAOP

Fixed: Connection status now shows before the install steps for self-hosted agents

Activating a self-hosted agent used to bury the “waiting for the first heartbeat” checklist below the worker token and install commands, so you had to scroll to see whether the agent had actually connected. It now appears right away, above the copy-paste steps — matching where it already showed up when importing an existing agent.Try it in KAOP

Fixed: Coding agent connect commands now paste cleanly into zsh

The claude mcp add and codex mcp add commands on the Coding Agents page carry a ?account= query string that pins the connection to your current workspace. Pasted unquoted into zsh — the default shell on macOS — the ? was read as a filename wildcard and the command aborted with zsh: no matches found before it ever reached the CLI.The URLs in every generated command are now quoted, as is the <YOUR_API_KEY> placeholder in the Codex snippet, whose angle brackets were being read as a file redirection. The commands are unchanged for anyone who was already running them in bash, and the quoting is safe to paste into PowerShell and cmd as well.Try it in KAOP

Improved: New agents can chat by default

The Chat capability toggle on the Agent instructions step now starts on instead of off, so a newly created agent shows up in Chat right away. You can still switch it off before creating the agent if you want it to run without a chat surface.Try it in KAOP

Improved: Activating a self-hosted agent now reads as two clear steps

The Activate step used to keep its “generate a worker token” framing — and a live button — after the token already existed, so the connection status and the install commands read as parts of token generation, and pressing the button again quietly replaced a token your cluster might already be running. Creating the agent is now its own phase, and once it exists the step leads with the connection status, followed by the install steps.A separate “Rotate token” action covers the case the shown-once token gets lost: it issues a new one without re-creating the agent’s triggers or re-applying its roles and secrets.Try it in KAOP

Fixed: A bad New Relic key is now refused when you connect

Connecting New Relic accepted any API key. The connection went green, said it was connected, and the key was never exercised — the failure only appeared later, when discovering MCP tools returned “401 Unauthorized” and pointed at a connection that still looked healthy.Connecting now calls a read-only New Relic tool with the key you supplied and refuses the connection if New Relic rejects it, so a wrong or expired key fails on the screen where you can still fix it. A key that works is unaffected.Try it in KAOP
NewImprovedFixed

Improved: Workflows now list in one deliberate order across the Incidents module

The Workflows table and the Overview comparison table used to sort differently — one by creation date, the other alphabetically — so the same workflows appeared in two orders on two screens of one module. Both now share a single order: active and paused workflows first, most recently fired at the top, then the drafts by when you last edited them. A workflow that has never run sorts to the bottom of its group instead of landing wherever an empty timestamp happened to fall, and the order is applied on the server, so it holds across pages.Try it in KAOP

Fixed: Creating an endpoint inside the workflow wizard now shows its URL and token

An endpoint created with ”+ Create new endpoint” on the New incident workflow wizard’s Endpoint step now reveals the endpoint URL, the one-time token and a ready-to-run curl command, exactly as the Endpoints page does. The token was previously issued and never shown, leaving a live endpoint no sender could authenticate against; the reveal now stays put while you move between wizard steps and clears only when you dismiss it.Try it in KAOP

Improved: A workflow row now offers one primary action plus a 3-dot menu

Every Incidents module workflow row — on both the Workflows list and the Overview comparison table — now shows exactly one primary action chosen by status (Activate on a draft or paused workflow, Pause on an active one), with Edit, Duplicate and Delete behind a 3-dot menu. Delete is last and marked destructive, so it is no longer one mis-click from Duplicate, and it still asks for confirmation by name. An action the workflow’s current status forbids is now shown disabled with the reason instead of silently disappearing.Try it in KAOP

Improved: The workflow Overview gets the same six-option window picker

Each incident workflow’s Overview tab now offers the same six windows as the module Overview — 1h, 3h, 24h, 7d, 14d, and 30d, defaulting to 24h instead of 7d. Both the run-statistics panel and the alerts funnel honour the selection, including the three sub-day options. The workflow Overview keeps its own window control, independent of the module Overview’s — picking a window here never changes what the module Overview shows, and vice versa.Try it in KAOP

New: A workflow now tells you whether it is healthy, not only whether it is on

Each incident workflow’s Overview tab now carries a health verdict on the alerts-funnel card, beside — not instead of — the Draft / Active / Paused badge. That badge says what someone configured; the verdict says what the runs since then actually did.It reads Pipeline healthy, Pipeline degraded, or Pipeline unhealthy from the share of the window’s runs that errored — under 5%, 5% to 20%, and 20% or more — and states the arithmetic in words beside the pill: “3 of 24 runs errored (12%) in the last 7 days.” The denominator is named on purpose: this is a rate over the workflow’s own runs, which is a different question from the module Overview’s incident-level ”% of errors”.Two cases are withheld rather than guessed. Under ten runs in the window it reads Not enough data — with that few, a single failure would decide the verdict on its own. A workflow that has never been activated reads Not activated, since it has dispatched nothing to judge. The verdict follows whichever window the picker holds, and it is a description of what happened, not an SLO.Try it in KAOP

Improved: Combine several conditions in one Slack routing rule

A routing rule can now match on a list of conditions rather than one idea. Pick all of or any of, then combine a mention, a channel message, a direct message, a keyword, or a named sender, and scope the whole thing to particular Slack channels. The routing list reads the rule back in words, so a rule says “mention or direct message” rather than showing a type name.The bot is also quieter. Once an agent has answered in a channel thread, further messages there are treated as conversation between people, so a team can talk through an answer without the bot replying to every line. Mention it again to bring it back, and direct messages are unaffected.Three things got safer alongside it. A channel has exactly one fallback rule, it is shown last because it is only ever tried after everything else, and switching which rule holds it now takes one step instead of two, so a failure halfway can no longer leave a channel with no fallback at all. A new rule starts out needing a condition, so saving one without thinking no longer creates a rule that quietly matches every message. And a freshly connected workspace gets one starter rule instead of two, which still answers a mention or a DM and nothing else.Try it in KAOP

Fixed: A run’s token count now adds up

The Tokens tile on a run’s page showed a total alongside an “in” figure that did not account for it — on a prompt-caching agent the caption described under 1% of the number above it, because it counted only fresh input and left out the tokens served from and written to the prompt cache.The caption now sums every input the run consumed, so it reconciles with the total beside it. Cost was always priced from the full picture and is unchanged.Try it in KAOP

Improved: An incident run now opens on the findings, not on raw JSON

The incident run page used to make you scroll past the step accordion and the full alert payload before you reached what the investigation actually concluded. Findings now leads the page, directly under the header and the banners you may need to act on. The steps panel follows it, and the Alert payload sits last, collapsed until you ask for it.The right rail was reordered to match: the investigation summary, then an Agent team card that names the orchestrator and each specialist by role — replacing the unlabelled list of run links — then Mark resolved, with the timeline closing the page.Try it in KAOP

New: See why a workload can’t be right-sized

The Right-Sizing table now shows an info icon next to a workload’s Automated toggle when it can’t be right-sized yet — hover it for the reason (managed by an HPA, too new, not seen recently, or a QoS guardrail) and, where applicable, a shortcut to update the policy.Try it in KAOP

New: A Cost & Optimization drawer for each Right-Sizing workload

Clicking a row in the Right-Sizing table now opens a drawer with that workload’s cost summary, a 30-day cost trend, and a per-container breakdown — searchable, with a 7/30-day time window and expandable rows showing CPU/Memory usage graphs. A “Right Sizing automation” card lets you add the workload to its resolved policy or jump to Settings to review policies.Try it in KAOP

Fixed: Quality Lab names a deleted agent instead of its id

A shadow-comparison verdict deliberately outlives the agent it graded, but the Agent column and an evaluator’s own row used to fall back to the opaque agt_… id once that agent was gone. Both now read Deleted agent, with the id still one hover away for anyone who needs it.Try it in KAOP

New: See what your agents’ memories are about, and filter by it

The Memory page has a Labels tab: the dimensions you declared for a space (deployment, cluster, env, team) and the keys agents invented while writing, each with a line saying what it holds and its values with the number of memories carrying them. Click a value to see exactly those memories, or a label on any memory to filter the list by it. A search now ranks a memory about your environment above a near-identical one about somewhere else.Agents fill these in as they work, and they write that line too. An agent saving a memory says what it is about and can say what a new key holds, so the next writer reuses the key you already have instead of inventing a synonym beside it. A key an agent invented can be promoted to a dimension of its own, keeping its key and its values.Stop tracking a dimension and nothing is lost. The values stay, the memories keep them, and they still count towards ranking. The key stops being one of the dimensions agents are asked to fill in, and stays on the tab marked Not tracked, with the date it stopped.Try it in KAOP

Improved: Incidents module workflows table gets aggregate columns; the feed is now a table

The Incidents module Workflows page now shows Status, Agents, Incidents (open), Duration (P95), Errors %, Cost, Created, and Last Seen for every workflow, scoped by a Window picker and a Workflow selector — with row actions moved into a 3-dot menu. The Feed is now a table with Title, Status, Started, Duration, Errors %, and Cost, replacing the card list.Try it in KAOP

Improved: A quieter filter sidebar on the Incidents feed

The Incidents feed’s filter sidebar was a three-group rail with icons on some rows and not others, which made it read as unfinished and buried the counts you actually scan. It is now a narrow two-group rail — Status and Priority — with no icons and the count sitting plainly beside each label. Filtering by workflow moved out of the rail to a workflow picker next to the search box, where it belongs: unlike the two filters above it, it narrows the incidents we fetch rather than the ones already on screen. A workflow’s own Feed tab now shows exactly the same sidebar as the module-wide feed.Try it in KAOP

Fixed: Filter the Incidents feed for Errored and Needs review

The Incidents feed’s Status filter listed only four of the six statuses an incident can hold, so an errored incident — or one that ran to completion with nothing left to retry — could only be found by scrolling the unfiltered list. Errored and Needs review now appear on the rail with their counts, on both the module feed and a workflow’s own Feed tab, in lifecycle order alongside the statuses that were already there.Try it in KAOP

Improved: An incident waiting on your decision now says so in the feed

An incident parked on a human decision used to be indistinguishable from one still being worked: it read Triaged like any other, and nothing in the feed said a person was the thing it was waiting for. Incident rows now carry a Waiting for response marker — on every incident feed and detail page — and the Status rail gained a facet that filters straight to them, so the incidents that need you are one click away instead of something you find by opening rows.The marker sits alongside the status rather than replacing it, because both facts are true at once: a parked incident really is triaged and waiting. It is a wait marker layered on the status, not a status of its own — which is why such an incident is still counted under Triaged too. The incident detail page says the same thing in words, under the status it already explained.One scope note: the marker tracks a pending remediation decision. A run held on a guardrail approval is a different kind of wait, and this marker does not claim to cover it.Try it in KAOP

Fixed: Incident findings always read as a report

An investigation whose alert asked the agent to answer in JSON used to land on the incident’s Findings card as the raw payload — a wall of {"status": …, "answer": "…\n…"} with every line break escaped, even though the root cause was written inside it. The answer is now unwrapped back into markdown, so headings, tables and evidence render the way they do for every other run. The card also stops printing “No root cause identified yet” above a narrative that plainly states one — that line now appears only when there really is nothing else to show.Try it in KAOP

Fixed: The incident feed row tells you which workflow triggered it, and which run it dispatched

Every row in the Incidents feed now carries a meta line under its title naming the module, the workflow that triggered the incident, and the orchestrator run it dispatched — the run id links straight to that run. An incident whose alert arrived with no severity is now tagged Unknown instead of showing a blank tag.Try it in KAOP

Improved: Fleet shows why a hosted agent never came up

A hosted agent whose pod is stuck — CrashLoopBackOff, ImagePullBackOff and friends — no longer sits on Deploying with nothing to say. Its status chip now reads the Kubernetes reason itself, and the agent’s panel shows the kubelet message together with the last lines of the crashed container’s log, so a bad image, a missing credential or a broken entrypoint is visible without kubectl. The status clears on its own once the pod is healthy again; editing, archiving and rolling back stay available throughout.Try it in KAOP

New: Findings now closes with what was applied to the incident — or why nothing was

The Findings card on an incident run ended at Next steps, which is advice — it never said what the workflow actually did. It now closes with a Resolution section stating the fact: the action that was applied and the step that applied it, or, far more often, why nothing was. A run that errored names the step it stopped at, a workflow that only proposes a remediation says an operator has not selected one yet, and a run still in flight says it is not resolved yet. The sentence is derived from the run’s terminal step and the workflow’s remediation mode rather than written by an agent, so it cannot report a fix that never happened.Try it in KAOP

Improved: Findings on a run leads with the root cause instead of everything at once

The Findings card on an incident run used to render every section expanded, so an agent-authored investigation could run several screens tall and push the rest of the run page out of view. It now opens on the root cause alone, height-capped to a few lines so a long one cannot defeat it, with Show all findings revealing the key evidence and next steps underneath. Everything behind the toggle is narrative — nothing you need to act on is hidden.Try it in KAOP

Improved: A remediation decision now shows the procedure, the verification and how to undo it

A parked remediation decision used to show you an option’s label and a single shell command. It now shows what the workflow actually froze behind the question: the finding that justifies asking at all, and per option the numbered procedure, the one check that says it worked, its cost or capacity impact, and the steps that reverse it. Where there is no undo, the option says which kind of no it is: an action recorded as irreversible reads as the fact somebody established, and an action nobody described reversing reads as missing information you should go and check. Neither is ever left as a blank section — before approving a production change, an absence you can see beats one you have to infer, and a gap nobody filled in should never look like a judgement somebody made. There is also an “Other” answer now: describe the action you would rather have reviewed, and nothing is applied. Everything shown is the copy taken when the question was asked, so a decision you answer a week later is the one you were originally offered.Try it in KAOP

Fixed: Run token and cost figures no longer double-count

A run’s recorded tokens and cost are now accurate in two cases that previously inflated them.A run whose worker was killed before it reported a result is requeued and executed again. Every span in that replay is new, so the replayed turn’s usage used to fold into the run’s total a second time — a re-executed run could report several times the tokens and cost it actually spent. Usage is now counted once per provider message, so a replay adds nothing it already counted.Separately, the per-turn usage an agent reports is now cumulative on every field. Mixed semantics previously let one run record a full conversation’s worth of cache reads beside a single turn’s input and output tokens, which made its estimated cost far larger than what it spent.Both apply to runs executed by agents on SDK 0.3.17 or later; figures already recorded are left as they were.Try it in KAOP
NewImprovedFixed

Improved: The workflow wizard gives remediation its own step

The New incident workflow wizard’s steps are now Name → Endpoint → Investigation team → Remediation → Review & activate. Picking the workflow template, choosing who remediates, and subscribing a notification sink used to be bundled into “Team & steps” alongside the orchestrator and investigate/verify specialists; they now live on their own Remediation step, so investigation staffing and remediation/delivery decisions no longer compete for the same screen.Try it in KAOP

Fixed: The workflow template picker only offers what it can actually save

The New incident workflow wizard offered a choice of workflow templates to every account, then rejected two of the three at save in environments where only one of them can run. The template a workflow runs is recorded on its workflow definition, and in an environment where nothing creates one, every workflow ends up on the operator-approved template no matter what was picked.The picker now asks the account whether a choice can be recorded at all. Where it cannot, it shows the one template the workflow will actually run and says why — either the workflow engine is not enabled for the account, or no step-orchestrator agent is deployed for it — instead of rendering buttons that fail on save. Where a choice can be recorded, nothing changes: the full list stays selectable.An existing workflow keeps showing its own template, including one this account could not create today.Try it in KAOP

Improved: The Steps tab now shows what Ingest and Group produce

The Steps tab’s “Before the workflow” group (Alert trigger, Dedup, Filtering) previously showed only a summary of each stage’s configuration. Alert trigger and Dedup now also list the fields they hand to the rest of the pipeline — alerts, started_at, title, service/environment for Alert trigger, and incident, merged_alerts, description, severity for Dedup — in the same Outputs chip list a real workflow step shows. Filtering declares none, since it only decides whether an alert is routed at all.Try it in KAOP

New: The workflow Overview now shows five headline stats

Each incident workflow’s Overview tab now shows five headline numbers above the alerts funnel: alerts ingested, noise suppressed (with the count behind the percentage), incidents created, RCA success, and remediation coverage. RCA success and remediation coverage are new: they read the run-statistics window already loaded for the tab, and each tile explains what it counts.Try it in KAOP

Improved: A workflow’s Incidents feed now has search and a Priority filter

The per-workflow Incidents feed tab now matches the module Incidents feed: a search box that matches incident title and alert summary server-side, and the Severity facet relabeled Priority (same field, same highest-first order). Every result stays scoped to the one workflow.Try it in KAOP

Improved: Edit a workflow step’s agents from the Steps tab

Every orchestrated step on a workflow’s Steps tab now has an “Edit specialists” link, in both List and Graph view. It deep-links straight to the wizard’s Team & steps page with the right specialist pool already open, instead of leaving you to find the wizard and the matching step yourself. Investigate and Verify — and, on a human-selected workflow, Propose and Apply remediation — share one binding, so the tab now says “Also feeds: …” next to the link whenever changing one step’s agents changes another step too.Try it in KAOP

Fixed: Review no longer reports a remediation agent a workflow will never call

The Orchestrated loop template has no remediation step — the orchestrator remediates inside Investigate and asks an operator from there. But if you picked a remediation agent first and then switched to that template, the pick stayed behind: the wizard’s last screen listed it as the workflow’s remediation, and the workflow was saved carrying a binding its graph could never dispatch.Switching to a template with no remediation step now drops the pick, and Review & activate reports remediation as handled by the orchestrator inside Investigate rather than naming an agent. Switching between two templates that both have a remediation step keeps your selection, as before.Try it in KAOP

Fixed: New Relic connections can pick their region

New Relic serves its MCP endpoint from three regional hosts, and a User API key issued in one region is rejected by the other two. The connection always dialed the US host, so an EU or JP account could paste a valid key, watch the connection go green, and then get a 401 from every tool call — a failure that reads as a bad key rather than as the wrong endpoint.Connecting New Relic now offers a New Relic region field alongside the API key. Existing connections are untouched and stay on the US endpoint, which is where they already were.Try it in KAOP

Improved: Loki setup asks which auth you use

Connecting Loki now starts by asking how your Loki is fronted — basic auth, a bearer token, or nothing — and shows only the fields that shape needs, instead of offering all three and rejecting the combinations that do not go together. Basic auth is the default, since that is how Grafana Cloud Logs works.Loki has no MCP server we host, so the next step has always been registering your own loki-mcp. That form now opens on the SSE transport when you reach it from a Loki connection — loki-mcp speaks the legacy SSE transport, and picking HTTP fails in a way that looks like an unreachable server rather than a wrong setting. The credential picker also says what binding actually injects, including the tenant header, and notes that loki-mcp queries the tenant in its own LOKI_ORG_ID, so a mismatch there is something you can see rather than discover through empty results.Try it in KAOP

Improved: Incidents Overview adds 1h/3h/24h windows, and defaults to 24h

The Incidents module Overview’s window picker now offers 1h, 3h, and 24h alongside the existing 7d/14d/30d, and defaults to 24h instead of 7d. The volume, resolution-time, automation-rate, and spend trends bucket hourly for a sub-day window and daily otherwise.Try it in KAOP

Improved: The Incidents feed’s Severity filter is now Priority

The Incidents module feed filters incidents by the same field as before — no new data, no schema change — but the facet in the sidebar is now labelled Priority instead of Severity, sorted highest-first. The underlying field is still severity on the URL and the API.Try it in KAOP

New: Search the Incidents feed by title or alert summary

The Incidents module feed now has a search box. It matches an incident’s title or its alert summary, filters server-side (so results stay correct once the list pages), is debounced as you type, and is reflected in the URL so a search is shareable and survives a refresh. An empty result now says which term didn’t match, instead of the generic “no incidents in this view” message.Try it in KAOP

Fixed: Creating an incident workflow no longer fails partway through

Creating a new incident workflow could fail on a later step with two error toasts, both saying the remediation mode (now the workflow template) “is chosen when the workflow is created and cannot be changed.” The wizard creates the pipeline as soon as you leave the Name step, before the template picker is even shown — every later autosave was then re-sending the template on every forward navigation, which the backend correctly refused once the pipeline had nothing to record it against.The wizard now only sends the template when you actually change it, so walking through the rest of the wizard no longer restates a value you never touched. A single failed save also now raises one error toast instead of two.Try it in KAOP

Improved: Every incident now says which module it came from

Incident rows in the module feed and the incident drilldown header now name the module that owns the incident’s workflow, alongside the workflow and the orchestrator run — the full Module → Workflow → Run trail. The module is read from the workflow’s own definition, so it stays correct without a copy stored on the incident, and it is what tells two modules’ incidents apart once more than one module ships workflows.Try it in KAOP

Improved: Incident KPI tiles now show the prior period and a delta

The Incidents module Overview showed a scalar and a sparkline per KPI, with no way to tell how a number moved without eyeballing the trend line.Every tile — Volume, P95 resolution time, % automated, and Spend — now also shows the prior period’s value and a signed delta, coloured by whether the move was good or bad for that KPI (or left neutral where a rising or falling count isn’t itself good or bad, like Volume).Try it in KAOP

Fixed: An imported agent can notify a sink when its runs finish

Importing an agent you already run offered no way to be told when its runs finish. The Triggers step’s second section — When the run finishes — was hidden for an import, on the reasoning that an import deploys nothing and so has no runs to hear about.That reasoning was wrong. Deploying is not what produces a run. AgentOps dispatches an imported agent’s runs, records them, and emits their lifecycle events exactly as it does for an agent it deployed itself — which is the same basis on which this wizard has always let you give an imported agent a schedule. A trigger fires, a run finishes, and a subscribed sink hears about it. Offering one half of that pair and hiding the other was the bug.The picker is now on the Triggers step for every kind of agent, and the sink you choose is applied when you import, alongside the roles, secrets and triggers.Two smaller things on the same screen. The Import agent button now states everything blocking it in one message rather than stacking one per blocker, and repeats it when you hover the disabled button. And a Not connecting? section lists what to check when your agent has not appeared yet — including the mismatch between AGENTOPS_AGENT_ID and the id the token was minted for, which is refused as an authorization error and so reads like a bad token when it is not.It is deliberately not a timed warning. Wiring up an agent in your own repository and pipeline can take days, and a page that called that broken after five minutes would be wrong nearly every time it said so.Try it in KAOP

Improved: Google Cloud now covers fourteen services, with mutating tools hidden by default

Connecting Google Cloud now offers fourteen of Google’s hosted MCP servers rather than three: Cloud Logging, Monitoring, Error Reporting, Trace, Service Health, Quotas, Compute Engine, GKE, Cloud Run, Billing, Recommender, Asset Inventory, Network Management, and Policy Troubleshooter. One service-account key covers all of them, and you pick which to register when you connect. The setup commands now grant roles/mcp.toolUser alongside roles/viewer — without it Google refuses every tool call while the connection still tests green, which previously looked like a broken integration rather than a missing role. Two services need one extra grant each, and the wizard says so: GKE wants roles/container.clusterViewer, and Billing wants roles/billing.viewer on the billing account. Servers registered through the wizard also hide the mutating tools these endpoints expose — creating, deleting or patching instances, clusters and services — so agents see the read-only surface and nothing that could change your project. You can widen that on the server itself if you want it.Try it in KAOP

Fixed: Gateway-routed run costs were priced off the wrong model

A run’s worker reports its own client-computed cost alongside its token counts, and the control plane trusted that figure once it had a nonzero cost, even when it came from a stale local price table with no entry for the model actually used — most visibly on Sonnet 5, priced at Sonnet 4.6’s rate and shown roughly 1.5x too high.Run cost is now priced from the account’s own gateway rates whenever the reported figure’s provenance is the client’s own guess rather than a real billing identity. A genuinely provider-billed cost is still never second-guessed. Existing runs correct themselves the next time they are viewed — no migration needed.Try it in KAOP

Improved: Incident Findings now show root cause, key evidence, and next steps

The drilldown’s Findings card used to render whatever markdown the investigating agent wrote, in whatever structure it chose — sometimes a root cause up front, sometimes buried, sometimes not named at all.It now reads the investigation’s own structured output and shows named sections: Root cause, Key evidence, and Next steps (all bulleted lists). Older incidents whose investigation recorded only a free-form narrative still show that narrative under What happened. An incident whose investigation produced no root cause says so explicitly instead of showing an empty section.Try it in KAOP

New: See which alerts merged into an incident, not just how many

The incident drilldown showed a “fired N times” count and nothing else, so there was no way to see which fires actually merged into an incident that refired.A refired incident now shows a Deduped Alerts card listing each merged fire with its timestamp. An incident that never refired shows no card.Try it in KAOP

Improved: Create a webhook endpoint without leaving the workflow wizard

The New incident workflow wizard’s Endpoint step now creates an endpoint in place: name it, pick which sender it is preset for, press Create, and the new endpoint is selected for the workflow. Previously the only way out of “no endpoints yet” was a trip to the Endpoints page, which meant abandoning a half-filled wizard and starting over. Everything already entered — the workflow name, the field mapping, the filters — survives the create, and a create that fails says so inline and leaves the draft untouched.Try it in KAOP

Fixed: Autoscaling a self-hosted agent now works, and keeps its tool access

Turning on KEDA autoscaling for a self-hosted agent could stop it calling any of its tools, reporting 401 Unauthorized from the control plane and ending every run as “orchestration incomplete”. Autoscaling gave every replica of the agent the same worker id, and each replica’s start-up replaced the one tool credential that id had — so whichever pod started first kept running with a credential that no longer worked, and only a restart cleared it.Each start-up now takes its own credential instead of replacing its siblings’, so replicas no longer cut each other off. The same shared id also broke the queue-depth reading KEDA scales on, which answered 403 and left the agent stuck at its minimum replica count — that now resolves the agent from the worker’s own credential, so scaling actually happens.Upgrade the chart to pick both up. Nothing to change in your values, and credentials already issued keep working.Try it in KAOP

New: Attach an agent to one MCP server, and pick its tools

The MCP tools step in the agent wizard now offers one server with the tools you choose, alongside the integration group it has always offered. Pick Connected integrations, then One server: choose a server, tick the tools this agent may call, and that is the agent’s whole tool surface.Scoping an agent to a single server used to mean creating an integration group with one member — shared, named registry state you then maintained for as long as the agent existed, with its tool list fixed for every agent bound to it. The per-server scope is the agent’s own: no group record, and the tools are chosen per agent. Ticking every tool records the tools that exist today, so a tool added to that server later is not granted automatically.The group option is unchanged and still the right choice when several agents should share one governed tool set — its panel now says why its tool list is read-only. The first option has been renamed No MCP servers, which is what it does on a fresh deploy.Available on both Komodor-cloud and self-hosted agents. Two fewer ways to over-grant: the wizard will not let you save a server with no tools picked, and a tool the server itself withholds cannot be ticked — the step names the layer that withheld it.Try it in KAOP

Improved: Pick the time window on an agent’s Observability tab

The Observability tab on an agent’s detail page was locked to a fixed 30-day view. You can now switch it to 7, 30, or 90 days, and the comparison follows your pick instead of always comparing against a fixed 30-day average — so a change from earlier this week shows up right away instead of getting averaged out.Try it in KAOP
NewImprovedFixed

New: See shadow A/B verdicts in Quality Lab, not just in Postgres

Shadow deploy (dual-deploy A/B) runs a candidate worker beside production under the same agent, with an evaluator grading the two arms head-to-head. Until now the only way to see a verdict was to query Postgres by hand.The Agent Score Results tab now has a Shadow Comparisons section listing every agent that’s been A/B’d, alongside the existing run-level scorecard. Open an agent to see comparison volume over time and an A/B Test Results table — production vs. shadow, per rubric, with the individual graded pairs and their un-blinded scores one click away. A low sample size (under 5 comparisons) is flagged rather than shown with false confidence.Try it in KAOP

Fixed: A run’s Output tab now renders like its Transcript

The Output tab on a run — the final answer, separate from the turn-by-turn Transcript — used a thinner markdown renderer than the Transcript did. Tables and inline code (like this) came out flat and hard to scan on Output even when the exact same text looked properly typeset one tab over, because both were rendering identical source text through two different, drifted implementations.Output now renders through the same Streamdown-backed component the Transcript already uses — syntax-highlighted code blocks, styled tables, the same typography — so a run’s answer looks the same wherever you read it.Try it in KAOP

Fixed: Run activity chart no longer flattens the days before your agent existed

An agent’s Run activity chart used to draw a flat line of zeros for every day before the agent was created, making a brand-new agent look like a dormant one that just woke up. Those days now show as a gap instead, and the week-over-week change on each tile compares only the days the agent could actually have run.Try it in KAOP

New: Preview a right-sizing policy right from the Right-Sizing table

The Right-Sizing table’s Policy cell now has a “Preview” link on every policy option — hover an option in the dropdown to see it, in place of its priority number. Preview opens a richer view of that policy: the workload it would apply to (cluster, namespace, workload, kind), the policy’s full configuration including its right-sizing guardrails, an “Edit policy” shortcut straight into the policy editor, and a “Manually Apply Policy” action to apply it to this workload on the spot.Edit policy requires the k8s_cost.policy.manage capability; without it, the button stays disabled with a tooltip explaining why.Try it in KAOP

Improved: The incident workflow wizard now reads its templates from the account

The New incident workflow wizard’s “Team & steps” step no longer hardcodes which workflow templates exist or which one is offered by default. It now calls the account’s own available-templates listing and renders the picker, the step graph, and each step’s name and description straight from that response — so an account that only has the operator-approved template sees exactly one option, and Komodor’s own account sees all three, with the create default always preselected correctly rather than assumed by position.An existing workflow whose template the account can no longer create — one made before an entitlement changed — still shows its own template rather than an empty picker.The remediation_mode field this wizard sent is renamed to template_key across the create/update API and the incident pipeline detail response, since it names a workflow template, not a runtime mode.Try it in KAOP

Fixed: Incident Slack notifications are no longer cut mid-sentence

An Incident triaged notification posted to Slack could stop partway through a word — in one report, mid-way through a run id — leaving a code span open so the rest of the line rendered as raw markup, and giving no sign that anything had been dropped. A truncated triage summary read exactly like a complete one that simply had little to say.Incident summaries now get roughly five times more room before anything is trimmed, so most reports arrive whole. When one genuinely is too long for Slack, it is cut at a word boundary, any open code block is closed, and the message ends with …(truncated) — with View incident still there for the full report.The same fix covers the resolved, failed, and needs-review notifications, which shared the behaviour.Try it in KAOP

New: Set your Enterprise Discount Program rates

K8s Cost → Settings has a new Discount Settings tab. Enter your AWS, Azure, and GCP Enterprise Discount Program rates so cost estimates across Allocation, Right-Sizing, and Pod Placement reflect your actual pricing. The tab also lets you set your account’s Pricing Fallback configuration — the default /CPUand/CPU and /GiB rates Komodor uses when real pricing data isn’t available.Try it in KAOP

New: Connect Loki from the Integrations page

Loki is now in the integrations catalog in its own right, so a Loki you run yourself — or a Grafana Cloud Logs endpoint — reaches your agents without a Grafana in front of it. Give it the Loki API base (not your Grafana URL), and the tenant if you run multi-tenant. Authentication is whichever shape your deployment uses: a basic-auth username and password, which is how Grafana Cloud Logs works — the username is your numeric Loki instance ID — or a bearer token, or neither, because a Loki running auth_enabled: false accepts an address alone. Testing the connection lists label names on your tenant, so a wrong address or a missing tenant is caught at connect time rather than on an agent’s first query. Point an agent at your own Loki MCP server and bind the connection to it, and AgentOps injects the credential and tenant on every call; agents that run their own MCP server receive them as LOKI_URL, LOKI_ORG_ID, LOKI_USERNAME and LOKI_PASSWORD/LOKI_TOKEN.Try it in KAOP

Improved: The alerts funnel now lives on the workflow it describes

The alerts funnel has moved off the Incidents module Overview and onto a workflow’s own Overview tab, where it is scoped to that one workflow. Open any workflow and the funnel sits between its configuration and its run statistics: every alert that reached this workflow, out to what became of it.It was on the module Overview because that is where it was built, not because that is where the question gets asked. “How much noise is this workflow filtering out, and is it filtering the right things?” is a question about one workflow, and it was being asked next to four KPI tiles about all of them — with a scope selector between the two doing the reconciling.One window control now drives the whole tab, so the funnel and the run statistics below it always cover the same days. The module Overview keeps its window and workflow-scope controls for the KPI tiles, and everything the funnel could already do — the drill-down behind each terminal, the values table, the conservation guarantee — is unchanged.Try it in KAOP
Improved

Improved: The routing list names its targets, and a pipeline’s trigger stays put

A channel’s routing rules now read as names rather than ids: an agent, workflow or incident pipeline is shown by its own name, falling back to the id only when it cannot be resolved. A channel’s status reads as a label (“Paused”, “Revoked”) instead of a raw value, and error and revoked are no longer missing from the frontend’s own list of statuses. A route that belongs to an incident pipeline is now shown read-only: editing, reordering, disabling, defaulting or deleting it here would have silently broken the pipeline’s Slack trigger, so those controls are gone and the pipeline owns the route. And a mention or DM that matches no route gets a short reply saying so, instead of silence.Try it in KAOP
NewImproved

New: The workflow Steps tab now shows what happens before your workflow runs

Every incident workflow’s Steps tab now opens with a “Before the workflow” group — three read-only cards for Alert trigger, Dedup, and Filtering, summarizing which endpoint feeds the workflow, the dedup-key strategy in force, and the filter conditions and rate limit every alert has to pass before an investigation ever starts. Each card links straight to the wizard step where that setting is edited.A workflow that hasn’t been provisioned with a step graph yet used to render a blank page. It now shows this same prelude plus a clear “No steps yet” message for the workflow section, so a pre-backfill workflow’s configuration is never invisible.Try it in KAOP

New: A run’s own alert trigger, dedup and filtering now show on its timeline

An incident’s run detail page now opens with the same “Before the workflow” group as the workflow’s Steps tab — Alert trigger, Dedup, and Filtering — but carrying this incident’s own facts instead of the workflow’s current configuration: which alert arrived, the dedup key it resolved to, and every fire that merged into it.A dedup key derived as random now says so plainly: the payload identified nothing, so this incident can never merge with a re-fire of the same alert — a re-fire opens a second incident instead. That used to be invisible on the run itself.A legacy incident with no workflow run — whose step breakdown is empty — used to render nothing at all here. It now shows this same prelude, with no step list beneath it.Try it in KAOP

Improved: The run timeline’s alert prelude now lives inside the Steps card

The “Before the workflow” group — Alert trigger, Dedup, Filtering — used to sit above the Steps card as three separate always-expanded cards. It’s now the first group inside the same Steps card, in the same collapsible row grammar as a workflow step: an icon in place of a step number, a dashed “Context” chip in place of a status chip, and no duration. A thin group header still marks it as pipeline context rather than a workflow step, with “Workflow steps” labelling the real steps below it.Try it in KAOP

Improved: Run cost now comes from the gateway that bills it

A run’s estimated cost is now read from your LLM gateway’s own per-model rates — the same table that meters the request — so the figure on a run and the figure the gateway bills you cannot drift apart. Cache reads and cache writes price at each model’s own published cache rate rather than one blanket multiplier.This replaces a price list that shipped inside AgentOps and had to be updated by hand. That list had gone stale: Opus runs were estimated roughly 3x too high and Haiku 4.5 runs roughly 4x too low. Because a run’s cost is recalculated when you open it, your existing runs show the corrected figure straight away — expect historical Opus costs to drop and Haiku costs to rise.A model your gateway does not serve now shows “Cost unavailable” rather than a number we cannot trace to anything that billed it, and the Cost tile says which model it could not price. Runs that never go through the gateway — an agent you host with your own API key, for instance — are in that group. A cost the worker reported itself is unaffected and still shown as reported.Try it in KAOP

New: See per-cluster pod placement savings, and turn it on right from the table

The Pod Placement page now has two tables below the resource allocation chart: Potential Savings lists clusters that could save CPU and memory by turning pod placement on, and Active Savings lists clusters where it’s already running. Sort any column, page through the results, and flip the “Pod Placement” switch on a row to enable or disable it for that cluster — turning it off asks for confirmation first.Requires the k8s_cost.policy.manage capability to toggle; without it, the switch stays disabled with a tooltip explaining why.Try it in KAOP

Improved: Pin an agent to any skill version, and preview a rollback before you run it

On a skill’s edit page, the Currently attached panel now lets you pin each agent to any published version — not just latest or the one it was already pinned to. Pick v2 for one agent while another tracks latest; publishing a new version never moves a pinned agent until you repin it.Rollback now shows you what you’re about to restore. Instead of a bare confirmation, Rollback opens a review with a diff against the current version (toggle to full content) so you can read the exact SKILL.md you’d republish before committing to it.Try it in KAOP

New: Incident workflows can ask you mid-investigation

A new incident workflow mode, Orchestrated loop, runs three steps — Investigate, Verify, Postmortem — and does the remediation inside Investigate rather than as steps of its own. The orchestrator investigates, delegates to your specialists, and pauses to ask you whenever the next decision is yours: which fix to apply, whether to apply one at all. You answer in Attention or on the incident, and the same investigation carries straight on with your answer instead of starting over. It can ask as many times as the work needs, and the time you spend deciding is no longer charged against the step’s time budget.Verification is mandatory on this mode, with one exception that matters: if nothing was actually applied — you declined every option, say — there is nothing to verify, so the workflow says so rather than spending a retry rediscovering it. That distinction comes from a record the workflow keeps of every change an agent makes, written before the change is attempted, so “this step succeeded” and “this step fixed the incident” are no longer the same claim. A failed verification re-investigates exactly once, and an incident that ends unresolved still gets written up.Try it in KAOP

Improved: An endpoint now shows how its payloads are read — and what a test request maps to

Integrations → Endpoints has a Test & mapping step. Send a test request (or wait for a real one) and the captured headers and body now come with what that request becomes: dedup key, status, title, SEV-1…SEV-5 severity, summary. The rehearsal runs server-side through the same extraction a real delivery uses, so it cannot disagree with what a live alert would produce, and it opens no incident.Beside it, in a collapsible section, the endpoint’s own field mapping — the mapping every incident workflow on that endpoint starts from. Each field arrives prefilled with the provider’s default path (with a link to that provider’s webhook-payload reference), so what you see is what gets stored rather than a suggestion, and switching provider rewrites only the fields you haven’t edited. Clearing a field hands that one decision back to the provider.The prefilled path is one that actually resolves, which is not always the path the docs quote: a Grafana payload stores fingerprint, not alerts[].fingerprint. And when a configured path resolves to nothing, the mapped fields now say so instead of quietly falling back to the default.The create wizard also has a Back button now. The endpoint is created when you leave the Authentication step, so going back there saves your changes rather than creating a second endpoint.Try it in KAOP

New: See where your alert volume actually goes

The Incidents module Overview now leads with an alerts funnel: every alert that reached AgentOps, out to what became of it. Filtered out by a workflow’s rules, dropped because nothing was listening, rate limited, folded into an incident that already existed, or opened as new work — and for the work that opened, whether it was fixed with no human, needed someone in the loop, was closed by hand, or resolved when the provider recovered.Until now none of that was answerable. An alert a workflow’s filters declined, or that arrived at a paused workflow, left no record anywhere — so “how much of our alert noise are we actually filtering out?” had no answer, and neither did “is a workflow quietly paused?”. Every arrival is now recorded with the reason it ended where it did.The numbers conserve: each terminal’s count sums exactly to the total that arrived, so a percentage read off one node means what it says. Click any terminal to see the alerts behind it, with the filter condition or rate-limit counter that put them there. The funnel shares the Overview’s existing window and workflow controls, and the four KPI tiles now sit below it.Two things the funnel is deliberate about: it counts only alerts that reached AgentOps, never a share of everything your monitoring fired; and the “Fixed, no human” figure carries its known over-reporting caveat, because a guardrail approval hold is not visible to the engine today.Try it in KAOP

Improved: Connecting a provider mid-wizard now says where its values come from

Deploying an agent from the catalog lets you connect a provider it needs without leaving the wizard — but that inline form asked for Azure’s four opaque GUIDs with nothing on the screen saying where any of them are found, while the Integrations page had said so since the start.The wizard’s Connect new form now opens with the same Before you start block: the exact command to run, a button to copy it, a line saying which output value goes in which field, and a link to the provider’s setup documentation. As on the Integrations page it comes from the provider, so it appears only where a provider has something to say.Try it in KAOP
NewImprovedFixed

Fixed: The Steps tab now names a human-selected workflow’s real graph

A human-selected incident workflow runs six steps — Investigate, Propose remediation, Select remediation (the decision an operator has to answer), Apply remediation, Verify, Postmortem. The Steps tab’s GET /incident-pipelines/{id}/steps endpoint used to describe every workflow as the four-step automatic graph regardless of which one it actually ran, so a human-selected workflow’s Steps tab named a “Remediate” step it didn’t have and never showed Select remediation at all.The endpoint now reads the workflow’s own instantiated definition — the same read the Team & steps graph view already used — so the Steps tab lists exactly the steps a workflow runs, in order, with the agents actually bound to each one.Try it in KAOP

Fixed: A human decision now sits where it ran on the run’s Steps card

On a human-selected incident workflow, Select remediation is the step where a person picks what gets applied. On the run drilldown’s Steps card it was the one step you could not place.It sorted after Apply remediation, Verify and Postmortem, with a blank duration — because a decision step dispatches no agent, so nothing ever recorded when it started, and the card orders by start time. An operator reading the page could not tell that the decision had run at all, let alone when, or how long it had waited for an answer.The moment the question goes live is now recorded as the step’s start. Select remediation sits in execution order between Propose remediation and Apply remediation, and shows how long it was open — whether it was answered or expired unanswered.Each row is also numbered by its position in the workflow graph rather than its place in the list, so the number matches the one on Team & steps even on a run that looped back to Investigate.Try it in KAOP

Fixed: A run now shows the guardrail that cleaned its input

A run whose only guardrail activity happened at the very start — a rule redacting something out of the payload before the run was allowed to begin — showed no Guardrails tab at all. The decision was recorded and enforced correctly; the run page just could not find it, because that check happens before the run exists and so is stamped a fraction of a second earlier than the run’s own start time.The tab now appears for those runs, and the row says what it is: “Before the run started”, marked on the payload this run started from. That is a stronger statement than the page can make about most other decisions — this run began from the cleaned payload, rather than merely overlapping with a rule firing nearby.Nothing changed about enforcement, and nothing changed for runs no guardrail touched — they still have no Guardrails tab.Try it in KAOP

New: Assign a policy and toggle automation right from the Right-Sizing table

The Right-Sizing table now has “Policy” and “Automated” columns, pinned to the right so they stay reachable without scrolling. Pick a policy to manually pin a workload to it, or switch it back to “Default” to fall back to whichever policy would otherwise apply. Toggle “Automated” to turn a workload’s right-sizing recommendations on or off — turning it off asks for confirmation first.Automating a workload requires a policy to already be assigned; the switch stays off until one is.Try it in KAOP

Fixed: Resource Requests charts open on the last 7 days, not 14

The Resource Requests chart on the Right-Sizing and Pod Placement pages was opening on a 14-day window. Every other cost page opens on the last 7 days by default — the chart now matches.Try it in KAOP

Improved: Reconnecting Slack brings your channel and routing back

Reconnecting a Slack workspace used to start you over: you got a new channel with default routing, and any rules you had customized stayed behind on the old one. Reconnecting now picks the original channel back up, so your routing rules come back with it. This applies whether you disconnected the integration yourself or the Slack app was uninstalled and reinstalled.If a rule pointed at an agent that has since been deleted, that rule cannot be restored — the channel comes back with default routing instead of routing nowhere.If the app was uninstalled from Slack, reconnecting also simply works again. It previously reported “this workspace is already connected to AgentOps” and left you with no working channel and no way forward.Workspaces disconnected more than 30 days ago reconnect fresh, with default routing. The older channel and its rules are still kept, not deleted.Try it in KAOP

New: Pod Placement now shows a Resource Allocation Optimization chart

The Pod Placement tab’s summary cards now sit above a “Resource Allocation Optimization” chart — current vs. optimal CPU and memory capacity across your selected clusters, plus how much of the gap between them is already being saved through automation and how much is still on the table.Switch between CPU and Memory, and between 7/14/30-day windows, the same way the Right-Sizing page’s own chart works.Try it in KAOP

Fixed: Operator-approved incidents keep their root cause in the postmortem

On a pipeline where an operator picks the remediation, the write-up at the end of the incident lost everything the investigation had established: the postmortem’s root cause and follow-ups read Not recorded, even though the investigation step had found both. The moment a person answered the question, the incident’s context stopped travelling with the run.Answering a decision now carries that context forward, so an operator-approved incident gets the same postmortem an automatic one does. A verification that fails and sends the run back to investigate also re-enters with the original alert instead of starting blind.Try it in KAOP

Improved: See what memory is learning, and from how much of the work

The Memory page now opens on an Activity tab that answers three questions at a glance: how many of the last seven days’ sessions taught an agent something, what became of every memory the worker reviewed, and which memories agents actually read while working. Each chart follows the space picker.Below them, the loop reads as one list instead of a Running and a History tab, with anything still running at the top. A run’s digest and its review are one row (memorizing, reviewing, then what the review decided), the running step is highlighted, and the list refreshes itself every minute while you watch it. Search, the memory list and the space bindings moved to their own tabs, and every tab is a shareable link.Try it in KAOP

New: Keep an agent’s tool traffic inside your own network

An agent you run in the same cluster as an Outpost can now be pointed at that Outpost directly, so its tool calls reach your private MCP servers without leaving your network at all. Pick the Outpost on the agent wizard’s MCP tools step and the values file it generates points there instead of at AgentOps.Self-hosted agents only — an Outpost’s listener is a Service inside your cluster, and a Komodor-cloud agent cannot reach it.These tool calls are not governed, and that is worth deciding on purpose. They never reach our gateway, so guardrails, policies and audit do not apply to them, and the Outpost’s local listener authenticates no caller: anything in that cluster which can reach it gets every tool in the groups that Outpost serves. A NetworkPolicy is what decides who may, and the chart now refuses to enable the listener until you have either let it create one or told it you manage your own. Guardrails you author at the tool gates say so when you have an Outpost that can serve a colocated agent.Two smaller changes came with it. A group that mixes servers behind the Outpost with servers reached any other way now works from a colocated agent, which previously saw only part of it. And an Outpost running an older configuration than we have rendered for it now shows the upgrade command on its own page, instead of only a badge.Try it in KAOP

Fixed: Updating a hosted agent now always takes effect

Saving a change to a hosted agent — a new prompt, a different image — could be accepted and then quietly never applied, leaving the agent serving its old configuration with nothing reporting a problem. Updates now roll out however many agents your workspace runs, and an agent that ends up unable to authenticate recovers on its own instead of sitting offline. One thing to expect: an update briefly takes the agent offline while it finishes any in-flight runs, rather than handing straight over to the new version.Try it in KAOP

Improved: The guardrail form asks what a rule does before asking what it matches

Writing a guardrail used to mean one screen holding every question at once — the name, where the rule stands, what it does on a match, and then a card per way-this-can-happen with its own tool globs and conditions underneath. That is now two steps. What it does is four short answers, each a name or a pick from a list. What it matches is the open-ended part, on its own.Nothing moved except where it is asked. The rule still reads the way it always did — on a gate, do an action, for every call that matches one or more criteria — and Add another way this can happen still adds a second criterion as an either/or.A new starting point on the empty state: “No payment card numbers reaching a model.” It arrives already pointed at what your agents send to a model, already set to redact rather than refuse, and already carrying the condition — so a rule that keeps card numbers out of your prompts is a name and a save.Two smaller things while we were in here. The operator that picks a class of value now reads “contains PII/secret” rather than “contains a”, which sat next to plain “contains” and read like a variation on it. And a Create button that cannot be pressed yet now says which step is unfinished, instead of leaving you to find it.Try it in KAOPThe recent-activity list under Govern → Guardrails told you a rule had fired and then left you to find the rest yourself. Now the agent name is a link to that agent, with its icon beside it, and the timestamp is a link to the run the decision was made in.The columns read in the order the question is usually asked: when · agent · outcome · rule · tool. The agent moved next to the timestamp because “which of my agents is this happening to” is what groups everything else, and the tool name — the longest value in the row — now ends it.A timestamp that is not a link means there is no run to open, not that something is missing. A rule that checks a trigger before a run has started has no run to point at, and neither does a check on the way to a model when the agent’s own framework does not tell us which run it was for.Try it in KAOP

Fixed: A failed incident step now tells you why it failed

Expanding a failed step in an incident’s run drilldown showed a red Failed badge and, when the step produced no output, nothing else — you had to reach for pod logs to find out what went wrong. The control plane had recorded the cause all along and returned it on every step; the drilldown just never displayed it.A failed or timed_out step now shows that reason in its expander, above the run link and the output. A step that timed out also reads Timed out instead of the raw timed_out, in the same red as a failure rather than the grey it shared with a step that hadn’t started yet.Try it in KAOP

Improved: Credential detectors now catch every GitHub, Slack and Anthropic token type

The credential detectors you can pick in a guardrail condition covered the common form of each credential and missed the rest. They now cover the family:
  • GitHub token — was personal access tokens only. Now also App installation tokens (ghs_, ghu_), OAuth (gho_) and refresh (ghr_). If your agents authenticate to GitHub through an app rather than a PAT — the usual case — this is the form that was slipping through.
  • Slack token — was bot and user tokens. Now also app-level (xapp-), configuration tokens (which mint other tokens), session credentials, and the legacy formats.
  • Anthropic API key — now also admin keys (sk-ant-admin01-), which administer the whole organisation.
  • Google Cloud API key — now also Google’s newer AQ.Ab8RN6… format.
  • OpenAI API key — now also the short-segment project and service-account form.
If you already have a rule armed on one of these, it will start catching more. That is the intended direction: a detector that misses a leaking credential fails at the one job it has. Nothing that matched before stops matching.Our credential patterns now track betterleaks rather than gitleaks, which its author has declared feature-complete and which points there itself. We take the patterns as data — never their option to verify a secret by sending it to the vendor, which in a redaction path would leak the exact value we were asked to contain.Try it in KAOP

Improved: Author a new skill without leaving the agent wizard

The Agent instructions step of the Create agent wizard now has a Create skill button beside the attach picker. When you have no authored skills yet — or just want a new one — you can author it inline instead of leaving the wizard for the Skills page and losing your draft.It opens the same create-skill dialog the Skills page uses, and the skill you create is attached to the agent’s draft at latest the moment you save — so it is on the agent from its first run.Try it in KAOP
NewImprovedFixed

Fixed: New workflows no longer start with remediation mode locked

Starting a brand-new incident workflow and clicking Next past the Name step used to silently create the backend workflow record right there — locking Remediation Mode to “Automatic” on the Team & steps step before you had a chance to choose. The workflow is now created only once you leave Team & steps, when the mode you actually picked is known, so both “Automatic” and “Human-selected” stay selectable until then.Try it in KAOP

Fixed: A tab left open across a deploy recovers instead of showing “Unexpected error

Each release ships its own set of content-hashed JavaScript chunks, so a tab that loaded the app before a deploy still points at chunk URLs the new release no longer serves. Opening any page from that tab was meant to reload quietly and pick up the new build. Instead it flashed a generic Unexpected error card reading Cannot read properties of undefined for the whole time the reload took, and the card’s own button did nothing while it was up.The recovery listener was suppressing the browser’s module-loading error rather than letting it reach the code that reloads, which left the page holding a module that had never actually loaded. Navigation from a stale tab now waits on the reload it already triggered, and anything that survives that reload gets a Stale UI after a deploy card with a working Hard reload.

Fixed: Setting up Slack MCP now asks for the access it actually needs

Slack’s hosted MCP server searches and reads as you, not as the AgentOps bot — it accepts your personal Slack token and refuses the bot’s. Connecting Slack asked for that separately, through a checkbox worded “Also connect my user account”, left unticked by default.So the requirement was met by accident or not at all. When it was not, nothing said so until tool discovery, which failed with “the provider rejected the stored credentials… add or replace the connection with a valid key” — advice that pointed at a credential that was working perfectly everywhere else Slack was used.Three things changed:
  • The MCP setup wizard states the requirement instead of offering it, and requests your access as part of the step. You still approve exactly what it grants — on Slack’s own authorization screen, which is where that decision was always really made.
  • A connection that cannot drive the endpoint is named before you bind it. One authorized for the bot alone shows as Bot only in the connection list and picker, with a one-click reconnect that adds your access; registering a server against it is blocked rather than left to fail later.
  • The error tells the truth. A refusal of this kind now reads “this connection was authorized for the bot only… reconnect the integration and include your user access” rather than blaming the credential.
Connecting Slack from the Integrations page for channels, notifications or messaging is unchanged — the bot token is all any of those use, and your personal access is still opt-in there.Try it in KAOP

Fixed: The Right-Sizing table shows every row, and the savings chart’s colors are fixed

The Right-Sizing workload table was capped at a handful of visible rows with its own scrollbar, even when you asked to see more per page. It now shows every row you’ve asked for, and only the page scrolls. Row-size choices are also back to the standard 10/25/50 you’d expect.The Resource Requests chart’s “Active Savings” and “Additional Potential Savings” bands had their colors swapped, so the chart under-represented how much was actually already saved — fixed.Try it in KAOP

New: Filter Right-Sizing recommendations by when a workload was last seen

The Right-Sizing table now has a “Last seen” filter, so you can focus on workloads that are actually still running instead of ones a cluster hasn’t reported on in weeks. Pick a window from the last 2 to 30 days and the table narrows to workloads seen at least once in that range.Try it in KAOP

Improved: Incident workflow steps get more time before they time out — and it’s now yours to set

Every orchestrated step of an incident workflow — Investigate, Verify, and the remediation step(s) — shared one hardcoded 300-second ceiling. Real investigations routinely used most of that budget, and some remediation steps timed out on their first attempt, only completing on a retry. Whether a step landed depended on retry luck more than on the work being tractable.The default is now 900 seconds, based on measured durations from real incidents. It’s also configurable per workflow: set Orchestrated step time budget in the wizard’s Team & steps step (or leave it blank for the new default), and — unlike Remediation mode — it stays editable after the workflow is created, so a ceiling that turns out too tight doesn’t require rebuilding the workflow.Try it in KAOP

New: Turn any OpenAPI document into agent tools

Add an integration and you can now point AgentOps at an OpenAPI document instead of registering an MCP server. Paste its URL or upload the file, pick the operations you want, and each one becomes a tool your agents can call, with the same groups, policies and credentials every other gateway server uses. Operations you leave off are not exposed. This was previously limited to accounts it had been enabled for.Try it in KAOP

New: Guardrail conditions can name what to look for, instead of a pattern

A condition can now say “contains PII/secret” and pick from a list — a payment card number, a US Social Security number, an IBAN, an email address, or one of eight credential types (AWS, Google Cloud, GitHub, Slack, Anthropic, OpenAI, JWT, PEM private key). No regular expression to write, and none to get subtly wrong.The reason to use it instead of a pattern is the checksum, not the typing. A regex for sixteen digits also matches order ids, invoice numbers and truncated timestamps — and a rule that refuses legitimate work all day is a rule you end up switching off. Payment card number verifies the card’s own check digit, so 1234-5678-9012-3456 is not one and 4111-1111-1111-1111 is.It works everywhere a condition works: on a tool call’s argument, on what your agent sends to a model (messages or system prompt), on what a tool returned, on what a model answered, and on a trigger payload before a run exists. Pair it with “Redact the value” to clean the value out and let the work continue, or with “Refuse the call” to stop it.Two things worth knowing before you test one. US Social Security number matches the separated form — 456-78-9123 — and deliberately not a bare 456789123, which is indistinguishable from an order number. It also does not match 123-45-6789 or 078-05-1120: those are published example SSNs, skipped so that the fixtures and tutorials in your own repositories do not trip the rule.Anything not on the list is still a pattern condition you write and own. Names, addresses and phone numbers have no reliable pattern — state those as a policy for a model to judge instead.Try it in KAOP

Fixed: Editing an incident workflow now shows the graph it actually runs

An incident workflow’s Remediation mode decides which step graph it runs — Automatic goes Investigate → Remediate → Verify → Postmortem, while Human-selected splits the middle into Propose remediation → Select remediation → Apply remediation so an operator approves the option before anything is applied.Opening an existing workflow for edit showed Automatic for all of them. The mode is fixed when the workflow is created and the graph is built from it, so there was nothing wrong with the workflow itself — but the wizard could not read the mode back, so it fell back to the default and drew the four automatic steps. A human-selected workflow was described as automatic, and its human decision step was invisible on the very screen you would go to in order to check for it.The mode is now part of what the API reports for a workflow, so Team & steps shows the mode the workflow was created with — still locked, because it cannot be changed after create — above the step graph that mode really runs.Try it in KAOP

Fixed: Disconnecting a Slack integration now releases the workspace

Disconnecting a Slack integration deleted the connection but left its channel behind — still listed as active, still holding the workspace it no longer had a credential for. Nothing said so, and the leftover row was not obviously anyone’s.That row is what a later connection attempt collided with, so reconnecting the same workspace, from this account or another one, came back as “already connected to a different AgentOps account” with nothing visible to disconnect. Disconnecting now revokes the channel and releases the workspace, so the next connect works.The channel is kept rather than deleted, so nothing about your setup is thrown away when you disconnect.Try it in KAOP

Fixed: Deleting an incident workflow no longer erases its incidents without asking

Deleting an incident workflow also deletes every incident it produced and the agent runs recorded against them — and it used to do that silently. Now the delete is refused until you confirm: the workflow you are removing tells you exactly how many incidents and incident runs would go with it, and the delete only proceeds once you tick the box that says so. A workflow that has no incidents still deletes in one click, and pausing a workflow remains the reversible way to stop it without touching the history it already recorded.Try it in KAOP

Improved: The connect form now says where its values come from

Connecting Azure meant filling four fields that all want an opaque GUID, plus a secret, with nothing on the screen saying where any of them are found. The one az command that emits every one of them was in the docs, and the link to those docs was not rendered anywhere.The connect dialog now opens with a Before you start block: the exact command to run, a button to copy it, a line saying which output value goes in which field, and a link to the provider’s setup documentation. The mapping matters more than it sounds — Azure’s command prints three of its four values, and the fourth is an argument you passed in, so it appears in no output at all. Fold the block away once you have read it.For Azure the block also states what the access it asks for actually grants — Reader can read every resource’s configuration and tags, your full network topology, all Log Analytics data, and Key Vault secret names, but never secret values, and it cannot write or run anything. That is the part worth reading before you approve a service principal, and it was previously only discoverable by reading Microsoft’s role definitions.The block comes from the provider, not from the form, so it appears only where a provider has something to say. A token you copy off a settings page still gets the form it had before.Try it in KAOP

New: Connect Dynatrace from the Integrations page

Dynatrace is now in the integrations catalog, so problems, DQL queries, logs, metrics and Davis AI reach your agents with no custom setup. Paste a platform token (dt0s16.…) and your environment URL — exactly https://<environment-id>.apps.dynatrace.com, not the classic .live. address — and AgentOps derives the hosted MCP endpoint for that environment itself. The classic API token (dt0c01.…) is an optional second field, only needed for the classic environment API v2. Agents that run their own Dynatrace MCP server instead receive the credential as DT_PLATFORM_TOKEN, DT_API_TOKEN and DT_ENVIRONMENT. Dynatrace SaaS only — Managed serves no hosted MCP endpoint.Try it in KAOP

Fixed: A workflow that ran out of retries is no longer reported as errored

An incident workflow can stop because a step genuinely failed, or because it ran every step to completion and simply had nowhere left to route — an exhausted retry cap, or an output shape no transition matched. Both used to show up on the incident as Errored, with a message telling you to retry a run whose retries were exactly what ran out.The two are now told apart. A step that failed still reads Errored, with a Retry button — retrying re-attempts it and might succeed. A run that ran out of retries now reads Needs review: no Retry button (retrying would only repeat finished work), and copy that says there is nothing left to retry — read the findings and resolve it manually, or update the workflow’s template.Try it in KAOP
NewImprovedFixed

Fixed: Editing a self-hosted agent no longer breaks its Helm upgrade

Upgrading a self-hosted agent after editing it could fail while Helm rendered the chart (nil pointer evaluating interface {}.enabled), if the release had been installed from an older chart version. The chart now tolerates the values of any release it supports, so the command the Edit wizard gives you works again — no change needed on your side beyond pulling the current chart. Your worker token and LLM key are carried forward as before.Try it in KAOP

New: Redact a value out of a tool call, instead of refusing the call

“Redact the value” now works at Tool call — the first of the two outbound gates, and the first place a rule changes what your agent sends rather than what it reads. The tool runs, with every value your conditions match replaced by [redacted] in its arguments.This is the one to reach for when a credential, a card number or an internal hostname keeps ending up in a ticket body, a search query or a webhook payload. Refusing those calls stops the work; cleaning them lets the work happen without the value leaving.One difference from the inbound gates, and it is worth knowing before you pick it. When we redact something an agent reads, it sees [redacted] and can reason about it. When we redact something it sends, it is not told — so it goes on to interpret whatever the tool made of a value it never sent. That is the right trade when keeping the value out of the tool matters most, and the wrong one when the agent has to act on the answer. “Refuse the call” tells it plainly; this does not.A redacting tool-call rule needs a condition naming the argument. Tool patterns on their own describe an action rather than a value, so there would be nothing to remove — that rule is a “Refuse the call”, and the form says so instead of saving something that could only ever refuse.The rules you already know are unchanged. All-or-nothing: if a matched value cannot be removed, the call is refused instead. Two redacting rules both apply. A rule that refuses still beats both — and so does one that holds for approval, so nobody is ever asked to approve a call whose arguments we quietly rewrote first.That leaves Model request as the last gate without it.Try it in KAOP

New: Redact a value out of the prompt — the last gate gets it

“Redact the value” now works at Model request, which was the last of the five gates without it. The request reaches the model with every value your conditions match replaced by [redacted] — in the messages, in the system prompt, or both, depending on what your conditions point at.This is the gate that can promise a value never went to a model provider at all. If a card number, a customer identifier or an internal hostname keeps arriving in a prompt, refusing those requests stops the agent working; cleaning them lets it work without the value leaving your account.The outbound trade applies here too, and it bites differently. When we redact something an agent reads, it sees [redacted] and can reason about it. When we redact something it sends, it is not told. At Tool call that means a tool acting on a value the agent did not send — the tool either works or it errors. At Model request it means a model answering a question with a hole in it, and a model will answer confidently either way. The hole shapes the answer and nothing in the exchange says so. Reach for “Refuse the call” when the agent has to act on what comes back.Redacting the system prompt is the sharper version of that. The system prompt is your text, not the conversation’s — so a rule pointed at it removes part of what the agent was told to do, rather than part of what it was told about. Worth pointing at messages unless you mean the instructions.Everything else is as it is at the other four gates. All-or-nothing: if a matched value cannot be removed, the request is refused rather than sent half-cleaned. Two redacting rules both apply, and each gets its own line in the audit trail. A rule that refuses beats both, and so does one that holds for approval — nobody is asked to approve a request whose prompt we quietly rewrote first.redact now stands at all five gates.Try it in KAOP

Fixed: Reconnect now repairs an integration instead of failing

When an integration’s key or token stopped working, the Reconnect button on it could not actually repair it: submitting new credentials came back with “an integration named X already exists”, blaming the connection’s name for what was a credential problem. The only thing that did work was disconnecting and setting the connection up again from scratch.Reconnect now replaces the credential on the connection you are already looking at. The name, and anything bound to it, stay as they were — so agents using that integration keep working, rather than quietly losing access the way they did when the connection had to be rebuilt. A key the provider rejects is refused before it is saved, which means a failed attempt leaves the old credential in place instead of replacing a working one with a broken one.This covers every integration you connect with a key or token, including AWS, Datadog and Grafana Cloud — useful whenever a secret expires or you rotate one on a schedule.Try it in KAOP

Fixed: A Slack workspace connects once, and reconnecting repairs it

Connecting a Slack workspace your account already had produced a second connection rather than an error. Both were marked active and both held a live bot token, so replies could go out through one while messages came in on the other — and the second attempt failed outright after the damage was done, leaving behind a channel that could never receive an event.An account now holds one connection per Slack workspace. Re-running the connect flow for a workspace that is already there does the useful thing instead: if its credential has been rejected — the app was removed from Slack, the token was revoked — the fresh authorization replaces the dead one in place, and the connection starts working again without you having to find it first. If it is healthy, the attempt is refused and says which workspace it already belongs to. A provider that cannot be reached counts as neither, so a network blip can never be read as a broken credential.Connections are also named after the workspace itself now. A row reads “Acme Corp” rather than “Slack 2”, so which is which no longer depends on remembering the order they were added in — though a name you set yourself is left alone.The same guard covers re-installing a GitHub App into an organization it is already installed in, which had the same duplicate behaviour.Try it in KAOP

Improved: Search and per-integration grouping in the integration-group picker

The Add group dialog for integration groups got three improvements:
  • The member-server checklist now has a search box, so picking a few servers out of a long list no longer means scrolling to find them.
  • The Allow/Deny tool list now scrolls correctly with a mouse wheel or trackpad — it previously only responded to keyboard navigation.
  • Tools in the Allow/Deny list are now grouped under the integration they belong to, each with its own “select all” checkbox, instead of one flat alphabetical list with no indication of which server offered which tool.
  • Each selected rule now names the integration it matches, and hovering it shows how many tools it covers on each one. A wildcard rule like *delete* reports its real reach across every member instead of reading as an unknown tool.
Try it in KAOP

Fixed: Incidents stop toasting “Not found

Opening the incidents inbox or an incident no longer raises a “Request failed / Not found” toast every 15 seconds. The page polls an inbox that is only present once module workflows are enabled for your account; where they are not, the affordances that need it are simply absent instead of erroring. A genuine failure still surfaces as before.Retry on an incident whose workflow engine is switched off now says so and leaves the incident where it was, instead of reporting success and quietly waiting on a step nothing would run.Try it in KAOP

Fixed: Incidents and runs get their names

An incident whose alert carried no title of its own now gets a short readable name instead of “Untitled alert”. A brief LLM summary of the alert payload is written over the placeholder in the background, as {workflow} — {five words}, so the inbox is scannable at a glance. Standalone agent runs get the same treatment on Fleet and History.This was built to work that way and never did: the model it calls was configured in no environment, so every attempt failed closed and quietly kept the placeholder. It now runs where the incident webhook and the run pipeline actually execute, and a failure is logged loudly instead of silently.An alert that does name itself keeps that name. The summary only replaces the placeholder, never a title your monitor already wrote — a Datadog or AlertManager alert reads exactly as it did before, and costs no model call.On the incident page, Orchestrator run and Specialist runs now show the agent’s name — Klaudia Investigator, rather than run_60b72c82c53865e17da3244b. The run id is still there on hover, and still the link.Try it in KAOP

New: See which notification sinks fired for each incident

Incident workflows and individual incidents now show a Notifications strip that tells you exactly which sinks were notified — and what happened when they were.On the workflow overview tab, the strip lists every notification sink subscribed to that pipeline. On an incident detail page, it shows which sinks fired for that specific incident, the event type (Triaged, Resolved, Errored), each delivery’s outcome and timestamp, and the error body when a delivery failed or was retried.Sinks that were skipped — because the sink was disabled or the incident event type was not in the sink’s filter — appear with a dash and a tooltip explaining why nothing was sent.Try it in KAOP

Fixed — “Hold for approval” on a model request let the request through

“Hold it for approval” could be set on a Model request rule through the API or IaC, and when it fired, the request went to the model anyway. Nothing said so. The rule appeared armed, its decision was recorded, and the traffic it was written to gate passed.The cause: an approval works because the boundary comes back to ask a second time. At Tool call it does — the call is refused naming an approval, and the agent’s retry finds it granted. The model-request boundary has no second ask; it refuses a request outright or sends it. So there was never anything to resolve a hold there, and the proxy — which checked “is this a refusal” rather than “is this something I know how to carry out” — treated the unfamiliar verdict as permission.Both halves are fixed. The model boundary now refuses anything it does not recognise, rather than passing it, so a verdict from a newer control plane can never again read as approval. And hold is no longer accepted on a model-request rule at all — the API answers 422, so nobody arms an approval that cannot be answered.If you had such a rule, it has been rewritten to “Refuse the call” — which is what it has been doing since the boundary fix, and what its author most likely wanted. Nothing was deleted. If you wanted the value kept out of the prompt rather than the request stopped, “Redact the value” now works at that gate.“Hold it for approval” is unchanged everywhere else, and still the right tool at Tool call inside a workflow step.Try it in KAOP

Improved: Setting up a provider’s MCP server walks one step at a time

Connecting a provider’s hosted MCP server used to show all of its steps at once, with the ones you could not act on yet greyed out. It now walks through Connect account, Expose tools and Route — the same shape as adding your own MCP server or importing an OpenAPI document.The tool list also reads what a server actually publishes. A server set to publish nothing, or everything except a named tool, used to show every tool ticked — and ticking anything would quietly republish the lot. Those are now shown as what they are, and a selection written as patterns is left alone rather than offered as checkboxes that cannot express it.Publish every tool is a switch on that step now. Left on, a tool the provider adds to its server later is published too; turn it off to pick from the list.Try it in KAOP

Improved: An agent-start rule now refuses a webhook at the door

A rule at Agent start is the only one that can say a value was never kept, rather than that we did not pass it on. That was true of the run, and not quite true of everything around it: an incident pipeline wrote the alert body, a review workflow wrote four fields off the pull request, a workflow webhook queued the whole payload and answered 202 before any rule was asked, and an endpoint capturing samples stored the body whether or not anything ran. The rule then refused the run, next to a copy of what it refused.Those four are now asked at the edge, on the body as it arrived. A refusal answers the sender 403 naming the rule, and nothing derived from the payload is written — no incident, no queued trigger, no sample. A redacting rule cleans the sample as well as the run, so the two no longer disagree about what arrived. A body that is not JSON is read too, since otherwise the way past every condition was to send one.One thing to know when you write the rule: leave it untargeted if the point is storage. A rule that names agents cannot be applied at the edge, because at the edge we do not yet know which agent will run — an incident fans out per pipeline and a review per bound reviewer, and picking one would refuse the others’ traffic. Those rules are enforced a step later, when the run is created, which is after the ingress row exists. An untargeted rule has no such ambiguity and is asked at both.Try it in KAOP

Fixed: The K8s Cost Overview page’s Cost Trend chart now shows real dates

Every bar on the Cost Trend chart was labeled with the same date, making the 30-day trend unreadable. The chart now shows the correct date under each day’s spend.Try it in KAOP

New: Connect an Azure subscription

Azure now sits in the connection catalog beside AWS: connect a subscription with a Microsoft Entra service principal, and agents you bind it to can read that subscription — Resource Graph, Monitor metrics, Log Analytics, and AKS cluster metadata.The three IDs are checked before anything is stored, so the two mistakes that used to be invisible are caught in the form rather than becoming a connection you would have to delete and rebuild: the Object ID pasted where the Application (client) ID belongs, and the tenant ID pasted into the Subscription field. Pressing Test names which of the four values to fix, and an expired client secret says so outright. Setup and what a subscription-scoped Reader role exposes are in the Azure integration reference.Try it in KAOP

Improved: Six kinds of config now take a name that is unique in your account

Review workflows, notification sinks, eval flows, golden scenarios, golden suites and triggers each had a name that was free to repeat. Nothing needed one to be a key until something tried to apply the same configuration twice and could not tell which row it meant.Names in each of those six are now unique per account. Existing duplicates were renamed once, in place: the oldest row of each set keeps the name it had, and every other gets a numbered suffix — Nightly, Nightly (2), Nightly (3) — skipping any suffix already in use. Nothing is addressed by id, so no link or webhook URL moved. A create or rename onto a name already taken now comes back as a plain conflict naming the name, instead of an internal error.One place does resolve a trigger by name: an incident pipeline’s triggerRef. Where an account had two triggers of one type sharing a name, such a ref matched whichever was created most recently; after the rename the surviving Nightly is the oldest one, so the ref settles on that row instead. Only an account that actually had duplicate trigger names is affected, and the ref is unambiguous from here on.Triggers are the one with a wrinkle worth reading. The name is unique per account per trigger type, so a schedule and a webhook may both be called nightly. Triggers your worker declares in its agent-spec.yaml are outside the rule entirely — the heartbeat owns those names and may repeat them — and a trigger with no name at all is unaffected.One behaviour genuinely changes. A trigger name declared in an agent’s manifest at create time is stored as if you had made it yourself, so it is covered: two agents in one account can no longer each declare a trigger named nightly. The second agent is still created; its trigger is reported as failed with the conflict named, and it succeeds once one of the two is renamed.Both trigger lists also take include_agent_spec=false now, which narrows them to exactly the rows the new rule covers. The default is unchanged, so nothing you read today looks different.

Improved: A catalog agent can state its scope before you name it

Some agents reach less far than their name suggests, and the wizard used to let you find that out after the install. A catalog agent can now carry a Notice on the first step of the Add agent wizard — the scope stated where you pick the name, with a checkbox you tick before Next will move.The first one is Kubernetes RCA: a read-only investigator for the single cluster it is deployed in, with Klaudia the recommendation for multi-cluster RCA, and a suggestion to postfix the agent name with the cluster name — because you deploy one per cluster, and the name is how you tell their runs apart.Only an agent that needs one gets one. A notice on every card would train everyone to tick without reading, which would cost the agent that genuinely needs it its only defence. Every other catalog agent’s first step is unchanged.The co-pilot cannot tick it for you either. Ask it to fill the wizard and it quotes the notice back and waits for your answer — the acknowledgement is yours to give.Try it in KAOP

New: Get notified when agents are created, deployed, or go unreachable

Notification sinks now support agent lifecycle events — six new event types you can subscribe to from the sink wizard:
  • agent.created — fires when a new agent is registered in your account
  • agent.updated — fires when an agent’s configuration changes
  • agent.archived — fires when an agent is archived
  • deployment.succeeded / deployment.failed — fires when a hosted-agent deployment completes or fails
  • agent.unreachable — fires when an agent goes offline
agent.created is account-wide: because the agent doesn’t exist yet when you configure the sink, the subscription step automatically adds an account-wide agent subscription and explains why it can’t be scoped to a specific agent. All other agent lifecycle events can be scoped to specific agents or set to receive events from all agents.Try it in KAOP

Improved: Add an integration from where you noticed you needed one

Two places used to stop with nothing you could act on when a workspace had no integrations yet.Building an agent, the MCP tools step offered to create an integration group — and the dialog it opened had nothing to put in one. It now asks for the integration first, because a group is assembled out of integrations, and only asks about the group once there is something to group.The Add integration group dialog says the same thing in its member list, and offers the way out rather than the sentence “Register a server first.”Both open the integration catalog over the page you were on. The agent you were part-way through building is still there when you come back — and if the flow you picked routed what it connected into a group on its way out, that group is already selected.Try it in KAOP
NewImprovedFixed

Improved: A detached credential goes back, and two more fields can be declared

An MCP server’s credential is now put back when it is detached outside your manifests. It was the one binding the operator pushed only when the manifest itself changed, so detaching it here left the resource declared and diverged at once, under a green condition. The next resync re-binds it, with no edit to the custom resource — and a server already bound to what it declares is still never re-written, so a cluster where nothing changed still costs no writes.An incident workflow can now name its endpoint by name. spec.triggerRef takes the endpoint’s name where spec.triggerId took an opaque trg_… id: readable in review, still right after the endpoint is recreated, and portable to a second workspace. The id form keeps working for an endpoint nobody named.And an MCP server on an OAuth client-credentials grant can now declare spec.oauthClientSecretRef. That was refused before, because AgentOps served no way to replace a stored client secret — so a manifest could set it once and never again. Rotating the Kubernetes Secret now reaches the server.Try it in KAOP

Fixed: Scheduled runs now reach the shadow arm

An agent whose real traffic is a schedule — a nightly report, a recurring sweep — now produces comparison pairs while a shadow is live, the same as every other way a run starts. Each cron fire creates the production run and its shadow together, linked by one comparison id.Before, a cron-fired run was the one path that never paired. The experiment looked completely healthy from the outside — the candidate read online, Fleet showed the production and shadow split — and the candidate claimed nothing, forever, with nothing anywhere explaining the empty comparison.Two limits are worth knowing. Scheduled runs pair from the database rather than the application, so the fleet-wide fan-out kill switch does not reach them: turning it off halts pairing everywhere else and leaves schedules pairing. And a paired fire counts twice toward an agent’s queue backlog, which can scale the production deployment up for work only the shadow may claim.Try it in KAOP

New: Remove a configuration cluster from its own row

A cluster you added to try the config-plane operator can now be deleted from the Configuration clusters list. Until now the row offered only Install and Pause/Resume, so removing one meant calling the API by hand.The confirmation names what deleting actually costs, because two thirds of it are easy to guess wrong. The resources the cluster declares are not deleted — they stop being managed and become editable here, and the dialog says how many. Its credential is revoked, so an operator still running in that cluster stops reconciling rather than carrying on quietly. And unlike a pause, nothing reverses it: the ownership rows are released in the same transaction as the delete, and the released count survives only in the audit record.Rotating a cluster’s credential is still not a row action. It invalidates the credential the running operator is using, and nothing in this view can carry the new one to the cluster.Try it in KAOP

New: Redact what the model answered, instead of withholding the whole reply

“Redact the value” now works at Model response, the last of the three inbound gates. Pick it and the agent reads the model’s answer with every value your conditions match replaced by [redacted], instead of an error where the answer should have been.At this gate the difference matters more than anywhere else, because blocking here is the most expensive refusal we have: the completion already exists and is already billed, so withholding it spends the tokens and delivers nothing. Redacting spends the same tokens and delivers the answer, minus the part your rule was written about. A long analysis that mentions one internal hostname is still a useful analysis without it.The rules are the ones you already know. All-or-nothing: if a matched value cannot be removed, the answer is withheld instead and the decision record says the redaction is what degraded. Two redacting rules both apply. A rule that withholds still beats both. And it takes conditions rather than AI judgment, because removing a value means knowing where it is.Streaming is unchanged — a streamed answer is still refused rather than inspected, because there is no whole completion to read. Redaction does not change that: a rule that cannot see the text cannot clean it.That leaves the two outbound gates — the tool call and the model request. They are a different problem rather than a smaller one: those change what your agent asked for, so they owe it a way to know its own request was rewritten before we can offer this there honestly.Try it in KAOP

New: Redact what a tool returns, instead of withholding the whole result

“Redact the value” now works at Tool response as well as Agent start. Pick it as the action and the agent reads the tool’s answer with every value your conditions match replaced by [redacted], instead of getting an error where the answer should have been.This is the softer sibling of blocking at this gate, and at this gate the difference is worth a lot. The tool has already run either way — its effects stand — so refusing the result buys nothing except the agent not seeing the content. Redacting keeps the part of the answer the agent actually needed. A file listing that happens to contain one private key is still a useful file listing without it.It reaches structured content too, not only text blocks. A tool answering in JSON is not a way around a rule about what your agents may read, and you should not have to know which of your tools do that.The rules that make it safe are the same ones you already know. It is all-or-nothing: if a matched value cannot be removed — including when removing it would break the result’s own JSON — the result is withheld instead, and the decision record says the redaction is what degraded. Two redacting rules both apply rather than competing. A rule that withholds still beats both. And it takes conditions rather than AI judgment, because removing a value means knowing where it is.Still to come at the remaining three gates — the tool call, the model request, and the model answer.Try it in KAOP

New: Redact a value instead of refusing the whole run

A guardrail at Agent start can now do something other than let the payload through or refuse it. Pick “Redact the value” as the action and every value your conditions match is replaced with [redacted] — then the run starts, on the cleaned payload. Dropping an entire order because one field held a card number is worse for the customer than removing the field, and this is the action that says so.Because the rewrite happens before the run’s row is written, the original value never enters your storage. This is the same property that makes Agent start the only gate that can promise it, and the reason redaction is offered here first: the other four gates stand on traffic passing through us, so the most they could offer is removing a value from what we forward.It is all-or-nothing, and it fails towards refusing. If a value your rule matched cannot be removed — because the excision would break the payload’s own structure — the run is refused instead of started. A value we found and could not remove is never let through dressed as a success. So a redacting rule can still stop a run, and the decision record says that is what happened.Two redacting rules both apply. They are not in competition: two rules removing two different values is not a contest between them, so what comes out is everything either of them found. A rule that refuses still beats both.Redaction takes conditions, not AI judgment. Removing a value means knowing where it is, and a judged policy answers whether something was violated, not where — so the wizard offers the action only on rules built from conditions, and says so rather than quietly doing something else. It is offered only at the gates that can carry out a rewrite; pick another and the wizard explains why instead of the option disappearing.Unlike “Allow it, and record it”, redaction is visible to the agent: what it reads has already been changed, so hiding that would leave it reasoning about a question it did not ask. It is told which rule acted and how much came out — never what the removed value held.Try it in KAOP

New: Reach an OpenAPI API that only your own network can resolve

An OpenAPI integration used to be refused the moment the API it calls was not reachable over public HTTPS, which left every internal service out. The Connect step of the import wizard now offers an Outpost: pick the relay already running in the network the API lives in, and the integration is reached through it.With an Outpost picked, the API base URL stops being optional and may be a private or plaintext http address, because the relay is the only thing that has to reach it. The document is unchanged in that respect: AgentOps still fetches it over public HTTPS, or you upload it.Every call the generated tools make then travels that relay’s tunnel instead of leaving the gateway’s own network, and the rule the relay is given is derived from the base URL you named, not from wherever the document is served. Binding one re-renders that Outpost’s configuration, so it reads Upgrade pending until its operator re-applies the install values, or Config unknown if the relay is not running yet.Clearing the Outpost puts the API back on the gateway’s own egress, so the base URL has to become a public HTTPS address, or be blanked, in the same edit. The wizard says so before it lets you save.Try it in KAOP

Improved: One bring-your-own tile instead of two

Add integration used to show a separate tile for registering your own MCP server and for importing an OpenAPI document, reading like two different products rather than two ways into the same one. They’re now a single Custom integration tile — its Add button opens a small chooser naming both paths, then hands off to the exact same wizard as before.Try it in KAOP

Improved: Connecting an MCP server is a guided setup now

Registering your own MCP server used to be one long form: thirteen fields at once, including two boxes asking for glob patterns like get_* to decide which tools the gateway would publish. It now walks through Connect, Configure and Review, the same shape as importing an OpenAPI document.Configure lists the tools the server actually advertises, so you tick the ones you want instead of describing them. A tool you leave out has no tool at all — an agent never sees it. The pattern boxes are still there for the cases that need them, and for a server that cannot be read until it exists.Review now asks which group the server joins. A group is the endpoint an agent connects to, so a server in no group reaches nobody — previously you had to know that and go to the Groups tab afterwards. It names what the group already reaches and how many agents gain the tools you just exposed.For a server that signs in through the gateway, the tools stay unreadable until it is saved and authorized. That step now happens in place: finish the sign-in and the wizard offers the tool list rather than closing.Static headers are name and value rows on both this wizard and the OpenAPI one, instead of a JSON object you had to get right by hand.Try it in KAOP

Fixed the MCP group dialog stretching past the screen

Adding or editing an integration group with a lot of MCP servers used to stretch the whole dialog downward, pushing Name, Allow/Deny, and Create hundreds of pixels off screen. The member checklist now scrolls within its own bounded box, so the rest of the dialog stays in view no matter how many servers there are.Try it in KAOP

Fixed cost reporting for Kimi and other non-Anthropic models

Runs on Kimi (and any other model reached through litellm rather than Anthropic directly) were showing roughly 8x their real cost. The run cost calculator was trusting a number the Claude Agent SDK invents for models it doesn’t recognize, instead of pricing them from real token counts at their actual published rate. Cost and spend numbers for these runs are now accurate going forward.Try it in KAOP

Fixed: Identical alerts now merge into one incident

A plain webhook that sends the same alert twice now updates one incident — refire count, timeline entry, and a merged_alerts list of every fire that landed on it — instead of opening a second incident.Providers that name their own alert cycle (Datadog, Grafana, AlertManager, PagerDuty, OpsGenie) always merged correctly. A generic endpoint that sent no dedup_key or fingerprint did not: the ingest gave each fire a random key, and a random key can never match, so every POST opened a fresh incident with a refire count of zero. An alert that put its name at the top level rather than under labels hit this even though it looked perfectly ordinary.Such a payload now gets a key derived from what the alert is and what it is about — alertname, service, environment, namespace, host and their siblings — deliberately ignoring anything that changes between fires, so a repeat hashes to the same key. Sending your own dedup_key or fingerprint still wins outright.If a payload carries nothing identifying at all, every fire still opens its own incident. That is now stated rather than silent: the incident reads dedup_key_source: random and opens with a Dedupe key unavailable timeline entry naming the fix.Try it in KAOP

Fixed: Human-selected remediation mode now activates

Choosing “human-selected” remediation mode on the Team & steps step of a new incident workflow now takes effect — Save Draft and Activate both used to fail with “Remediation mode is chosen when the workflow is created and cannot be changed to ‘human-selected’”, because the workflow was silently created as automatic three steps earlier, before you ever reached the mode picker. The workflow is now created once the mode (and the rest of the team) is actually chosen, so the pick you make is the one that ships.Try it in KAOP

New: End a hosted shadow from the agent’s page

A hosted agent running a shadow now has Remove shadow on its page, the same one-click teardown an agent you deploy yourself already had. It removes the shadow’s release, its worker token and its GitOps directory; the main agent keeps running on its own token, with its config untouched.Before, hosted was the one arm with no teardown in the product. Fleet showed the experiment running — the replica count split into production and shadow — but the shadow has no row of its own to act on, so ending one meant an API call. The two paths that did exist in the UI were far too blunt: deleting or archiving the main agent sweeps the shadow away with it, and ends the agent too.The confirmation says what actually happens here. Ending a shadow you deploy yourself tells you your candidate is still yours to remove; for a hosted shadow AgentOps removes the release itself, so that sentence would send you after a cluster you do not have.One thing worth knowing if you scripted around this: the API teardown, DELETE /api/v1/hosted-agents/{customer}/{agent_id}-shadow, used to refuse with “archive the agent before deleting it”. That guard is there to stop a live agent being decommissioned by accident, and it was never meant for a shadow — a shadow’s teardown leaves the agent serving. It no longer applies to the shadow arm, so the call now succeeds directly, from the UI and from chat alike.Try it in KAOP

Improved: The attention banner collapses when a lot needs fixing

An account with more than a handful of failing credentials used to see all of them listed at once above the Integrations table, pushing the table itself below the fold. The banner now shows the four most important failures — widest blast radius first — with a Show N more toggle for the rest.It can also be dismissed with the × in its corner. Dismissing only hides the failures you’ve already seen — a credential that fails afterward reopens the banner rather than staying silently tucked behind an old acknowledgement.Try it in KAOP

Fixed: An incident now says when its agent is offline

An alert routed at a workflow whose orchestrator agent had no worker running was accepted and then did nothing — the incident sat at Open with no error, no run, and a timeline holding only “Alert received”. Such an incident now carries an Agent offline badge in the feed and a timeline entry explaining that the investigation is queued until the agent reconnects, and a workflow bound to an agent that is not running shows that agent as Offline on the workflow detail page. When the bound agent no longer exists at all, the incident is marked Errored with the reason instead of waiting indefinitely.Try it in KAOP

Fixed: A workflow’s run duration reads as a duration again

The Run statistics panel on an incident workflow now shows the mean duration of the runs in the window — 1m 30s, or an em dash when no run recorded both ends.Before, that row read NaNs at every window. The panel asked the API for a field name it had stopped sending, got nothing back, and formatted the nothing. Every other number on the card — runs, tokens, cost, errors — was correct the whole time, which made the one broken row easy to read as a real measurement of something.Durations elsewhere on the card are hardened the same way now: a value that is missing or is not a number renders as an em dash rather than as arithmetic on it.Try it in KAOP
NewImproved

New: See the incidents one workflow produced

An incident workflow now has a Feed tab beside Overview and Steps, listing only the incidents that workflow created — severity, status, who it is waiting on, and the investigation run, in the same rows as the incident inbox. Answering “what has this workflow actually caught?” no longer means going back to the inbox and reading the workflow column row by row.Try it in KAOP

Improved: Every server sits with the integration it belongs to

The separate Servers section is gone. Integrations already listed every server — the ones sharing an integration’s credential grouped under it, the rest as rows of their own — so the two screens were showing you the same records and you had to know which one to be on to do anything.Now the row is where the work happens. Open the ⋮ menu on any server for Manage tools, Edit, Delete, and, on servers that sign themselves in, Login — the same forms as before, reached from the row you were already looking at instead of a second screen. Test stays where it was.Two things the row now tells you without being asked. A server built from an OpenAPI document is marked OpenAPI, so it is obvious which ones are edited through the spec wizard rather than the server form. And a server reaching its upstream through an Outpost names that Outpost — flagged not routing when the Outpost has been switched off, which until now looked healthy from every angle while refusing every call made through it.Reload gateway moved to the top of the page, beside Add integration, because it always applied to the whole gateway rather than to one server.Old links to the Servers section still work — they land on Integrations.Try it in KAOP
NewFixed

New: Import an OpenAPI document by uploading it

An OpenAPI integration no longer needs its document published at a public HTTPS URL. The Connect step now asks where the document comes from, and Upload the document takes the file itself: a spec generated into a build artifact, exported from an internal portal, or simply handed to you.Everything after that is the same. The same operation picker, the same tools, the same groups, policies and credentials.An uploaded document is stored with the integration, so opening the editor lists every operation it declares without fetching anything, and you can widen the selection later. Choose a file again to replace it. Refresh, which re-reads a document from its URL, reports that there is nothing to re-read for an uploaded one and points you at re-uploading instead.One thing the chat co-pilot cannot do for you: it can choose the upload source, but it has no way to attach a file, so it will ask you to pick it.Try it in KAOP

Fixed: Shadow an agent you deploy yourself

Deploying a shadow of an agent you run yourself — your own image, your own infrastructure, connected through the SDK — now hands you what that actually takes: a worker token for the shadow, and the environment to start your candidate build beside the one already running.Before, it handed you a values.yaml and a helm install for our generic agent-base image. Running that started a different agent under your agent’s id. Both arms completed, both showed up as a comparison pair, and the shadow side was never your candidate — a comparison that looked right and was not. Only agents whose configuration AgentOps itself renders still get the chart command, because for those it is the right answer.The hand-off is also explicit that the candidate must run on the shadow’s own token. Deploying it with a copy of the main agent’s token is the quiet failure: it registers as an ordinary production replica and starts answering real callers, and nothing stops it. Your existing deployment needs no change.Remove shadow on the agent’s page ends the experiment and revokes the shadow’s token — the one thing that was previously impossible without an API call. Your candidate deployment is still yours to remove.Try it in KAOP

New: Pick which AgentOps tools your agent sees

An agent connecting to the AgentOps MCP endpoint received all 47 tools whether it used one of them or forty. Add ?toolset= to the URL and it sees only the groups it asked for — ?toolset=slack_bot,knowledge serves eleven tools instead of forty-seven, leaving the rest of the context window for the actual work.Name specific tools with ?tools=, combine both to union them, and send no parameter at all to keep the full set exactly as before. whoami always answers, so a narrowly scoped agent can still tell you which workspace it is in.Slack bot tools moved onto the same endpoint as the slack_bot group. Anything pointed at the old /slack-bot/mcp address should now use /internal/mcp?toolset=slack_bot.

New: Guardrails on the payload that starts a run, before the run exists

A guardrail can now stand at the entry. Pick Agent start as the gate on the first step and the rule is checked on the payload that would start a run — a webhook body, a schedule’s input, whatever an API caller sent — using the same conditions and operators every other gate already uses. With this the wizard covers five gates: what arrives, a tool call and what it returned, a request to a model and what the model answered.This is the only gate that can promise a value never entered your storage. The other four stand on traffic passing through us and refuse to pass it on; this one is asked as the run is created, before its row is written. So “no card numbers in anything that arrives” is a rule you can actually keep, rather than one that catches the number after it has been saved.The card is the conditions and nothing else — there is no key to name. A payload’s keys belong to whoever sent it: a Datadog webhook, a GitHub event and a schedule agree on nothing, so a rule naming one key would mean something different per source. The conditions read the whole payload, nested values included.Refusing means refusing. The run never starts, no run appears in your history, and the sender gets an error naming the rule that refused it. A webhook that retries has its retry refused the same way — the rule is deterministic and the second delivery is the first payload. Ask-me-first is not offered here, and for the opposite reason to the two inbound gates: there, the call already happened and there is nothing left to approve; here nothing has happened at all, so an approval would have no call to name.If you are not sure what a rule would catch, arm it as “Allow it, and record it” first — every trigger it matches is recorded and nothing is refused, so you can read the matches before the rule changes anything.Try it in KAOP
NewImprovedFixed

Improved: A closed co-pilot stays closed

Closing the wizard co-pilot is now remembered. Reload the page, or open a different wizard, and the panel stays shut until you bring it back.Until now the choice lasted only as long as the page: every wizard reopened with the co-pilot expanded, so anyone who prefers the form on its own had to close it again each time.The Co-pilot button that brings it back has moved to the bottom-right corner of the window, where a floating assistant is usually found. It used to sit in the top-right of the wizard, on top of that wizard’s own header.Try it in KAOP

New: See which resources a cluster declares, and when it last synced

Anything declared from a Kubernetes cluster — a guardrail, an MCP server, group or policy, a credential, a workflow, an incident workflow or a knowledge base — now carries a Cluster-managed badge, so you find out before you try to edit it rather than from the error afterwards. The badge names the cluster and where in it the declaration lives, and where the edit really would be refused the row’s own switch and delete button are disabled with the reason. It also distinguishes a cluster that is enforcing its manifests from one that is paused — there your edit sticks until the cluster resumes and reasserts what it declares.Adding a cluster now hands you its credential together with the two commands that consume it — the Secret to create and the Helm install, with your URL already in them, and every row has an Install action that shows those commands again later. Settings now has a Configuration clusters page listing every cluster running the config-plane operator: whether it is enforcing, its namespace and operator version, how many resources it declares, and when it last reported in. “Never” there is the first thing to check on an install that looks dead.Try it in KAOP

Improved: The review history puts the newest pass on top

A review’s history listed its passes oldest first, so on a PR that has been pushed to a dozen times the pass you came to look at — the one running now, or the one that just finished — was at the bottom of a list that grows with every push. Newest first now, with the latest marker on the top row where it belongs.Try it in KAOP

Improved: Notifications are one group

The Integrations rail groups what belongs together. Webhooks and Channels now sit under Notifications — everything that carries events out of AgentOps — while Integrations, MCP servers and Integration groups sit under Agent tools, which is everything an agent calls.“Notification sinks” is now just Webhooks. The old name described the row in the database rather than the thing you set up, and every one of them is a signed POST to an HTTPS endpoint.Channel routing is gone from the rail. It was a read-only view of routes you can already see — and edit — on a channel’s own page, so it is one place now instead of two that disagreed about which was authoritative.Inbound endpoints left the rail too. They receive events rather than send them, so they never belonged beside the outbound half; they are reachable from Webhooks and from the incident pipeline wizard while they find a permanent home closer to the agents they trigger.Try it in KAOP

New: Guardrails on what a model answers, before your agent reads it

A guardrail can now stand on the model’s answer. Pick Model response as the gate on the first step and the rule matches on the text a model produced — so “never tell my agent to drop a database” is one rule with one condition, using the same operators every other gate already uses. With this the wizard covers all four gates: a tool call and what it returned, a request to a model and what the model answered.This is the gate whose damage lands on a person rather than on a system. A model can answer with advice that harms whoever acts on it, or raise something your organisation has decided a model must not raise at all — and neither of those appears in the prompt you sent or in anything a tool returned, so no other gate can see it. That is also why deterministic conditions fit it well: a command template or a keyword is something you can name, explain, and get the same answer from twice.The card is the simplest of the four — the conditions and nothing else. There is no tool filter here, because an answer has nowhere to have come from but the model.Say what it does and does not do, because “refuse” means something weaker here: the answer has already been generated and billed by the time the rule is asked. What the rule withholds is the text, and the agent gets an error naming the rule instead of the content. Ask-me-first is not offered at this gate, for the same reason it is not offered on a tool result — there is nothing left to approve.One thing to know before you arm one: a streaming answer is refused outright. A streamed reply never exists as one whole text, so there is nothing for a condition to read — and rather than let it past unchecked, the request is refused before any tokens are spent, with an error naming your rule. If your agents ask for streaming answers, arm the rule as “Allow it, and record it” first: a recording rule never causes that refusal, so you can see what it would catch before it changes anything.Try it in KAOP

Fixed: Closing the co-pilot no longer leaves a button on top of the wizard

Closing the wizard co-pilot left its Co-pilot reopen button floating in the top-right corner of the page, on top of whatever the wizard’s own header had there. The button now takes its own space above the form and lines up with the content, so nothing it covers is lost behind it.Try it in KAOP

Improved: Import an OpenAPI document from the catalog

Add integration now offers Custom integration (OpenAPI) beside the bring-your-own MCP server. Point AgentOps at an OpenAPI document and its operations become tools your agents can call — no MCP server of your own required.The wizard is the one that shipped last week. What changed is where you find it: importing a document was reachable only from the Servers panel’s Add menu, which is somewhere you go to manage servers rather than to connect something new.Try it in KAOP

Improved: Change which keys an agent receives, after the credential exists

A multi-key credential’s per-agent key scope is now editable from the Access panel. Open Access on a credential, and every bound agent shows what it actually receives — All keys, or 2 of 3 keys — with an Edit keys control to narrow or widen it. Granting access to a new agent can carry a restriction from the start.Until now the checkboxes existed only in the create wizard. A scope set there was permanent in practice: the Access panel showed no restriction and had no way to change one, so the only route to a different key set was to remove the agent and grant it again — which handed over every key.That mattered most right after replacing a value. When a replace removes keys, AgentOps prunes them from every binding scoped to them, and deletes a binding left with none — deliberately, so an agent scoped to a key that no longer exists does not silently inherit access to new ones. The notice telling you to re-grant that agent now points at a panel that can re-create the narrow scope, and it names the agents rather than printing their ids.One quieter change with the same shape: re-granting a credential from anywhere else in the product — the agent wizard’s permissions step, an agent’s secrets card — no longer clears a key restriction. Widening a binding to every key is now something a screen has to ask for explicitly.Also: the Edit dialog says whether a credential holds one value or a named key set, the new-credential wizard reports an over-limit or malformed key set at the field instead of failing on submit, and a multi-key credential is no longer offered where only a single string can work — as an LLM key, or as an MCP server’s bearer token.Try it in KAOP

New: Answer what is waiting on you

Attention is one list of everything in the workspace that has stopped and needs a person, with the act on the row rather than three screens away:
  • A held tool call. A guardrail armed to hold refuses a call and registers it for approval. The row shows the tool and its arguments — the whole point of an approval — with Approve and Deny. A held call made inside a workflow step resumes on its own once you approve it.
  • A remediation selection. When a workflow parks on a choice, the options it froze when it asked are the options you pick from, “take no action” always among them. Answering advances the run down that branch immediately.
  • A run that gave up. Anything that stopped and needs review, linked back to the incident it belongs to so you can retry it there.
Incident rows and the incident page now carry Awaiting approval, Awaiting your selection and Needs attention badges, and the header bell counts all three rather than just approvals.Try it in KAOP

Improved: Creating an agent takes two fewer steps

Scaling and Notify me are no longer steps of their own. KEDA autoscaling is now set where you choose to self-host, and the notification sink is now the second half of the Triggers step — so the step count no longer changes under you when you switch between self-hosted and Komodor cloud.The last step is called Activate rather than Review: it drops the configuration JSON and the permissions recap that only restated what you had already filled in, leaving the one action that actually creates the agent. Capabilities now sit next to the instructions that decide what the agent does with them.Try it in KAOP
NewImproved

New: Jump from an incident straight to the alert that started it

An incident that came in from Datadog or PagerDuty now shows an Open in Datadog / Open in PagerDuty button on its detail page, linking to the monitor or incident on the provider’s side. The link is read from the alert payload itself — for Datadog, make sure your webhook template includes $LINK. When the payload carries no link, no button is shown.Try it in KAOP

Improved: Model-request guardrail decisions now show on the run

A run’s Guardrails tab used to list only the decisions made about its tool calls. Decisions made about what the agent sent to a model were recorded in the audit trail and shown nowhere near the run they belonged to, so the first place anyone looks when a run behaves oddly was the one place that did not mention them. They now appear alongside the tool-call ones.They come with a during this run mark, and the distinction is worth knowing. A tool-call decision is made inside a run-scoped session, so it names the run it stopped. A model-request decision is made in the LLM path, where the request is built by the agent’s own harness — nothing there can attach a run to it, so the tab matches it by the agent it names and the time it landed. That supports “this fired while your run was open” and not “this stopped your run”, and the mark is which of the two you are reading. A decision that names a different run is left out rather than shown unmarked.The audit list gained the filters behind it: you can narrow it to one agent, and to a time span, without writing a query by hand.Try it in KAOP

Improved: MCP setup picks the group for you to choose, and saves as you go

The last step of MCP setup used to be a blank box you typed a group name into. It is now a picker over the groups you already have — each one listing the servers it holds — with an explicit option to create a new one. Choosing an existing group names the agents already on it, so you can see who gains the tools you just exposed before you commit. Typing a name that is already taken, or one the group endpoint could not carry, is now said out loud instead of failing on save.Exposing tools no longer needs a save button: ticking a tool writes straight through, and the step says so. Unticking every tool now means every tool is withheld, which is what it looked like it meant.Finishing is one button. Done saves the group and closes; if the servers have no group yet it says so — with a Close anyway next to it, since an ungrouped server is registered and safe, just not reachable by any agent until you group it.Try it in KAOP

New: Integrations show the servers they feed

Integrations is one list now. A connection and the MCP servers that authenticate with it sit together — a Datadog connection feeding five servers is one row you expand, instead of one row under Connections and five unrelated rows under MCP servers. Custom servers you registered yourself are rows in the same list rather than a separate surface.Every row says who is using it. “Used by” resolves through the groups your servers belong to to the agents scoped to those groups, and names them rather than counting them. It is also honest about the case that used to be invisible: an agent with no MCP group can reach every server on the gateway, so a row the whole fleet can call now says so instead of reading as unused.A failing credential is called out above the list, with what broke, when, and how many servers are serving unauthenticated because of it. Reconnect appears on the row while it is actually failing, and nowhere else.The catalog moved behind Add integration — a searchable panel by category, with the servers you run yourself first. An account with nothing connected still sees it inline, since there is nothing else on the page to look at yet.Try it in KAOP

New: Guardrails on what your agents send to a model

A guardrail can now stand in front of a model request, not only a tool call. Pick where the rule stands on the first step: Tool call checks a tool before it runs, Model request checks what the agent sends to a model — the prompt, before the model sees it. A rule at the new gate matches on the messages or the system prompt, with the same operators tool-call rules already use, so keeping an internal codename or an account number out of a prompt is one condition rather than a code change.The form also authors the two criterion shapes it previously could only display. Model judgment states a policy in plain language — “no customer identifiers leave the account” — and names which of your models answers it, for the worry no pattern can describe; you choose how long it may hold the call open before an unanswered policy counts as unreachable. A rule carrying either shape now opens for editing instead of refusing, which is the more useful half of the change: until now a rule written through the API could be read in the console and not changed there.One thing the new gate does not get is the replay. The backtest works by re-reading the tool calls your agents already made, and nothing records what they sent to a model, so there is no history to test a prompt rule against — the review step says so rather than showing an empty result. Arm such a rule as “Allow it, and record it” to watch what it would catch on live traffic, then switch it to refuse.Try it in KAOP

New: Guardrails on what a tool returns, before your agent reads it

A guardrail can now stand on the way back. Pick Tool response as the gate on the first step and the rule matches on the text a tool returned — so “no private keys in anything a tool returns” is one rule with one condition, using the same operators tool-call rules already use. It reads every content block a tool sends and its structured JSON answer too, so a server replying in JSON instead of text does not slip past.The fields are the other way round at this gate, which is the point of having it. What comes back is what the rule is about, and the tool picker is demoted to a filter: Only from these tools, where leaving it empty means any tool. That is the reverse of a tool-call rule, where naming no tool covers nothing.Say what it does and does not do, because “refuse” means something weaker here: the tool has already run by the time this gate is asked, so its effects stand. What the rule withholds is the result, and the agent gets an error naming the rule instead of the content. Ask-me-first is not offered at this gate for the same reason — there is nothing left to approve. If a result arrives in a shape the gate cannot read at all, it is withheld rather than passed.There is no replay here either: nothing stores what your tools returned, and that is deliberate — a tool result is not kept in the audit trail either, so no decision can hand it back to you. Arm the rule as “Allow it, and record it” to watch what it would catch on live traffic, then switch it to refuse.Try it in KAOP

New: Connect any REST API from its OpenAPI document

Point AgentOps at a service’s OpenAPI document, tick the endpoints worth exposing, and those endpoints become tools your agents can attach. Nobody has to write or host an MCP server for that API first, so connecting an internal service or a vendor’s API stops being a development task.Add integration now offers Custom integration (OpenAPI) alongside the MCP server option. The wizard reads the document, lists everything it declares, and lets you filter and pick — reads start ticked, anything that writes starts off. Credentials bind exactly as they do for an MCP server, and the generated tools appear in the catalog under the integration’s name.Come back to the same wizard any time to change the selection, add a note to an operation, or re-point the credential. Refreshing re-reads the document and tells you which endpoints the upstream has stopped offering.Try it in KAOP

Improved: The agent wizard shows what an MCP group actually grants

Picking an MCP group for an agent used to tell you the group’s name and nothing about what the agent could do with it. The step now lists the group’s tools per member server, and marks the ones the agent will not get — separating a tool the server itself denies from one a group rule withholds, because they are fixed in different places.Each server folds away, and a server with a long tool list scrolls in place instead of pushing the rest of the wizard off screen.Try it in KAOP

Improved: A tool pattern that matches nothing now shows you what would

Writing a tool pattern by hand, the wizard already warned you when it matched none of your tools — but not what your tools actually look like, which is the thing you need to fix the typo. It now names a few, one per namespace rather than the first handful in catalog order, so a catalog full of k8s_* tools cannot hide the fact that the prefix you meant was datadog_.Try it in KAOP
NewImprovedFixed

Fixed: Tables and menus sit straight

Providers fits on screen again. A status pill like “Verification failed” used to wrap inside its own border, and the verified-at column carried a full timestamp that took two lines, so every row was half again as tall as it needed to be and the table ran past the bottom of the window. The pill stays on one line, the timestamp is the compact form the rest of our tables use — the full value is still there on hover — and nine models now fit where six did.Every row menu in the console had its icons pressed flush against their labels. They no longer do, and the menus follow the same spacing, cursor and label styling as the dropdowns beside them.Fleet’s “Last change” no longer reads as a dash for an agent you just created. Until the agent’s definition changes for the first time, the column shows when it was created, which is the only change there has been.Try it in KAOP

Fixed: Rows and filters look clickable

Hovering a review row or a sidebar filter on the Code Review inbox showed a plain arrow, as if the row were inert. Every inbox built from the same parts — Incidents and the Quality Lab included — now shows a pointer over its rows, its filters, and the collapsible filter groups.Try it in KAOP

New: Get notified when your agent’s runs finish

The agent creation and edit wizards now include a Notify me step. Choose any notification sink from your account and it will receive run lifecycle events — run started, run succeeded, run failed — for every run that agent starts.The step shows a searchable dropdown of all your configured sinks with their kind (Webhook, Slack, MS Teams), and once you select one it lists exactly which events that sink is subscribed to so there are no surprises. The step is optional: skipping it leaves the agent with no sink attached.Try it in KAOP

Fixed: The MCP group you pick now sticks to the agent

Build an agent from scratch, pick an MCP group, and the group is now recorded on the agent itself. It always reached the worker — the group is in the values.yaml you download — but the control plane’s own record of it stayed empty, so Fleet showed no group and nothing that reads an agent’s tool group could see one. The only screen that showed it correctly was the Edit wizard, which was reading back its own notes.Changing the group from the Edit wizard now takes effect too, instead of being remembered by the wizard and nowhere else. Setting an agent to use every gateway MCP server clears the group as you would expect.Picking a group that no longer exists is now refused when you deploy, with the group named, rather than accepted and quietly ignored.Try it in KAOP

Fixed: An imported Claude Agent SDK worker now shows its trace, not just its cost

Importing an existing Claude Agent SDK agent through the wizard produced a worker that registered, connected, and reported real usage and cost — but its runs’ Trace Timeline stayed empty. No tool calls, no message content, nothing to show what the agent actually did, even on a run that succeeded.The generated wrapper now merges AgentOps’ hooks into the agent’s own ClaudeAgentOptions before calling query(), the same two-line change the “importing an existing agent” skill now documents. Nothing about the agent’s own loop, prompt, or options changes — the hooks only observe. A freshly imported Claude Agent SDK agent’s tool calls and message content now show up in its Trace Timeline like any other agent’s.Try it in KAOP

Fixed: A slow deploy is no longer a failed one

An agent that took a few minutes longer than usual to roll out would report “Deploy timed out — the agent did not come online within the expected window” and keep reporting it forever, even while the agent was up and serving runs. The deploy really had missed its window; what was wrong is that nothing ever revisited the verdict once the agent arrived.Two changes. The window a rollout is given now covers the whole chain — the build, the sync, the image pull and the pod starting — rather than expiring while the pod is still being pulled. And when a deploy does time out, it is no longer the last word: as soon as the agent comes online, the deploy settles as done on its own and the banner clears.A deploy that failed for a reason the agent cannot answer — the change never merged, its credentials were never provisioned — still reads as failed, which is what you want it to say.Try it in KAOP

Improved: A guardrail reads in one glance

The guardrails list gave its widest column to a sentence describing each rule, which is fine for one rule and unreadable for a page of them. A row now answers the question in three words instead: the gate it stands at, the kind of matcher it uses, and whether a match is refused or only recorded. The patterns and the description are still there, on the rule itself.Guardrails also shows how many you have in the Govern menu, like every other section of the catalog does.Try it in KAOP
NewImprovedFixed

New: Outposts say whether they still run the config AgentOps rendered

An Outpost relay reports a digest of the configuration document it loaded, but nothing compared it against anything: drift was recorded on every connect and detected by nobody. An Outpost could be enrolled, connected, and quietly serving traffic through an allowlist nobody in AgentOps had written.Every Outpost now carries a config verdict in the list and on its detail page, beside its enrolment and connection badges. AgentOps records the digest of each configuration it renders install values for, and compares it against what the relay says it is running.The verdict separates the two cases that matter, because they call for opposite responses. Upgrade pending means the relay runs a configuration AgentOps did render, just not the newest one — ordinary after any change on our side, and fixed by running the upgrade command on the Install tab. Config differs means it runs a configuration AgentOps never rendered: edited in the cluster, or overridden by a chart value. That one is worth a look.Two more verdicts exist so the badge never overstates what it knows. Config mixed appears when replicas of one relay disagree with each other, which is normal mid-rollout; a fleet part-way through an upgrade no longer reports as though every replica agreed. Config unknown means there is nothing to compare yet — nothing live is reporting, or install values have never been rendered.Turning the local listener on does not move the verdict. It changes how the relay is deployed, not what it may route or reach, and a badge that cried drift over it would be one you learned to ignore.Try it in KAOP

New: Incident and run names are short summaries now

New workflow incidents are named {workflow} — {short summary} instead of the raw monitor title from Datadog or AlertManager, so the inbox is scannable. Agent runs in Fleet and History get the same kind of five-word summary as soon as the run is queued. Chat sessions are unchanged — they already use your first message. Existing rows are not rewritten.Try it in KAOP

Improved: Connection checks now work for Datadog, GitHub, AWS and more

When connections started reporting whether their credentials actually work, two providers could never report a pass — including Datadog, the most connected provider. The check was talking to their tool servers without opening a session first, so those servers refused it and the connection showed a problem that was ours, not yours. GitHub had a second version of the same fault: the check called a tool that server does not offer.Both are fixed, and the check now covers AWS, Buildkite and Grafana Cloud too. Between them and Slack and Google Workspace, every provider anyone has connected can now say whether its credentials are accepted — where before, most connections had never been checked at all.Each of these was verified against a real account, twice: once with working credentials to confirm it passes, and once with deliberately broken ones to confirm it fails. That second check is the one that matters, and it changed two of them. Buildkite’s obvious endpoint needed a permission many tokens do not have, so a working connection would have been reported as broken; GitHub’s needed an endpoint that suits an app rather than a person.Providers nobody has connected yet still say they cannot be verified, which is the honest answer rather than a green tick nobody earned.Try it in KAOP

Improved: Agents are named, not numbered

Screens that referred to an agent by its internal id — agt_9f3c2d1e8b4a — now use the agent’s name. The Workflow details drawer was the worst of them: the orchestrator, every specialist and the remediation agent were all listed as ids. Guardrail decisions, a golden run’s evaluator, the agents bound to a credential and the run history on Cost and Overview have all changed too.The ids have not gone away, because they are what you paste into an API call or a support thread. Hover any agent name to see its id and its handle, with a button to copy the id.Two agents can share a display name, since that name comes from the worker rather than from you. Where that happens on one screen, both are shown with their handle so you can tell them apart.Try it in KAOP

Fixed: The agent edit dialog no longer offers to rename an agent

Editing an agent used to show Name and Description boxes you could type into. Nothing you typed was ever saved — the dialog reported success and the old values came back a few seconds later.Those two come from the agent itself, published by its worker, and they are refreshed every time the agent checks in. The dialog now shows them as read-only and says where they come from. To change them, update the agent’s card where the agent is defined and redeploy it.Turning automated triggers on and off, which is the part that always worked, is unchanged.Try it in KAOP
Fixed

Fixed: Pull requests are reviewed whatever branch they target

A pull request opened against another topic branch rather than the default one — the middle of a stack, master → branch-1 → branch-2 — was silently skipped: no review, no run, no commit status, and no comment on the pull request. A review workflow only accepted pull requests targeting a single branch, and when its base branch filter was left empty it fell back to the repository’s default branch instead of accepting every branch.An empty filter now means what the wizard always said it did: every base branch is reviewed. Naming a branch in the filter still restricts reviews to pull requests targeting exactly that branch, and a filter that was set can now be cleared again.Re-pointing a pull request at another branch also triggers a fresh review. GitHub does that automatically the moment the branch under a stacked pull request merges, and the review that follows describes the changes the pull request actually contains now rather than the ones it carried against its old base.Try it in KAOP

Fixed: A connected MCP server’s tools reach your agents right away

Connecting an MCP server and attaching its credential is now all it takes — its tools show up on the gateway for every agent within seconds, with nothing to press and no wait for a restart. Previously a server could pass its connection test, list every tool in Manage, and still be invisible to your agents.Manage → Refresh gateway tools republishes one server’s tools on demand, for when you have changed something on the provider’s side and want the gateway to pick it up immediately.Try it in KAOP

Fixed: Incident findings are no longer cut off mid-sentence

The Findings card on an incident used to stop around two thousand characters, often in the middle of a heading or a sentence, because we stored only a short excerpt of the orchestrator’s write-up. The page now keeps the full report. Slack still posts a short preview with a link back to the investigation.Try it in KAOP
NewImprovedFixed

New: Filter which alerts start an investigation, and see what happened to the rest

An incident workflow used to investigate every alert its endpoint received. Narrowing that was a Datadog-side monitor change, or a second endpoint — and if an alert never became an incident, there was nothing to look at to find out why.The workflow’s Endpoint step now has a filter: rows of severity, title, status, labels.<key> or raw.<any.payload.path> compared with is / is not / contains / excludes / at least / is any of, and so on. All of them have to hold, and no conditions still means every alert.contains is the one that matters for real payloads. Senders routinely bury the value you want inside a human-readable string rather than giving it a field — a Datadog webhook on the stock template arrives with its priority in the title as [P2] ..., and a PagerDuty orchestration event carries the alert name in event.data.incident.summary. raw.title contains [P2] and raw.event.data.incident.summary contains [FIRING:1] are now expressible; before, the closest you could get was pasting a whole per-incident string that would match exactly one alert.Severity comparisons work across the ways providers spell it. severity at least SEV-2 ranks correctly whether the payload says P2, critical or high, and severity is SEV-4 now matches PagerDuty’s Sev-4 and Datadog’s P4 — the value shown in the test panel is the value that matches.Every workflow also gained an Events feed showing each alert its endpoint delivered and what the workflow did with it — routed (with a link to the incident), filtered out (naming the condition that rejected it), throttled, or deliberately ignored. Integrations → Endpoints → History shows the same thing from the endpoint’s side, for every workflow listening on it. An alert that was dropped on purpose no longer looks the same as a webhook that never arrived.If one endpoint feeds workflows that read its payloads differently, a workflow can now override the field mapping — the dedup key, title, severity and transition paths — and preview it against a real captured event before activating, so a wrong path shows up as a visibly empty title instead of silently mis-parsed incidents.All of it is available to the co-pilot and to aops workflows filter set|test, using the same <field> <operator> <value> phrasing.Try it in KAOP

Improved: The create wizard now waits for your agent, instead of telling you to go and look

Creating an agent used to end on a promise: “your agent will appear in Fleet once it reports its first heartbeat”. True, but unactionable — and indistinguishable from a deploy that quietly failed.The last step now watches for that heartbeat and tells you which of the two happened. It fills in on its own as the agent comes up, and once it connects it shows you the card the agent published about itself — its name, model, version, capabilities and skills — so you can see it started with the configuration you deployed. You can leave the page and come back; the agent is already registered either way. If nothing arrives after a few minutes you get a short list of things worth checking, not a spinner and not a failure: an agent that has not connected yet is an install you have not finished, and it waits as long as you need.Komodor-deployed agents no longer jump you to the agent page the instant the deploy is accepted — they stay put and report progress, so you can watch the thing actually start before moving on.For agents you run yourself, the two artifacts now arrive in the order you have to use them: values.yaml first, then the helm install command that reads it. The command references that file, so a screen that offered it first was asking you to run something that could only fail.Try it in KAOP

Improved: Test an MCP connection before the setup wizard registers it

Setting up an MCP server from a connected account used to mean registering it first and finding out afterwards whether the address was right and the credentials worked. If they didn’t, the wizard had already created the server and had to go back and delete it.There is now a Test connection button beside Register. It reads what the address advertises before anything is created, and it says which question it actually answered: with an integration connection selected the test sends that connection’s credentials, so a clean result means those credentials were accepted — not merely that something responded. Where several servers are registered in one click, each endpoint is tested and reported on its own line. A result disappears the moment you edit the URL, switch service or change connection, so what you are reading always describes what is on screen.Testing is advisory, never a gate: a server reachable only through an Outpost cannot be probed before it exists, and those still need to be registerable.Signing in to a provider got two fixes on the way through. The login window is now opened by your click rather than a moment later, which is what browsers’ pop-up blockers were stopping. And once you finish signing in, the connection is re-tested and the result shown in place — with a little tolerance for a token that has just been written and has not settled yet, which used to surface as a freshly authorized server appearing broken.Try it in KAOP

Fixed: The Providers page tells you when a provider refused

Adding models to a provider used to look the same whether the provider had nothing to offer or had rejected the key outright, and the page would show a “Known models” badge over an empty list. It now states the provider’s own reason where the list would have been, and keeps the hand-typed model id available.Model ids are also offered exactly as the provider spells them, so ids that carry their own provider prefix (groq/compound) register and verify instead of failing with a not-found that looked like a permissions problem. Models you tick are tied to the credential they came from, so switching credentials starts the selection over rather than saving the previous provider’s picks.Try it in KAOP

Fixed: A server whose credential was deleted now says so, instead of “permission denied

Disconnecting an integration destroys its credential, but an MCP server bound to that connection kept its binding. The server’s status then read permission denied — the wording for a credential the upstream refused — which sent people looking for a missing scope on a credential that no longer existed.Those servers now report credential missing, and testing the connection says plainly that the credential is gone and that the fix is to reconnect the integration or pick another credential under Authentication. The same wording appears whether the status came from a connection test, a startup check or live agent traffic, so the row and the detail under it agree.Choosing a credential that has since been disconnected while adding a server is now caught the same way, before the server is saved.Try it in KAOP

New: Match a tool argument with a regular expression

A guardrail’s Only when conditions could compare an argument four ways: equals, is one of, contains, and a glob. That covers a value you can name, and misses the rules people kept asking for — any destructive statement, any role outside our own account — because those are shapes, not values. The workaround was contains, one condition per verb, and a rule broader than it read.Conditions now take a fifth operator, matches regex. The pattern is found anywhere in the argument, so (?i)\b(drop|truncate)\s+table\b is one condition where four contains rules used to be an approximation. It is case-sensitive unless the pattern starts with (?i), and unanchored unless you write ^…$ — the same expression you would test anywhere else. Back it with Try it before arming it, which replays the rule over real traffic and is worth more here than anywhere: a regex is the one condition you cannot check by reading it twice.Two guarantees behind it. A pattern that does not parse is refused when you save it, not quietly stored as a rule that never fires. And a pattern is given a fixed budget of time on each call, so a badly-behaved expression cannot slow down the agents the rule was written to protect.Try it in KAOP

Fixed: Disconnecting an integration now asks first — and says what it takes down

Disconnect on a connection used to be a single click with no way back. It destroys the stored credential and removes the connection outright, and the MCP servers authenticating with that connection did not stop working in any way you could see: the gateway carried on connecting to them without credentials, so the first sign was a tool call failing somewhere else with what looked like a permissions problem.Disconnect now opens a confirmation that names the MCP servers bound to that connection, says they will keep running but unauthenticated, and states plainly that the credential is destroyed and that reconnecting means authorizing the provider again. Cancel leaves everything exactly as it was. The dialog also says what it cannot see — Slack channel routing and agent credential bindings using the same connection are not listed yet — so the list it does show can be trusted as far as it goes.Try it in KAOP

Fixed: Connections now tell you whether the credentials actually work

A connection went green the moment you saved a credential, whether or not the provider would accept it. Paste a Datadog API key with a typo in it and the row read “active” — and pressing “Test connection” agreed, because the test only checked that the fields were filled in on our side. Nothing had ever tried the credential against the provider.Now something does. Where we can verify a provider’s credentials, we do: a key the provider rejects is refused at the point you save it, with the provider’s own reason, instead of becoming a connection that looks fine and fails later. Testing an existing connection makes a real call too, and the result is stored on the row rather than vanishing with the toast.Connections show this next to the existing status, because the two answer different questions — one is whether the connection was set up, the other is whether its credential works today.Where we cannot verify a provider yet, the row says so plainly rather than showing a green tick nobody earned. That is deliberate: telling you a credential is unproven is honest, and telling you it works when nothing checked is how this went wrong in the first place. More providers will move from unproven to verified as we confirm each one against a live account.Try it in KAOP
NewImproved

New: A run now shows which guardrails it tripped

Finding out why a guardrail refused a tool call meant going to Settings → Audit, filtering to the guardrail category, and matching rows to a run by their timestamps — which is guesswork as soon as two runs overlap.Guardrail decisions are now recorded against the run that caused them, and a run with any decisions gains a Guardrails tab showing what was refused or recorded, which rule decided it, and which tool it was about. The count is on the tab, so you can see whether a run tripped anything without opening it.Nothing changes about what guardrails do — this is the trail catching up to the enforcement.Try it in KAOP

Improved: Upgrading an Outpost no longer discards your own Helm values

An Outpost’s install panel now gives you a separate upgrade command with its own values document, and the upgrade merges into the release instead of replacing it. Settings AgentOps does not render — a CA bundle, the local listener, resources, image pull secrets — now survive an upgrade rather than being dropped, which previously took the relay off the air whenever the lost value was the CA in front of AgentOps. Needs Helm 3.14 or newer.Try it in KAOP

Improved: Connected MCP servers now report their own health, and recover on their own

A connected server’s status used to describe the last time someone pressed “Test connection”. A token that expired the next day, a scope someone revoked, an upstream that went down — none of it showed until the next manual test.Real traffic now keeps that status honest. When an agent’s tool call is refused because the credential no longer works, or the server cannot be reached, the server says so on its own; when a later call succeeds, it clears itself, so a server that recovers needs nothing from you. Each row shows when its status was last observed and what observed it, so a green tick you can trust is distinguishable from one nobody has confirmed in a week.Two things this fixes along the way: the status badge now appears for every connected server rather than only ones using OAuth sign-in, and a server whose only safely callable tool is something like a version endpoint no longer reports a clean bill of health — calling it proved the server was up, never that your credentials work, and it now says so and points you at setting a validation tool.Try it in KAOP

New: Grafana Investigator can connect to Grafana Cloud’s own hosted MCP server

Until now, a Komodor-cloud deploy of Grafana Investigator could only reach Grafana through its bundled MCP tools — reliable, but not a fit if you’d rather point at Grafana Cloud’s own hosted MCP endpoint instead.Deploying to Komodor cloud now offers an MCP tools step where you can bind a gateway-connected MCP server — register Grafana Cloud’s hosted server once, then pick it here. Self-hosted deploys are unchanged: the bundled Grafana MCP and the “use your own in-cluster MCP” option still work exactly as before.Try it in KAOP

New: See your whole fleet at once on the new Analytics tab

Fleet answered “what do I have” but not “how is it doing”. Analytics now leads the tabs, ahead of Agents, Triggers and Workflows: outcomes and reliability, cost, mean latency and versions deployed across the window, then runs, cost and latency plotted day by day, a worst-error-rate and a slowest-agents board ranking five agents each (with the run count beside every error rate, so you can see how much evidence a rate rests on), and the model mix behind it all as a donut whose legend names every slice and what it cost. A Today / Last 7 days / Last 30 days selector sits in the Activity tier’s header and governs that tier.Above all of that is Live status, which is about right now rather than the window and says so: fleet composition (online, offline, draft), replica capacity, queued work and its oldest item, channel reachability and the SDK versions your workers are running. Every composition bucket links straight into the agents list already filtered, so a count you want to act on is one click from the agents behind it.The agents list gained a matching Activity (30d) column: runs and cost over error rate and latency for each agent, sortable by any of the four from the column’s header menu. Narrow the fleet with the label, status or replica filters and Analytics follows, so the numbers describe the same agents the rail has selected.Try it in KAOP
NewImprovedFixed

Improved: Run the Remediation agent in your own cluster

Remediation can now be deployed self-hosted, not only to Komodor cloud. Pick it from the worker catalog and the deploy step offers both homes; nothing else about the agent changes.It is the same agent either way. Remediation is plan-only — it reads a finished investigation and returns a decision, and it has no shell, no cluster access and no credentials for the systems it reasons about — so a self-hosted deploy is not a limited version of the hosted one. What moves is where its prompt runs and where its model credential lives: your cluster instead of ours.The install is the usual one: AgentOps mints the worker token, renders a values.yaml for the agentops-agent-base chart, and gives you the helm install to run. No cluster RBAC is granted — this agent reads nothing from Kubernetes.Try it in KAOP

Improved: Testing an MCP connection now proves the credentials actually work

Testing a connected MCP server used to check only that we could reach it and list its tools. A token could pass that check and still be refused the first time an agent used it — because it was valid but missing a scope or a permission.The test now runs a second check: it invokes one read-only tool for real and reports whether that succeeded. The two results are shown separately, so a server that is reachable but not yet permitted to do anything is obvious at a glance, and a permissions problem no longer looks like a connection problem. The tool used is picked automatically, and you can name a specific one per server when the automatic choice is not the one you want.Try it in KAOP

New: Give an agent more than one trigger

The Triggers step asked one question — schedule, webhook, or neither — and took one answer. An agent that should run every morning and on an inbound event had to be created with one of those, then finished off by hand somewhere else.The step now holds a list. + Add trigger appends a card, each card has its own type and fields, and they are created together with the agent in a single call. Deploy an agent with two schedules and a webhook and all three exist the moment it appears in Fleet — each reported individually, so a schedule that was created never reads the same as one that wasn’t.Two limits are the agent’s, not the form’s, and the form now says so up front rather than letting you find out at deploy: an agent may hold one webhook, so that option greys out in the other cards once one has it, and Manual / on-demand means no automated trigger, so it cannot sit alongside others. Remove the card holding either and it becomes available again immediately — nothing here is one-way. There is always at least one card, because zero triggers is what Manual already says.Editing an agent opens the step on the triggers it actually has, which is the part that used to go wrong. Switching a schedule to a webhook once left the schedule firing and quietly added the webhook beside it; the wizard warned you and pointed at a cleanup screen. Now removing a trigger removes it, changing its type replaces it, and both wait for Save — the step tells you what is about to be deleted before anything is. Triggers an agent declares in its own agent-spec.yaml show up read-only, since that file owns them and the next heartbeat would overwrite an edit made here.Try it in KAOP

Fixed: Deploying an agent no longer looks like it did nothing

Pressing Deploy to Komodor cloud on the wizard’s Review step took you to your new agent — except when it didn’t. Anything that went wrong after the agent had already been created was swallowed, and all you saw was the configuration summary quietly disappearing from a page that otherwise looked untouched, Deploy button and all. It read as “nothing happened”, so the natural next move was to press it again.The agent’s existence is now recorded on screen the instant it is created, before anything else is attempted. You still land on the agent’s page as before; if you don’t, you are left looking at Agent deployed and a link to it, rather than at a Review step that appears never to have been used. The Deploy button cannot come back either way, so a retry can no longer deploy a second agent.Try it in KAOP

New: Build a code reviewer without writing a worker

Adding a custom reviewer meant writing and hosting an SDK worker. That is still the right answer for a reviewer that needs its own GitHub plumbing or incremental re-review — but it was the only answer, even for a reviewer whose whole job was “read this diff against our conventions and say what you find”.An agent built in the Add agent wizard can now declare Code review beside Chat. It runs the shared agent image, so its instructions are the only thing you write: give it an MCP group with GitHub, tell it what to look for, and it appears under Reviewers with a Custom badge, ready to add to any workflow next to the built-in reviewer.Worth knowing: the toggle makes the agent selectable, and nothing more. How good the review is comes entirely from its instructions, and a reviewer that returns no structured verdict shows up as a run link with no findings breakdown rather than an error. Both are also declarable over the API and the agent-management MCP, so an agent onboarding another agent can set them too.Try it in KAOP

Fixed: A label typed in capitals now works

Labels are matched by exact equality — everywhere they are used, a filter, an access grant, or a workflow looking for a particular agent compares the label character for character. A label saved as ROLE: Scenario-Runner therefore matched nothing that looked for role: scenario-runner, and nothing said so: the save succeeded, and the Labels list shows keys in capitals regardless, so a broken label looked exactly like a working one.Labels are now lowercased when they’re saved, keys and values both. Type ROLE and you get role — no error, no correction to make by hand — and the label meets the grants and filters written against it. Existing agent labels have been converted, so one that never matched before starts matching now.The visible change is in values: an agent labelled owner: AgentOps now reads owner: agentops. Keys look the same as they always did, since the list displays them in capitals either way. If two labels on the same agent differed only in case, they were always two separate labels, and they now collapse into one.Try it in KAOP
NewImprovedFixed

New: Try a guardrail before you arm it

A guardrail had one outcome: refuse the call. That left one question unanswered — is this pattern actually right? — and only one way to find out, which was to arm it on live traffic and see whose run broke.A guardrail can now be set to allow the call and record it instead. Same rule, same matching, no refusal: the call goes through untouched and shows up under Recent decisions with the agent and tool that triggered it. Read a day of that, then switch the same rule to refusing once you can see what it catches. It is also the honest way to run a rule you only want to watch — a scope boundary you would like reported, not enforced.Two things worth knowing. A call covered by both a recording rule and a refusing one is still refused, so adding a watcher can never quietly disarm a control you already rely on. And a recorded call is invisible to the agent — told it had been flagged, an agent would start working around whatever it was, and the rule would stop measuring the thing you armed it to measure.Try it in KAOP

New: See what a guardrail would have stopped, before you turn it on

The form could already tell you which of your tools a pattern matches. It could not tell you the thing you actually want to know: would this rule have refused work my agents were really doing? So the only way to find out was to arm the rule and wait.The review step now has a Check the last 7 days button. It replays your unsaved rule — the same matching the gateway uses, argument conditions included — against the tool calls your agents have already made, and answers with the count and the agents and tools behind it: “would have refused 3 of 240 tool calls in the last 7 days”.It only counts calls a guardrail could ever have decided, so a worker’s own shell commands don’t pad the total, and it never shows you a call’s arguments — those can carry credentials, and an authoring screen is no place for them.Try it in KAOP

Improved: See which agents have a shadow running, without opening them

A shadow deploy runs every request twice — the candidate version alongside the live one — for as long as it stays up. Until now the only way to notice one was already running was to open that agent’s Instances tab.The Fleet list now shows it directly: an agent with a live shadow splits its replica count into N production · M shadow right in the row, and a new Shadow active filter in the left rail narrows the whole fleet down to agents currently running one. Works the same way for both self-hosted and cloud-hosted shadows.Try it in KAOP

New: Scope an agent’s roles while you build it

Every agent has always started with the built-in agent role — the full set of control-plane and MCP capabilities a worker needs. Narrowing that meant creating the agent first, then finding it on the Access page and editing its roles separately, with a live agent running on the default set in between.The Agent Builder wizard now has a Permissions step, right before Review, on every flow — scratch, catalog, self-hosted and cloud. Leave it alone and the agent gets the same built-in role it always has. Choose Custom and pick a narrower set of roles instead, and the agent is scoped to exactly that from the moment it’s created — no gap where it runs wide open.The same step also lets you bind the credentials the agent will need, so an operator scoping access and an operator wiring up secrets aren’t two separate trips through Settings.Custom roles replace the built-in set rather than add to it — the built-in role already covers everything a worker needs, so adding to it wouldn’t narrow anything. Assigning roles is an admin-level action; the step only appears for an operator who holds it.Try it in KAOP

Improved: Catalog agents get a Model step and a Scaling step

Deploying a catalog worker (AWS Investigator, Grafana Investigator, and the rest) now walks the same shape of wizard as building an agent from scratch. Any catalog worker with its own model credential gets a real Model step to pick Opus, Sonnet, or Haiku — previously no catalog worker ever offered one. Installing a catalog worker into your own cluster now offers the same Scaling step (KEDA-based autoscaling) a self-hosted custom agent already had.The AWS Investigator’s IRSA option (letting a self-hosted pod inherit AWS access from its own IAM role, with no credentials stored in AgentOps) is now shown alongside the normal connection picker rather than replacing it, and is clearly disabled — not hidden — when you’ve chosen to deploy to Komodor cloud instead of your own cluster, since IRSA only works self-hosted.The Model step now also picks up the same provider/model picker the from-scratch wizard uses, and is the one place you set how the agent authenticates — no more seeing the same credential asked for twice, once on Model and again on Integrations. “Bring your own key” and “use an existing secret” are hidden for catalog workers, since a pinned catalog image can only authenticate through a named credential reference. Integrations now comes before MCP tools in the step order, matching the order the rest of the wizard reasons about a worker’s setup.The Integrations step no longer shows up empty for a worker that has nothing to configure there — the Orchestrator, which authenticates entirely on the Model step and reaches other workers directly rather than through a bound connection, now skips straight from Model to Review (or Workers, if you’ve picked which agents it may delegate to). It also no longer offers a wizard-created schedule or webhook, since it’s meant to be invoked over chat or ad-hoc runs instead.Grafana Investigator can now run on Komodor cloud too, not just in your own cluster — pick “Where it runs” like any other dual-capable worker. “Use your own in-cluster MCP server” stays disabled (with an explanation) when you choose Komodor cloud, since there’s no cluster of yours for it to reach from there. Grafana, Datadog, and Kubernetes RCA Investigators can now be chatted with directly, joining the Orchestrator and Klaudia Investigator — and now so can the AWS Investigator (chat support lives in each worker’s own image, not a wizard setting).The Card step’s Capabilities section (Chat, Code review) is scratch-agent-only: a catalog worker’s capabilities are baked into its pinned image, not chosen in this wizard, so the whole card is hidden when deploying a catalog worker rather than showing a toggle with nothing behind it.Try it in KAOP

Fixed: A tool call that failed no longer reads as still running

A run transcript closed a tool call only when its result arrived, and a call that failed never recorded one — so it kept a spinner and read as still running, in a run that had long since finished. The Trace tab said error at the same time. The first place this showed up was a guardrail refusing a write: the agent narrated the refusal in plain words, and the tool row above it claimed the call was in flight.A failed call now records its result like any other, carrying the error, and the row shows Failed with the reason. And because a result can also go missing for reasons nothing can record — a worker killed mid-call — an unpaired call in a finished run now reads No result instead of spinning: not a success, not a failure, just the honest statement that nothing came back.Try it in KAOP
NewImprovedFixed

New: Try a change on a cloud agent without changing it

Shadow deploy now works for agents we host for you, not just self-hosted ones. Edit a cloud agent as usual, and on the last step choose Deploy as a shadow instead of updating it: AgentOps deploys your edited version beside the live one and sends it a copy of every run — while callers keep getting only the live agent’s answer.Nothing about the running agent changes. It keeps its own prompt, its own worker token and its own traffic; the shadow gets a fresh token and its own deployment, both provisioned for you. There is no command to run.Each request then produces a comparison pair — the candidate’s run badged Shadow next to production’s, linked both ways, with a side-by-side view in the run console.A shadow may differ from the live agent in its prompt, model, skills, image or replica count — the things that live in its own config. It shares the live agent’s credentials, tool group and triggers, so those stay as they are.When you have seen enough: apply the change for real by editing the agent again and choosing Update the main agent, then ask the assistant to remove the shadow (or call DELETE /api/v1/hosted-agents/{customer}/{agent_id}-shadow). Archiving the main agent pauses both.Try it in KAOP

Improved: Storing a key you already have no longer makes a second copy

AgentOps now recognises a secret it already holds, whatever you call it. Adding a credential whose value your account already stores no longer creates a duplicate — you get the existing credential back, and nothing is written.That used to be impossible to catch. The only check was on the name, so the same API key stored as datadog-prod and dd-key looked like two unrelated secrets. Rotating one of them quietly left the other live and still handed to whichever agents were bound to it.Two things you will notice. Adding a value that is already stored says so, and names the credential that holds it — which may not be the name you typed, so that is the one to pick in the agent wizard. And a name clash now means something sharper: the name is taken by a credential holding a different value, which is worth a look rather than a shrug.Local credential discovery gets the same treatment: running it a second time creates nothing, even when a teammate already stored that key under another name.Deliberately unchanged: replacing a credential’s value still stores exactly what you give it, since rotating a key is not the same as merging two credentials that agents depend on separately.Try it in KAOP

New: Guardrails — stop an agent from taking an action, before it takes it

An agent that can call a tool can call it in production. Until now the only ways to prevent that were to withhold the tool from the agent entirely, or to find out afterwards from the audit trail. Guardrails are the rule in between: name some tools, optionally name which agents and under which arguments, and a matching call is refused at the gateway before it runs. The agent is told which rule refused it, by name, so it can escalate to a human instead of retrying blindly.The rule most people want is not “never delete pods” — it’s “not in production”. So a guardrail can condition on the call’s arguments: k8s_delete_pod stays available for staging while the production call is refused. Conditions compare on JSON text, so 0 and "0" mean the same thing to a rule, and a call that omits the argument entirely is not covered.Authoring is a picker, not a syntax. Guardrails lists your real tools, grouped by server, so nothing has to be typed from memory — and if you do write a glob by hand, it tells you how many of your tools it currently matches. A pattern that matches nothing is the mistake worth catching: it produces a rule that looks active and protects nothing.Or describe the rule and let the co-pilot draft it. The form sits beside a conversation: say “stop my agents deleting anything in production” and it fills in the tools, the targeting and the conditions from your account’s real catalog, for you to check and save. It never submits — the rule exists when you press the button.Targeting one particular agent is a pick from a list too, by the name you gave it. The rule stores the internal id underneath, because that is what the gateway checks a call against — but nobody should have to know that, or type it.Every block is audited — which guardrail, which tool, which agent — and the page shows recent blocks, so “did my rule actually fire?” has an answer on the same screen where you wrote it.Managing guardrails needs the guardrail.manage capability; it is an admin verb, because a guardrail is a control over other people’s agents.Try it in KAOP

Improved: Editing a self-hosted agent now tells you what each change will actually do

The Edit wizard could tell you a field wasn’t editable, but not much else. It now explains itself at each step, and the final screen says what your edit will do before you commit to it.The Review step lists your changes, grouped by when they take effect. Some of what this wizard saves is live the moment you continue — a schedule or webhook is a Komodor-side resource and starts firing straight away. The rest only reaches your agent when you run the generated helm upgrade in your own cluster. The button now reads Save and generate command, because that is the step that writes; the command screen after it is just the part we cannot run for you.Switching a trigger to Manual does not delete it. This wizard only creates and updates triggers, so an existing schedule keeps firing after you switch away from it. The step now says so, and points you at Automations to remove it.Agents built from your own image are recognised more reliably. An agent whose agent-spec.yaml doesn’t set a version was previously reported as one we couldn’t classify, so the wizard declined to say anything about it — even though its repository and source path were right there on its card. Most agents don’t set a version, so this affected a lot of them. They’re now identified from where their spec came from.Instructions you can’t edit are still shown. When your agent’s prompt lives in a baked agent-spec.yaml, the field used to be locked and empty. It now shows the instructions the running agent is reporting, so you can see what it’s actually doing while editing it stays a repository change.Smaller things: the MCP step’s controls are genuinely disabled when they can’t be changed rather than only looking that way; the agent id field no longer shows a validation error for an identifier you can’t edit; and the two overlapping “we don’t know much about this agent” notices are now one.Try it in KAOP

Fixed: Editing an agent no longer drops the MCP servers you installed yourself

An agent you created over the API — rather than through the wizard — could be edited here in two ways that quietly did the wrong thing.Your own MCP servers survive an edit now. If you picked the group MCP mode, the generated helm upgrade wrote the AgentOps gateway into agent.mcpServers — the same list your original helm install used for your own servers. Helm replaces lists rather than merging them, so that upgrade removed every server you had configured there, and nothing in the wizard said so.AgentOps never sees those servers’ definitions. Their transport, URLs, headers and environment stay in your cluster deliberately — a free-form environment block would otherwise copy your secrets into our database — so the wizard cannot list them back to you, and it should never have written a list it could not see. It now writes only its own entry, under agent.managedMcpServers, and the chart merges the two. Your list is left alone.This needs the agentops-agent-base chart at 0.1.5 or newer, which the generated command already points at.A bring-your-own Anthropic key can no longer be skipped. Opening a BYO agent showed an empty key field with Next still enabled, so you could walk through to Review and produce an upgrade command carrying no key. The field is empty because the key goes straight into your cluster’s Secret and AgentOps keeps no copy of it — it was never lost. The wizard now says exactly that, waits for you to type it in, and puts what you typed into the command instead of discarding it.An agent that never recorded how it authenticates is left alone rather than being asked for a key it may not use.Try it in KAOP

New: Create an agent and its trigger in one call

POST /api/v1/agents now takes triggers and autoscaling, so an agent can be created fully configured without a wizard and without follow-up calls. Scripts and the assistant reach the same field set — the MCP manage_agent create tool shares it.
Name a schedule and re-posting edits it. The name is that schedule’s identity on the agent, so sending the same create again with a different cron updates it instead of adding a second one — and an identical retry changes nothing rather than firing your agent twice a tick.A trigger that fails is never reported as made. The agent is created before its triggers are, so if one fails you get 207 instead of 200, with the real error and the endpoint that creates just that trigger. Retry the trigger, not the create — re-running the create rotates the agent’s worker token and would invalidate a values.yaml you have already applied.A webhook returns its URL path, header and one-time token in the same response, so it is usable immediately. webhookType: "github" is supported too, given a signing secret.Autoscaling here is what you declared, not what your cluster is doing. It is recorded on the agent and written into the Helm values you install, but AgentOps cannot see your cluster — whether KEDA is installed, or whether it is actually scaling anything.In the Add-Agent wizard, the Triggers step now tells you when you don’t have permission to create one, instead of accepting a schedule the deploy would go on to refuse.Try it in KAOP

Fixed: An archived draft agent stays archived

Archiving a draft agent now actually retires it everywhere. Because a draft never sends a heartbeat, and an agent’s status is derived from its last heartbeat, an archived draft kept reporting itself as a draft forever — so four screens went on treating it as unfinished work long after it had been deployed and archived.Most visibly, the Add-Agent wizard offered it back as “You have an unfinished agent”, and its Continue setup button really did deploy it a second time from a configuration that had already shipped. The Fleet’s Archived tab linked those rows straight back into the same wizard, and an archived draft’s own page told its operator to go and deploy it.Now the offer only ever names a draft you can genuinely still finish, archived rows in the Fleet open their agent page like every other archived agent, and that page says plainly that the draft was archived — with no setup link to follow. Unarchive it from the Fleet if you do want to pick it back up.Try it in KAOP

Fixed: A chat waiting on a busy agent now says so, instead of looking stuck

Asking an agent a question while it was already working on another run left the chat spinning with no explanation, and after five minutes it gave up with what looked like a network error. The run was fine the whole time — it was queued, waiting for the agent to free up, and nothing said so.The wait is now visible. While your question is queued you see “Waiting for a free worker”, with how many runs are ahead of it when there are any — counting down as they finish. The notice clears the moment a worker picks your question up, and the reply streams in as normal.And waiting your turn no longer fails the turn. The timeout that ended those questions exists to catch a run that has genuinely stalled, but it was counting the queue wait against the same budget — so a healthy question that simply had to wait was reported as failed. Being picked up by a worker now resets that budget, so the time your question spent queued is no longer held against the answer.One case is still open: a question that waits more than five minutes for its turn is still given up on, so if the run ahead of yours takes that long you will still see an error.Try it in KAOP
NewFixed

Fixed: The Edit wizard now tells you which changes can actually reach your agent

Two agents can both be self-hosted and behave oppositely. One built from the generic agentops-agent-base image takes its instructions, model, capabilities and MCP servers from the config the chart writes — so editing them here and running the generated helm upgrade genuinely changes how the agent behaves. One built from your own image carries its own agent-spec.yaml and never reads that config at all.The wizard used to present both the same way. Editing the prompt of an agent in the second group produced a success message, a values.yaml, a helm command, and no change in behaviour whatsoever — plus a Redeploy pending badge in Fleet that nothing could ever clear, because the running agent was never going to move toward what you saved.Now the wizard works out which case your agent is in and says so:
  • Fields your image owns are shown read-only, with the reason and a pointer to the repository whose agent-spec.yaml actually defines them. Change them there and rebuild the image.
  • What still applies from here stays editable — the display name, triggers, replicas and scaling all reach your agent regardless of how it was built.
  • Fleet distinguishes Redeploy pending (you have changes to deploy, and the badge clears once you do) from Spec mismatch (your image disagrees with the saved config, and no redeploy will reconcile it).
  • Values your worker will ignore are no longer recorded as if you had chosen them, so the badge that could never be cleared can no longer appear.
When your agent hasn’t reported enough for us to tell which case it is, nothing is locked and no promise is made either way — you get a note saying so rather than a guess.Try it in KAOP

Fixed: Model lists come from the provider, not from a built-in list

Adding models to a provider you configured earlier used to show a short built-in list for that provider kind, because the key you had saved could not be read back to ask with. The list looked like an answer, but it was the same handful of ids for everyone — including, for Bedrock, models AWS has since retired.AgentOps now asks the provider itself, using the key you already gave it: Bedrock (foundation models and inference profiles), Gemini, Cohere, Azure deployments, and every OpenAI-compatible provider answer with what your key actually serves. When a list still cannot be confirmed — a provider with no list to ask for, or a credential saved before this change — the page says so instead of passing the built-in list off as live. Saving that credential’s key once makes it live.Try it in KAOP

New: Jump to any page with ⌘K

Press ⌘K (or Ctrl+K) anywhere in the app to open the command palette: type a few letters, hit Enter, and you’re on the page — Members, API keys, Quality Lab evaluators, an Integrations section — without reaching for the sidebar. It searches everything you can navigate to, including sub-pages and tabs, understands aliases like token for API keys or slack for channels, and keeps your five most recent jumps one keystroke away. There’s also a Search… button in the header if you prefer the mouse.
ImprovedFixed

Fixed: A trigger set in the Add-Agent wizard is now created when you deploy

The Triggers step used to collect your cron and drop it: the agent deployed, the wizard reported success, and no schedule existed until you reopened the agent and added it again. Now the trigger is created together with the agent, for cloud and self-hosted deploys alike — a schedule fires from day one, and a token webhook is created with its endpoint URL and one-time inbound token shown right on the review step.If the agent deploys but its trigger can’t be created, the wizard says exactly that — the agent exists, the trigger doesn’t, and why — with a one-click retry. No more silent success. One deliberate exception: a GitHub-type webhook needs a signing secret the wizard doesn’t collect, so instead of creating an endpoint that would reject every delivery, the wizard points you to Integrations → Endpoints to finish it with the secret after deploy.Try it in KAOP

Improved: Reuse a model credential you already stored, whatever you named it

Deploying a catalog worker no longer asks you to paste a key the account is already holding. The Credentials step used to look for a credential named exactly what the worker expected, so a key you had stored as anthropic-prod read as Missing and the only way forward was to paste it again under the expected name — leaving two copies of the same secret to rotate.Now that step offers the credentials your account already has, and picking one answers the requirement: no second copy, and no renaming. Adding a brand-new credential is still there, presented as the last resort rather than the only option. If you don’t have permission to list the account’s credentials, the step says so instead of showing you an empty list with no explanation.Try it in KAOP

Improved: Search for a credential or timezone instead of scrolling and typing

Creating an agent no longer means hunting through an unsorted credential list — the model step’s credential picker now has a search box and lists credentials alphabetically with their descriptions. The Triggers step’s timezone field is a searchable picker too, rather than a text box you had to spell an IANA name into by hand: it shows every zone with its current UTC offset, puts your own zone and UTC at the top, and finds a zone by offset (+05:30) or by its modern name as well as its older one. Both changes also apply when you edit an existing agent, and the same timezone picker now backs the schedule forms in Fleet.Try it in KAOP

Fixed: History search now looks at your whole workspace, not just the newest page

Searching History used to filter only the 100 most recent sessions the page had already loaded, so a run or chat older than that came back as “No matching history” — for something that was plainly still there. Search now runs against the whole workspace, and matches on the run or session id, its title or prompt, and the name of the agent behind it, so an agent’s name finds its work however far back it goes.History is also honest about the page now: when the list is showing as much as it can, it says so instead of quietly stopping, and the tiles above it (Runs, Active, Run cost) report the last 30 days across the workspace rather than a tally of the rows on screen.Try it in KAOP

Fixed: Editing an agent no longer creates a GitHub webhook that rejects every delivery

Choosing Webhook → GitHub in either agent edit wizard used to create the endpoint immediately, without a signing secret. GitHub deliveries are verified against that secret, so every one of them was rejected — while the trigger row existed and Fleet showed a configured webhook. The only visible symptom was an agent that never ran, which reads as a broken agent rather than a webhook that was never valid.The edit wizards now behave like the Add-Agent wizard: a GitHub webhook isn’t created there at all, and the step says why. Create it under Integrations → Endpoints instead, which collects the signing secret and routes the endpoint to your agent. Everything else about the edit saves normally, and static-token webhooks are unaffected.If you already have one of these, you can spot it: Integrations → Endpoints now marks a GitHub endpoint with no signing secret with a No signing secret badge. Fix it in place — edit the endpoint and add the secret.Try it in KAOP

Improved: A searchable agent picker for eval rules

Setting up an eval rule used to mean scanning an unlabeled wall of pills for the agent you wanted. The Targets step now has a searchable, alphabetically-sorted list instead — type a few letters to find an agent, and the count of what you’ve picked stays visible with a one-click Clear.The step also drops the “Run filter” control: every rule now evaluates successful runs, which is what every existing rule already did. If the agent you picked has never had a successful run, the review step says so before you activate the rule.Try it in KAOP
ImprovedFixed

Fixed: Create agent always starts fresh — and your draft is still there if you want it

Clicking Add agent now always opens a clean wizard, so you can start a second agent, or switch between building from scratch and deploying from the catalog, without the previous attempt following you in. Your unfinished work is not thrown away: it is offered as a “Continue where you left off” card on the start screen, naming the agent and the step it reached, with a Discard beside it. And if you simply reload the page mid-wizard, you land straight back where you were — no card, no click.Try it in KAOP

Improved: Create agent asks where it runs up front, and only shows the steps that apply

The agent wizard now asks where the agent will run at step two, right after you name it — your own cluster, or Komodor cloud. Previously the question waited until the very last screen, which meant you could work through Scaling and set up KEDA autoscaling for an agent you then chose to have Komodor run, where those settings do not apply and were quietly dropped.Answering up front shapes the rest of the wizard: pick Komodor cloud and the Scaling step is not there at all, because there is nothing for you to scale. Pick self-hosted and it is, as before. The Review screen no longer asks a second time — it opens straight on the deploy path you chose. Catalog workers get the same question when they can run in both places, and skip it when their image only runs in one.Drafts already in progress are unaffected: an unfinished agent resumes on the step you left it on, still set to run in your own cluster.Try it in KAOP
Fixed

Fixed: Change which specialists an orchestrator and an incident workflow use

An orchestrator’s delegatable workers are now shown on its page in Fleet, and you can change them there — previously the list was chosen at deploy time and could never be seen or edited again, so an orchestrator could not gain a new specialist after launch. Incident workflows follow the same rule as code review workflows: pause one and it becomes editable, so you can add a specialist to a workflow that has already gone live.Try it in KAOP
Improved

Improved: See the impact and cost of your code review program

Code Review now opens with a 7-, 30-, or 90-day view of review volume, open findings, spend, webhook coverage, and performance by repository and reviewer. Compare the typical first pass with incremental re-reviews, and hover the engineer-time estimate to see exactly how it was calculated.Try it in KAOP
New
Deploy the AgentOps Orchestrator straight from the agent gallery — a lead agent that, given a problem, decides which specialists in your fleet should handle it, delegates a scoped sub-task to each, and synthesizes one combined answer. Chat with it directly or trigger it as a run.Try it in KAOP

New: The Code Review Platform is now available to every account

Point a reviewer agent at your repositories and every new pull request gets an automatic review — no waitlist. Set up a review workflow, watch reviews land in the inbox as GitHub pull requests open and update, and see each reviewer’s verdict alongside the PR.Try it in KAOP
NewImproved

New: Test a new worker version in shadow, next to production

Dual-deploy (A/B) is here for self-hosted agents: deploy a candidate version of a worker next to the serving one, and every request runs on both — while callers keep getting only the production answer. The candidate can never slow down, break, or change what a user hears.Turning it on is one environment variable on the candidate deployment — same agent id and worker token, plus AGENTOPS_VARIANT_ROLE=secondary. No changes to the serving version.Each request then produces a comparison pair you can spot everywhere:
  • History — the candidate’s run is tinted amber and badged Shadow, production’s is badged Production, and each row links its sibling. A pair_… id correlates the two.
  • Run console — an amber band compares both arms side by side: status, duration, cost, and output.
  • Agent page — the Instances tab marks the candidate replica, and the Versions tab splits traffic into production · shadow instead of inflating the count.
For a local end-to-end loop, aops shadow spins up two demo workers (aops shadow start --agent demo-ab), fires one message at both (aops shadow trigger), and verifies the pair ran on both versions (aops shadow validate).Try it in KAOP

Improved: Choose which agent remediates an incident workflow

The Remediation step of the incident workflow wizard is now a choice rather than a fixed panel. Pick any remediation agent you have deployed, or Investigate only to stop at the conclusion and hand off to a person. The investigation itself is unchanged either way — it still produces the same evidence-backed findings.A workflow binds at most one remediation agent, because the orchestrator hands the findings to exactly one worker once triage converges. The step lists every remediation agent in your fleet, offline ones badged with their status — your pick is the only remediation worker the workflow’s runs can reach, so an unreachable one means remediation does not run, never that another agent stands in for it. An existing workflow whose agent has left the fleet says so instead of failing quietly.Try it in KAOP
New

New: Describe what you want and the form fills itself

Every multi-step form in the app — new agent, credential, integration provider, evaluation flow, code-review and incident workflows, incident pipeline, maintenance agent — now opens with a co-pilot beside it. Tell it what you are trying to set up in plain language and it fills the fields in, walking you to the step it changed, so you review a filled form instead of starting from an empty one. It never submits and never touches a secret field: you type the secret yourself and you press create, so nothing exists until you say so.Try it in KAOP
NewImproved

Improved: Every Slack answer names the runs behind it

An AgentOps reply in Slack now ends with a footer naming each agent that contributed to the answer, with a link to its run. When the assistant delegated the question to a specialist, the footer names the specialist and the assistant that routed it — so a summarized answer is one click from the full output it was built from.The footer is written by the control plane, not by the agent, so it stays accurate even when the reply is a paraphrase.Try it in KAOP

New: Run the Kubernetes RCA agent inside your own cluster

Kubernetes RCA now installs into your own Kubernetes cluster. Pick it from the marketplace and the wizard hands you a worker token and a ready-made helm install — the agent runs on your side, investigates the cluster it is installed in, and reports findings back to AgentOps. Nothing about your cluster leaves it except the finished analysis.This is the only way to deploy it, by design: the agent’s evidence is read-only kubectl against the cluster it runs in, so a Komodor-cloud copy would faithfully investigate the wrong cluster.The generated values grant read-only access — a binding to Kubernetes’ own view role plus the cluster-scoped reads it omits, such as nodes and storage classes. Secrets and pod exec are deliberately not granted; both are documented opt-ins you enable yourself if you want them.Investigating clusters connected through a Komodor agent instead? Kubernetes RCA (Actions) still runs on Komodor cloud and needs nothing installed.Try it in KAOP

Improved: Continue an investigation from a finished run

Beneath a finished run’s output there is now a Continue this investigation composer. Type your own follow-up, choose who answers, and press Continue — a chat session opens with your question already asked against that run.It replaces the old Review in chat button, which could only ask one fixed question of one fixed agent. Now the run’s own agent answers when it supports chat; otherwise the AgentOps Assistant steps in and the composer tells you why. Leave the box empty and it simply asks what to look at next, rather than assuming something went wrong.You can also hand the investigation to Claude Code or Codex, or copy the prompt for any other coding agent. Those handoffs carry only the run’s id and link — never its output, logs or traces — so nothing sensitive ends up in a desktop URL. The coding agent reads the run through AgentOps MCP using your own permissions.Try it in KAOP

Improved: Filter one agent’s run history by status

An agent’s Run history tab now has the same status filter the global run ledger has, so you can pull up just that agent’s failures — or just what is queued right now — without paging through everything it has ever done. The choice lives in the URL, so a refresh keeps it and a link you paste to someone else opens on the same filtered view.Try it in KAOP

New: See a month of an agent’s runs at a glance

An agent’s Run history tab now carries a 30-day density strip in its header: one thin bar per day, stacked by what those runs did — blue for running, yellow for queued, red for failed, green for succeeded, grey for cancelled. Every colour is drawn on the same pixels-per-run scale, so a segment’s height is a count you can compare between days, and a day with nothing at all is a flat grey tick rather than a gap. Run activity, the metrics-and-trend dashboard, has moved to a tab of its own, so the history tab is just the runs.Try it in KAOP
NewImprovedFixed

New: Ask the AgentOps assistant to review a finished run

A finished run now has a Review in chat button. Press it and the AgentOps assistant opens with the question already asked — it reads the run’s output, logs and trace and tells you whether anything actually went wrong, and what to do about it.It is offered on every completed run, not only failed ones, because a run that could not do its work still finishes as succeeded — an agent that found no tools bound, an integration that was never authenticated, or a bad credential all report success. If the run did what it was asked, the assistant says so.Try it in KAOP

Fixed: Pausing a review workflow now actually lets you edit it

Editing a review workflow that had ever been activated used to fail with “Only draft workflows can be updated. Pause the workflow first.” — advice that could not be followed, because pausing left the workflow paused rather than back in draft, so the same error came back every time. A paused workflow is now editable, which makes pausing the real unlock it always claimed to be. An active workflow still cannot be edited underneath a live webhook, and the detail page now hides its Edit button until you pause, instead of offering an action that was going to be rejected on save.Try it in KAOP

New: See how an agent’s volume, spend and reliability moved

An agent’s Run activity now charts a metric over time instead of showing only today’s totals: invocations, spend, cost per invocation, error rate, and the days it changed version. Pick a tile to plot it, and each tile carries its week-over-week move — so “it costs more this week” and “it started failing on Tuesday” are things you can see rather than infer. Version changes are marked on the chart, which is usually the first place to look when a line bends.Try it in KAOP

Improved: See agent outcomes, latency, cost, and changes together

Agent run history now opens with successful investigations and errors stacked by day, with bar or trend views grouped by day or version. Cost shows total and per-invocation spend, latency is tracked from completed runs, and selecting a change bar opens that exact revision’s field-level diff in Versions.Try it in KAOP
NewImproved

Improved: Check connectivity now reaches the agent itself

Check connectivity now sends a request down the agent’s own connection and waits for the agent to answer, showing you what it said. Previously it started a run that was answered without ever reaching your agent’s code — so it could report success for a worker that could not actually do anything. It also tells you which thing is wrong: nothing connected, not responding, or connected but reporting a problem of its own.Try it in KAOP

New: Agents can now cite your knowledge base

Agents you build from scratch, and the investigators in the Komodor catalog, can now search your knowledge base while they work — the same runbooks, service docs and postmortems you uploaded on the Knowledge page. An answer that leans on one comes back with the document title and section it came from, so you can check the source instead of taking the agent’s word for it. There is nothing to turn on: an agent picks this up the next time it runs on a current image.Try it in KAOP
NewImproved

Improved: A review now shows every pass it made, not just the latest

A pull request that gets pushed to repeatedly is reviewed repeatedly, and until now the review page only showed where things stand right now — the newest run per reviewer. A new Review history section lists every pass the review has dispatched, newest first: which revision it looked at, what it concluded, and a link to the run. Passes that produced no verdict are in the list too, labelled by why — a run that failed reads differently from one that finished but returned output the platform could not parse, and previously both simply left a gap where a verdict should have been. Selecting a revision in the history switches the reviewer cards above to that revision, so the two always agree about what is on screen.Try it in KAOP

New: Connect Google Workspace and pick which products agents can use

Google Workspace is now in the integrations catalog. One connection covers Drive, Docs, Sheets, Slides, Calendar, Chat, Contacts, and Gmail — you choose which of them to include when you connect, and only the scopes those products need are requested on the consent screen. Drive is selected by default; Gmail is never preselected, since it needs a restricted scope. Each selected product is registered as its own MCP server behind the same credential, so agents get exactly the Workspace tools you approved. You can come back later and add a product to an existing connection — the consent screen only asks for the new scopes. “Test connection” now makes a real read-only call to Google rather than just checking that a token is stored, so a revoked token or a Workspace API that was never enabled for your Google Cloud project shows up as a failure with the reason.Try it in KAOP
Improved

Improved: Run transcripts now read like the agent’s actual session

A run’s transcript now shows what the agent said between its tool calls, in the order it happened — reasoning, then the tools that reasoning led to, then the next step. Previously the agent’s narration was collected up and dumped as a single block of text at the end of the run, so a long run looked like a wall of bare tool calls followed by one paragraph that mentioned all of them at once. The run’s Output card is unchanged: it still shows the final answer, and while a run is in flight it previews what the agent is working on right now rather than everything it has said so far.Try it in KAOP

Improved: See every model you can call — and what it costs — in one table

Providers has moved under Connect and is now a single table of models rather than a tab per tier and a card per credential. Each row shows the provider, who manages it (Komodor or one of your own credentials), the upstream model behind it, its spend, its requests and tokens, and whether it is working — so searching and filtering covers the models Komodor manages and your own at the same time, with your gateway budget above it. Adding a provider now ends on a verify step where you can send a model a real prompt and read its reply, and Baseten and Nebius Token Factory join the list of providers you can bring.Try it in KAOP

Improved: Reviewers now pick up where they left off when you push a fix

A reviewer used to start from a blank sheet on every push, so findings could appear and disappear between passes, ones you had already declined came back, and a fix was never acknowledged. A re-review now carries the reviewer’s own previous findings: it works out which ones your new commits fixed, keeps the rest open unchanged, and reviews the new commits in full — widening beyond them whenever the change warrants it. Each review reports the ledger — what is new, what is still open, what you resolved — and a fixed finding gets answered on its own comment thread instead of only in a fresh summary you have to cross-reference. A push that changes no content at all, such as an amend or a rebase in place, no longer produces a second review of identical code. Hitting Re-review still gives you a full pass from scratch.Try it in KAOP
NewImproved

Improved: Configure a whole Slack routing rule from the agent spec

A slack_channel trigger in agent-spec.yaml now carries the whole routing rule, not just the channel list. Pick what the agent answers with rule — mention (the default), channel for every human message, keyword for messages containing a term, or private for direct messages — add keywords, opt into other apps’ posts with allow_bots: true, and order the route against your other rules with priority. Declaring it in the spec is enough; no follow-up step on the channel page.Try it in KAOP

New: Pick one of your own providers when you create an agent

The new-agent wizard no longer offers a fixed Anthropic model list. Alongside the models Komodor manages, it now shows the provider deployments you registered on Providers, under your own names, with the provider kind beside each one. A deployment you cannot call yet is still listed but not selectable, and says why — it is still verifying, or your account routes straight to Anthropic instead of through your providers. The model you choose is checked again when the agent is created, so a deployment that stops being reachable is refused at creation rather than at the agent’s first run.Try it in KAOP
Improved

Improved: Page through the reviews inbox instead of scrolling all of it

The Reviews inbox now shows 10 reviews per page instead of every review at once, so the page stays a readable length as your review history grows. Pick 25, 50, or 100 per page if you’d rather see more, and your page and page size stay in the URL — so a refresh keeps your place and a copied link opens on the same page. Narrowing the list with a status/repo filter or the search box takes you back to the first page so you never land on an empty one.Try it in KAOP

Improved: Connect AWS Bedrock with a Bedrock API key

Adding a Bedrock model no longer forces you to paste an IAM access key pair: a single AWS Bedrock API key now works on its own, with the access key pair, session token and runtime endpoint left optional beside it. Each field explains what it is for, and Providers tells you up front that a Bedrock credential needs either the API key or a complete access key pair before it will save.Try it in KAOP
Fixed

Fixed: Open a review workflow to see its status, not the setup wizard

Clicking a review workflow used to relaunch the setup wizard, landing on the last step every time. It now opens a read-only detail page showing status, repositories, webhook health, and reviewers at a glance — with an explicit Edit button when you actually want to change something.Try it in KAOP
Improved

Improved: Search reviews, track revisions, and see reviewer cost

Find a review instantly by searching a PR number, title, author, or repo. When a PR gets pushed to again, switch between revisions on the review page to see what each reviewer said about that specific commit, and a “Merged” badge now shows alongside merged PRs so you don’t mistake them for a plain close. Reviewer cards on the Reviewers tab also show 30-day spend, and review detail now renders full-width for easier reading.Try it in KAOP
Improved

Improved: Filter the fleet by team, tier, or any other label

Fleet now lists your agents as a scannable table — name, status, last run, last change, and instances — with the filter sidebar every other platform page uses. Each label key you use (team, tier, env, role, …) becomes its own multi-select category ordered by how many agents carry it, so you can narrow to team: sre or team: platform and tier: production in two clicks and share the resulting URL. Credentials filters the same way.Try it in KAOP
New

New: Quality Lab — evaluate agent runs with judge panels

Quality Lab is now an evals platform. Set up an eval rule — pick which agents’ runs to evaluate, how often to sample them, and a panel of judges (any agent with the grader capability) — and every matching run is graded automatically. Browse the verdicts in the Evals inbox, open a run to read each judge’s score and reasoning side by side, and see the Judges roster and a single quality score per agent on the Agents tab. Human star ratings sit right alongside the judges’ verdicts.Try it in KAOP
New

New: Uninstall managed agents from the Fleet

Agents you deployed from the catalog now show a Managed badge in the Fleet, and you can tear one down without leaving the product. Archive the agent, then use Uninstall on its detail page — this decommissions the deployment and revokes its worker token. No more agents lingering in the deployment after you archive them.Try it in KAOP

New: Deploy catalog workers to your own infrastructure

Catalog workers like AWS Investigator can now be deployed self-hosted — on your own Kubernetes cluster instead of Komodor cloud. The marketplace review step generates a worker token and a ready-to-use Helm install command with your credentials. Choose between bringing your own Anthropic API key or using the managed LLM proxy.Try it in KAOP

New: Pause, resume, and delete code review workflows from the list

Each review workflow now has inline actions on the Review workflows list. Pause an active workflow to stop it from reviewing new pull requests, resume a paused or draft one to provision its webhooks again, and delete a workflow you no longer need. Delete only appears once a workflow is paused (or still a draft), so an active workflow can’t be removed out from under in-flight reviews.Try it in KAOP
New

New: Automate incident response with alert-triggered workflows

Create an incident pipeline that wires any alert source — Datadog monitors, AlertManager, or a generic webhook — through routing rules to an orchestrator and its specialist team. Each firing alert becomes an incident, gets deduplicated and severity-mapped, and dispatches an investigation automatically.Try it in KAOP

New: Prove a new agent works the moment it’s deployed

Every agent now ships with two built-in self-check commands. Use Check connectivity for an instant, model-free confirmation that a freshly deployed agent is reachable, or Sample run for a real run through the model — so you can prove a new agent works in one click, right after deploying it.Try it in KAOP
New

New: See what shipped recently, right here

AgentOps now has a What’s New page — this one — with everything recently shipped, newest first. Watch for the badge on “What’s new” in the sidebar: it counts the updates you haven’t seen yet and clears when you visit.Try it in KAOP

New: Knowledge base is now a real wiki

The Knowledge page now reads like a wiki: pages render as formatted Markdown with working links between documents, a navigation tree grouped by document type, and a Raw view of the original file. Drop several files at once anywhere on the page to upload them in one batch — OKF frontmatter (title, type, tags, owner, source) is picked up automatically, and you can watch indexing progress live.Try it in KAOP
NewImproved

New: Cancel runs that are queued or in flight

You can now cancel a run from History or the run detail page — whether it is still queued, claimed, or already executing. Cancellation interrupts the worker itself rather than just marking the row, so a runaway run stops doing work (and spending tokens) the moment you cancel it.Try it in KAOP

Improved: Search and filter the audit log

The audit log now supports free-text search across actors and actions, plus a status filter, with server-side sortable columns. Finding who did what — and whether it succeeded — no longer means paging through the whole ledger.Try it in KAOP