Skip to main content
Agent spend answers what your agents cost, where the money went, and whether it is trending up. The Komodor Agentic Operation Platform (KAOP) builds every figure from the usage each run reports, so a total on this page can always be traced back to specific runs. Figures are reported from the usage each run emits, not sampled or estimated after the fact, so any total can be traced back to the specific runs behind it. Find it under Improve → Cost.

What you see

Pick a window of Last 7 days, Last 30 days, or Last 90 days — 30 by default. The breakdown can be pivoted by Agent, Model, or Provider.
Cost also appears where you are already looking: the account Overview shows Spend (30d), Fleet health shows cost per finished run across the fleet and window cost for each agent, Insights shows known cost and the per-model mix, and every run’s detail shows its own usage.

Where the numbers come from

Every run carries a usage record: the provider, the model, the token breakdown, a cost amount, a currency, and a cost source. The usage is merged from what the worker reports during the run — token events, the run’s output, and its metadata. Each run’s cost then falls into one of three classes. Runs with no usage record at all contribute runs but no cost.
Why the split is shown rather than hidden. A single “total spend” number silently mixes figures a provider billed you for with figures KAOP computed from token counts and published rates. Those have different error bars, and treating them as one is how a cost review ends in an argument. The split is on the tile so you always know how much of the total is measured and how much is modelled.

When a worker’s own figure is discarded

If a run names a provider KAOP does not price directly and the model’s rate is resolvable, the worker’s own cost figure is discarded and the cost is re-estimated from those rates. That is deliberate: a worker-supplied figure for a provider the platform cannot corroborate is not more trustworthy than a rate-based estimate, and mixing the two conventions within one total is worse than either. Where cache rates are not published for a model, they are derived from its input rate — cached reads at a fraction of it, cache writes at a premium to it.
If no rate can be resolved for a model at all, that run’s cost is left unset rather than guessed. It shows up as runs with no cost, which is a data-quality signal about the model, not evidence that the work was free.

Making your agents’ cost show up

Agents built on the KAOP SDK adapters report token usage automatically. If you build a worker yourself, include a usage object in the run’s output or metadata. Several field aliases are accepted, so most existing usage shapes work unchanged: Report a cost amount to get reported cost. Report a model plus token counts to get an estimate. Report neither and the run contributes volume but no spend.
Reporting model and tokens is usually the better choice even when you could compute a cost yourself. It keeps one pricing convention across the fleet, so agent-to-agent comparisons on the breakdown are comparisons of usage rather than of two teams’ accounting.

Reading spend through the API

The Spend view is backed by the account activity endpoint, which you can call directly.
Both return run counts split into succeeded, failed, and active; cost totals split by source; daily buckets; and the per-model mix. window_days accepts 1 to 365.

Budgets

A budget caps what an agent may spend in a time window, and optionally stops it when the cap is reached. You set one on the agent itself. Go to Operate → Fleet, open the agent, and pick its Budgets tab — it offers Add first budget until that agent has one. Budgets are per agent, so there is no fleet-wide cap: an agent with no budget is uncapped regardless of what the others have. Every agent’s budgets are also listed together on the Budgets view, which summarizes them across the fleet and is where you compare them.
Budgets are enabled per account. If the Budgets tab is not on your agents, ask your Komodor contact to turn it on — see Support.

Setting one

1

Pick the windows

Under Spend cap by window, enable any combination of Hourly, Daily, Weekly, and Monthly, each with its own cap amount. Each window is an independent budget that resets on its own clock, all of them in UTC — hourly at the top of the hour, daily at midnight, weekly on Monday, monthly on the first.
2

Set the notification threshold

A percentage of the cap. Crossing it is what makes the budget read Near cap rather than OK, before it is actually reached.
3

Choose what happens when the cap is hit

Notify only or Block new runs.

What “Block new runs” does

The behavior is specific, and worth reading before you enable it:
  • The agent’s schedule, webhook, and channel triggers stop firing until the window resets. A run that was already queued is cancelled.
  • A run already in progress is terminated when the cap is hit.
  • Manual invocation still works. You can always run the agent yourself from the console.
Blocking is the right default for a runaway-cost guard and the wrong default for anything on your critical path. A blocked agent stops responding to its triggers entirely, which for an incident-facing agent means it is absent exactly when it is needed. Start with Notify only, and move to blocking once you know the agent’s normal spend shape.

Reading the Budgets view

Three tiles summarize the picture: agents with a budget, the combined cap across all windows, and how many budgets are at or near their cap. The table then lists each budget with its Cap & window, Spend this period, Status, and Resets in. Spend against a budget counts finished runs, so a run still in flight is not yet charged against the window — with one exception: the enforcement check for a blocking budget includes runs currently executing, because a cap you can exceed by starting enough runs at once is not a cap.
A budget cannot be created for an agent that brings its own model key. KAOP does not see that agent’s spend, so a cap it could neither measure nor enforce is refused rather than accepted and quietly ignored.

Bringing spend down

1

Find the concentration

Pivot the breakdown by Agent, then by Model. Spend is almost never evenly spread — one agent or one model usually dominates.
2

Read cost per success, not total

A high-volume agent with a low unit cost is working. A low-volume agent with a high unit cost is the one to look at.
3

Open the most expensive runs

Each one links to its evidence. An outlier run usually explains itself — a retry storm, an oversized context, or a tool loop.
4

Settle any change as an experiment

Swapping to a cheaper model is a quality decision, not a billing one. Prove it against real history before you apply it.

Next steps

Fleet health

Cost beside reliability, per agent and per window.

Runs & evidence

A single run’s usage record, and what it did to earn it.

Observability & OTel export

Token volume as an operational signal in your own backend.