What you see
Pick a window of Last 7 days, Last 30 days, or Last 90 days — 30 by default.
The breakdown can be pivoted by Agent, Model, or Provider.
Cost also appears where you are already looking: the account Overview shows Spend (30d),
Fleet health shows cost per finished run across the
fleet and window cost for each agent, Insights shows known
cost and the per-model mix, and every run’s detail shows its own usage.
Where the numbers come from
Every run carries a usage record: the provider, the model, the token breakdown, a cost amount, a currency, and a cost source. The usage is merged from what the worker reports during the run — token events, the run’s output, and its metadata. Each run’s cost then falls into one of three classes.
Runs with no usage record at all contribute runs but no cost.
Why the split is shown rather than hidden. A single “total spend” number silently mixes figures a
provider billed you for with figures KAOP computed from token counts and published rates. Those have
different error bars, and treating them as one is how a cost review ends in an argument. The split is
on the tile so you always know how much of the total is measured and how much is modelled.
When a worker’s own figure is discarded
If a run names a provider KAOP does not price directly and the model’s rate is resolvable, the worker’s own cost figure is discarded and the cost is re-estimated from those rates. That is deliberate: a worker-supplied figure for a provider the platform cannot corroborate is not more trustworthy than a rate-based estimate, and mixing the two conventions within one total is worse than either. Where cache rates are not published for a model, they are derived from its input rate — cached reads at a fraction of it, cache writes at a premium to it.Making your agents’ cost show up
Agents built on the KAOP SDK adapters report token usage automatically. If you build a worker yourself, include a usage object in the run’s output or metadata. Several field aliases are accepted, so most existing usage shapes work unchanged:
Report a cost amount to get reported cost. Report a model plus token counts to get an
estimate. Report neither and the run contributes volume but no spend.
Reading spend through the API
The Spend view is backed by the account activity endpoint, which you can call directly.window_days accepts 1 to 365.
Budgets
A budget caps what an agent may spend in a time window, and optionally stops it when the cap is reached. You set one on the agent itself. Go to Operate → Fleet, open the agent, and pick its Budgets tab — it offers Add first budget until that agent has one. Budgets are per agent, so there is no fleet-wide cap: an agent with no budget is uncapped regardless of what the others have. Every agent’s budgets are also listed together on the Budgets view, which summarizes them across the fleet and is where you compare them.Budgets are enabled per account. If the Budgets tab is not on your agents, ask your Komodor
contact to turn it on — see Support.
Setting one
1
Pick the windows
Under Spend cap by window, enable any combination of Hourly, Daily, Weekly, and
Monthly, each with its own cap amount. Each window is an independent budget that resets on
its own clock, all of them in UTC — hourly at the top of the hour, daily at midnight, weekly on
Monday, monthly on the first.
2
Set the notification threshold
A percentage of the cap. Crossing it is what makes the budget read Near cap rather than
OK, before it is actually reached.
3
Choose what happens when the cap is hit
Notify only or Block new runs.
What “Block new runs” does
The behavior is specific, and worth reading before you enable it:- The agent’s schedule, webhook, and channel triggers stop firing until the window resets. A run that was already queued is cancelled.
- A run already in progress is terminated when the cap is hit.
- Manual invocation still works. You can always run the agent yourself from the console.
Reading the Budgets view
Three tiles summarize the picture: agents with a budget, the combined cap across all windows, and how many budgets are at or near their cap. The table then lists each budget with its Cap & window, Spend this period, Status, and Resets in.
Spend against a budget counts finished runs, so a run still in flight is not yet charged against the
window — with one exception: the enforcement check for a blocking budget includes runs currently
executing, because a cap you can exceed by starting enough runs at once is not a cap.
A budget cannot be created for an agent that brings its own model key. KAOP does not see that
agent’s spend, so a cap it could neither measure nor enforce is refused rather than accepted and
quietly ignored.
Bringing spend down
1
Find the concentration
Pivot the breakdown by Agent, then by Model. Spend is almost never evenly spread — one
agent or one model usually dominates.
2
Read cost per success, not total
A high-volume agent with a low unit cost is working. A low-volume agent with a high unit cost is
the one to look at.
3
Open the most expensive runs
Each one links to its evidence. An outlier run usually explains itself — a retry storm, an
oversized context, or a tool loop.
4
Settle any change as an experiment
Swapping to a cheaper model is a quality decision, not a billing one. Prove it against real
history before you apply it.
Next steps
Fleet health
Cost beside reliability, per agent and per window.
Runs & evidence
A single run’s usage record, and what it did to earn it.
Observability & OTel export
Token volume as an operational signal in your own backend.