Skip to main content
Use the sections below to find the issue you’re experiencing in the Komodor Agentic Operation Platform (KAOP) and work through the recommended checks. Run details, error messages, and agent health information can help you identify where to start.
For recurring failures, check Insights first. It groups related failures and includes the reported failure reason, which may help you identify a common cause.

An agent appears offline

What you might see: an agent appears as Offline in Fleet, or remains in draft after you start its worker. An agent’s status reflects its worker’s connection to KAOP. The control plane checks the connection every 30 seconds and marks the agent offline after 90 seconds without a response. A draft agent becomes live when its worker first connects. Work through these checks:
1

Check that the worker is running

Review the worker’s startup logs for errors. If startup fails before the worker connects, the agent may remain in draft.
2

Check connectivity to KAOP

Confirm that AGENTOPS_URL points to your account’s host and that the worker can make outbound HTTPS connections. Workers do not require inbound connections.
3

Check the worker token

The worker authenticates using Authorization: Bearer <worker token>. If the token was rotated or revoked, update the worker with a valid token. You can generate one from the agent’s Permissions tab under Worker token, then restart the worker.
4

Check the agent ID, if configured

If the worker sends an X-Agent-Id header, it must match the agent associated with the token. You can also omit this optional header and let the token identify the agent.
5

Review the last heartbeat

Open the agent’s Metadata tab and check Last heartbeat. This can help you determine whether the connection stopped recently or the worker has been disconnected for longer.
A revoked token and an agent ID mismatch both return 401. If these checks do not resolve the issue, contact Support for help identifying the cause.

An agent is online but does not receive work

An online status confirms that the worker is responding to connectivity checks. Fleet health provides additional information about whether it can receive work.

API or webhook requests return 401 or 403

What you might see: a request returns an authentication or permission error. Start by identifying which type of credential the request requires. Webhook endpoints use their configured authentication method. An API key does not replace the endpoint’s webhook credential.
  • GitHub-type endpoints expect X-Hub-Signature-256.
  • Bearer-type endpoints expect Authorization.
  • If the configured header is absent, KAOP also accepts Authorization: Bearer <token> as a fallback for senders that cannot use custom headers.
Also check whether the webhook’s signing credential is active. A deactivated credential fails verification even if the token value has not changed.
A 403 response can also indicate that the caller lacks the required permission or that the account is suspended.

A run reports a missing credential

What you might see: a run fails, or a tool is unavailable, with an error that names a credential. Start with the error message in the run details. It may identify the missing credential and the configuration needed. For example, a Datadog error may name DD_API_KEY and DD_APP_KEY. Then check the following:
1

Confirm the credential is bound to the agent

Credentials are assigned to individual agents. A credential configured for your account still needs to be bound to the agent that uses it.
2

Check that the credential is available

A deleted or empty credential can cause the connection’s health badge to report a missing credential.
3

Check whether the integration needs user access

Some MCP servers require authorization for your user as well as the application. If the error requests user access, reconnect the integration with that access and test it again.
4

Check the execution context

Credentials are provided during a run. A tool or search helper called outside a run may be unable to request them.
Credential values are not displayed in run evidence or agent logs. To troubleshoot, use the credential’s binding, status, and connection test results.

An integration or MCP server is failing

What you might see: a tool call fails, or a connection displays a warning badge. Check both the connection’s status and its most recent test result:
  • Lifecycle status shows whether the connection is pending, active, in error, revoked, or disabled.
  • Health badge shows the result of the most recent credential check.
An active connection may still need attention if its latest check failed.
Connection checks do not run continuously in the background. Select Test to get a current result.
After you reconnect an integration or update its secret, the previous result is cleared. Test the connection again to verify the updated configuration.

A trigger does not start a run

What you might see: a schedule or endpoint is configured, but the expected run does not appear. Start with the guidance for your trigger type:
  • Is the trigger enabled? A schedule is only registered while it is enabled and has a cron expression. Disabling it unregisters it without deleting it.
  • Is the agent enabled? A disabled agent does not fire triggers. It still answers manual and chat invocations, which is why a disabled agent can look healthy while its schedule is silent.
  • Is the cron expression what you meant? An invalid expression is rejected at write time with the expression quoted back at you. So is a cron-and-timezone pair with no single UTC equivalent, which is how daylight-saving edge cases surface.
  • Did the request pass verification? A request that fails the endpoint’s authentication is rejected before any run is created. See API or webhook requests return 401 or 403 above.
  • Is the target enabled? A request aimed at a disabled agent or a disabled workflow is refused with a conflict rather than queued.
  • Is the body the right shape? For an endpoint, the whole request body is the agent’s input. The invoke and schedule APIs differ: they wrap the object under an input key.
  • Does the target exist? A missing endpoint, trigger, agent, or workflow returns a not-found response, which is worth distinguishing from an authentication failure.
A channel rule fires on a mention by default, and messages posted by bots do not fire it unless the rule explicitly allows bots. That is the usual reason an automated Slack post is ignored while a person’s identical message works.
If a run was created but remains queued, continue to the next section.

A run is queued or fails before work begins

What you might see: a run remains queued, or fails without showing agent activity. The failure category can help you choose your next step. Runs also have time limits:
  • A queued run can fail if no worker becomes available within the allowed time. This may happen sooner when the target agent is already offline, in draft, or has failed to deploy.
  • A claimed run can fail when it reaches its execution time limit, including when a worker stops responding.

Knowledge search returns no results

What you might see: an agent’s knowledge search returns no matches, or the Knowledge page reports that search is unavailable. Check the following:
1

Search availability

Semantic search requires an embedding provider. If the error says the embedding backend is not configured or cannot be reached, contact your Komodor representative for help.
2

Document indexing status

Newly uploaded documents become searchable after indexing finishes. If a document has an error status, re-index it.
3

Search filters

Filters for service, environment, document type, or sensitivity may exclude relevant documents. Try broadening the filters.
4

Knowledge base access

Confirm the documents are in a knowledge base available to the agent. Search is scoped to your account and the agent’s configured knowledge bases.
An empty result can also mean the search completed successfully but found no matching documents. Check the response to distinguish this from search being unavailable.

HTTP status codes

Use this table to help interpret an API error.
KAOP’s control plane does not return 429 rate-limit responses. If a run reports 429, check the response from the upstream provider used by the agent.

Need more help?

If the issue continues, contact Support. Include the run ID and exact error message, if available, along with the checks you have already tried.

Next steps

Limits & quotas

Review limits that may affect your request.

Runs & evidence

Learn how to inspect run details, spans, and logs.

Support

Find contact information and details to include in your request.