AI monitoring for AI agents
AI monitoring gives you visibility into how well a deployed AI agent's responses meet quality standards, and how much it consumes in token resources. Collibra monitors agents by reading signals, such as quality evaluations or invocation outcomes, directly from the AI platform each agent is integrated through. What signal is read, and how often, depends on the AI integration; see AI Monitor assets and Pass and fail criteria below. Monitoring results feed into the Operational Health theme of the AI Trust Score for AI Agent Version assets, and are surfaced in the agent Quality tab and the Agents dashboard.
For a conceptual overview of how AI monitoring fits into the broader operational trust framework, go to Operational trust for AI agents.
| Data collected | Description |
|---|---|
| Judge pass count | The number of agent interactions that received a passing verdict from the LLM judge in the hour. |
| Judge fail count | The number of agent interactions that received a failing verdict from the LLM judge in the hour. |
| Judge no-assessment count | The number of agent interactions for which the LLM judge returned no verdict in the hour, for example, because the judge had not yet completed evaluation at the time of sync. |
| Prompt token count | The total number of prompt (input) tokens consumed by agent interactions in the hour. |
| Completion token count | The total number of completion (output) tokens generated by agent interactions in the hour. |
Enabling the feature
AI monitoring must be enabled. In Settings, under AI Command Center, switch on Operational trust for AI agents. For complete steps, go to Set up AI Command Center
How AI monitoring connects to deployments
AI monitoring operates at the deployment level. A specific AI Agent Version is linked to an AI Endpoint through an AI Agent Deployment, which is a complex relation. One or more AI Monitor assets can be linked to each deployment.
Whether monitoring data is shared across an agent's versions, or scoped to one version's deployment, depends on the AI integration:
- Databricks AI: An AI Agent's MLflow experiment ID does not vary by version, and Collibra uses that experiment ID as the stable join key for monitoring data. As a result, operational trust data is anchored at the AI Agent level, not captured separately per version. A single AI Agent can have multiple versions, each with its own deployments and endpoints, but they share the same underlying monitoring signals.
- AWS Bedrock AI: Each AI Monitor asset is linked to one AI Agent Deployment, which pairs one AI Agent Version with one agent alias. Monitoring data is scoped to that deployment rather than shared across the agent's versions, so versions using different aliases can have different monitoring results.
- Gemini Enterprise Agent Platform: Monitoring data is attributed to a specific AI Agent Version on a best-effort basis: Collibra infers the current version from the agent's live traffic configuration together with its highest-numbered linked version. If the agent has a manual traffic split active across multiple versions, this attribution is not possible, and monitoring data is posted against the AI Agent itself instead of any specific version.
- Azure AI Foundry: The agent's model deployment endpoint is used as the stable join key, and this endpoint is tied to a specific AI Agent Version. As a result, operational trust data is anchored at the AI Agent Version level, not shared across an agent's versions. A single AI Agent can have multiple versions, each with its own deployment and monitoring results.
| Relation | Public ID | Kind |
|---|---|---|
| AI Agent Version + AI Endpoint → AI Agent Deployment | AIAgentDeployment | Complex relation |
| AI Agent Deployment is monitored by AI Monitor | DeploymentMonitoredByAIMonitor | Explicit relation |
AI Monitor assets
An AI Monitor asset represents a monitoring source for an agent deployment. AI Monitor assets are created automatically when an integration syncs; you do not create or configure them manually. What one AI Monitor asset represents, and how it is named, depends on the AI integration:
| AI integration | What one AI Monitor asset represents | Asset name |
|---|---|---|
| Databricks AI | One LLM judge (scorer) configured on the agent. | The name of the LLM judge in Databricks. |
| AWS Bedrock AI | One agent alias. | The alias name, followed by "CloudWatch Monitor". |
| Gemini Enterprise Agent Platform | A single anchor asset per agent, used only to post pass/fail and token metrics. It has no relation to any AI Agent Deployment and does not represent drift detection. | The agent's display name, followed by "(pass/fail monitor)". |
| Azure AI Foundry | A single anchor asset per AI Agent Version, used only to post pass/fail and token metrics. Created only for agents built through Azure's newer Agents API on Foundry project resources; agents from the legacy Assistants API or a classic ML workspace are cataloged without an AI Monitor asset. |
A single AI Agent can have multiple AI Monitor assets, one for each judge or alias that monitors the agent's responses, depending on the integration. AI Monitor assets are created in the domain that you configured for the relevant integration capability, the same domain used for the rest of that integration's synced assets. The following attributes are defined on the AI Monitor asset type.
| Attribute | Description | Public ID |
|---|---|---|
| Data Drift Detection Enabled | Whether data drift detection is enabled on this monitor. | DataDriftDetection |
| Prediction Drift Detection Enabled | Whether prediction drift detection is enabled on this monitor. | PredictionDriftDetection |
| Schedule | The frequency at which monitoring data is synced from Databricks into Collibra. | Schedule |
| Alert Configuration | The threshold conditions that trigger an alert, and the notification channels used to notify stakeholders when a threshold is breached. | AlertConfiguration |
| URL | The URL of the LLM judge in Databricks. | Url |
Connection to the AI Trust Score
Monitoring results feed into the Operational Health theme of the AI Trust Score for AI Agent and AI Agent Version assets. The score for this theme is the 7-day average monitor pass rate: the percentage of monitoring checks that passed over the last 7 calendar days, expressed as a value from 0 to 100. As described above, whether this score is shared across an AI Agent and its versions or attributed separately per version depends on the AI integration.
If no monitoring data is available for this agent for the past 7 days, for example because no interactions have occurred or the agent has no AI Monitor assets configured, the Operational Health theme is excluded from the trust score entirely. This can vary per agent, depending on usage and monitor configuration. For complete information, go to AI Trust Score: contributing factors.
Pass and fail criteria
What counts as a pass or fail check, and how the pass rate is calculated, depends on the AI integration.
Databricks AI
Databricks AI monitoring metrics run on their own monitoring schedule, separate from and more frequent than the regular Databricks synchronization schedule. Databricks does not guarantee how quickly LLM judge results become available after an agent interaction, so there can be a lag before these results appear in Collibra.
Each monitoring check corresponds to one LLM judge evaluation of one agent interaction. The result of each check is one of the following:
- Pass: The judge returned a passing verdict for the interaction.
- Fail: The judge returned a failing verdict for the interaction.
- No assessment: The judge ran but MLflow could not map its response to a pass or fail outcome. This typically happens because LLM judges are themselves non-deterministic. They can occasionally return unexpected output (extra text, an ambiguous answer, or a value outside the expected set) instead of the constrained verdict they were instructed to produce. A no-assessment result can also occur if the trace was present but the judge had not yet evaluated it at the time of sync.
The pass rate displayed in Collibra is calculated as follows:
Pass rate = passes ÷ (passes + fails + no-assessments)
This differs from the pass rate shown in Databricks, which excludes no-assessment results from the denominator. As a result, the pass rate in Collibra will always be equal to or lower than the pass rate shown in Databricks for the same judge and time period. The Collibra UI shows a tooltip on pass rate values to clarify this difference.
For example, if a judge evaluated 100 interactions and 90 passed, 2 failed, and 8 returned no assessment:
- Collibra pass rate: 90 ÷ (90 + 2 + 8) = 90%
- Databricks pass rate: 90 ÷ (90 + 2) ≈ 98%
For complete information about the interaction trace data this is based on, go to Integrated Databricks AI data.
AWS Bedrock AI
Each monitoring check corresponds to one invocation of a Bedrock agent alias. Collibra computes daily pass, fail, and error counts from invocation, client error, server error, and throttle data read from AWS CloudWatch. This is computed only for Bedrock agent aliases, not for model deployments that are not part of an agent.
Pass/fail and token consumption data for an agent alias only appear after synchronization has run at least once for that alias.
For complete information about the CloudWatch data this is based on, go to Integrated AWS Bedrock AI data.
Gemini Enterprise Agent Platform
Each monitoring check corresponds to one request handled by the agent. Collibra computes daily pass and fail counts from the request_count and response_code metrics that Google Cloud Monitoring records automatically for the agent, with no customer configuration required. This is a simpler signal than an LLM judge verdict: it reflects whether a request completed successfully, not whether the response met a quality standard.
Token consumption works differently for this integration. Google Cloud Monitoring has no built-in metric with per-agent token attribution, so token data only appears if the customer's own agent code emits a specific custom metric.
For complete information about this integration, go to Integrated Gemini Enterprise Agent Platform data.
Azure AI Foundry
Each monitoring check corresponds to one pass/fail result from Azure's continuous evaluation feature for the agent. These results are computed daily, on their own monitoring schedule, separate from the regular synchronization schedule.
This is only available for agents built through Azure's newer Agents API on Foundry project resources. Agents from the legacy Assistants API or a classic ML workspace are still cataloged, but without an AI Monitor asset, and therefore without pass/fail or token consumption data.
Token consumption is collected automatically from Azure Monitor (input and output token counts); it does not require customer self-instrumentation. Token usage can only be backfilled for roughly two months (61 days).
For complete information, go to Integrated Azure AI Foundry data.