Drift Watch Agent
Monitors every production model against the thresholds set at validation: PSI on the score, CSI on each feature, and performance decay once outcomes arrive. When a threshold is breached it opens a revalidation case straight away, rather than waiting for the annual cycle, and proposes a costed retrain to the business owner.
An agent executes the process end to end and logs what it did. No human is in the path, because there is no discretion to exercise and no consequence to carry.
- Acts for
- MLOps Engineer
- Action class
- advisory
- Decisions / 30d
- 1.2K
- Safe-state returns / 30d
- 2
Processes covered
- Drift watch and revalidation trigger
- Retrain proposal with costed compute
- Suspension and low-usage retirement review
Stages:Operate & charge backMonitor & revalidateSuspend & retire
Memory
Rolling 90-day telemetry window per asset; older aggregates archived to the audit store.
Safe state
Raises a 'monitoring gap' alert to MLOps and pauses drift verdicts for the affected model when metrics stop arriving.
Reviewer time per case: 0 min with this agent, 30 without.
Ownership and sensitivity
Who approves access
- Owner, Platform Engineering, Hafiz Othman
Business-sensitive. Models and data products scoped to named business units.
Entitlement per business unit, approved by the owner; conditions attach.
- Aggregated only
- No row-level customer data; aggregates with small-cell suppression.
Tools
- Vertex AI Model Monitoring (PSI, CSI, performance)
- Validation threshold reader
- Retrain cost estimator
- Revalidation case writer
- Owner notifier
Data access scope
- Read-only: production monitoring metrics and reference distributions
- Write: drift alerts, revalidation cases and retrain proposals
- No access to customer-level inputs or outputs beyond aggregated statistics
Guardrails
- Never starts a retrain or rolls back a version; the business owner accepts the cost first
- A breach on a High-risk model pauses new entitlements until the owner responds
- Alerts carry the threshold, the observed value and the window, never a bare warning
Human-in-the-loop
- The business owner accepts the retrain cost before any job starts
- Model Risk Management and the owner decide any suspension
Orchestration
Every tool call is scoped by the declared data access above; the orchestrator cannot reach systems outside it. Owning business unit: Group Data & AI.
Live demo run
Watch the agent execute a real scenario step by step, every tool call, validation, and human checkpoint is traced and auditable. Typical run: ~5.1s.
Used in workflows
Retiring this asset would require these chains to be re-pointed first.
Audit
Updated 2026-07-27. Run traces retained 24 months for audit under the platform governance policy.
Community · 0 threads
Questions, findings and requests from the business units that use this asset. Owners reply here; threads with upvotes surface to the owning team's inbox.