Lakehouse Data Quality Agent
Runs after each lakehouse load. It reads the Dataplex data quality scan results for gold-layer tables, compares them with the previous 30 runs, works out which models and dashboards are affected, and raises an issue to the data steward with a proposed hold or release. It never edits data.
Tools
- Dataplex data quality scan results
- Dataplex lineage API
- BigQuery read (table profiles only)
- Gemini on Vertex AI (policy-rag-copilot)
- Issue tracker: data steward queue
Data access scope
- Read-only: data quality scan results, table profiles and lineage for gold-layer tables
- No row-level reads of Strictly Confidential tables
- Write: issues to the data steward queue and a hold flag proposal on the table
Guardrails
- Can propose a publishing hold; only the data steward can apply or lift it
- Never edits, deletes or backfills data
- Escalates to the Data Governance Committee secretariat if the same rule fails three runs in a row
Human-in-the-loop
- Data steward decides on every hold or release
- Monthly review of agent-raised issues by the Data Governance Committee
Orchestration
Every tool call is scoped by the declared data access above; the orchestrator cannot reach systems outside it. Owning business unit: Group Data & AI.
Live demo run
Watch the agent execute a real scenario step by step, every tool call, validation, and human checkpoint is traced and auditable. Typical run: ~8.0s.
Used in workflows
Retiring this asset would require these chains to be re-pointed first.
Owning business unit
Updated 2026-07-15. Run traces retained 24 months for audit under the platform governance policy.
Community · 0 threads
Questions, findings and requests from the business units that use this asset. Owners reply here; threads with upvotes surface to the owning team's inbox.