Getting started: data stewards

The data products in the lakehouse. What the role can do and a first-week path through the screens you will use.

For: Data Steward · 1 min read

Your role on the platform

You own data products in the lakehouse: their quality, their documentation and who may use them. Stewards publish datasets through the intake pipeline and answer the access requests that name their data.

What you can do

  • Nominate and onboard lakehouse data products
  • Prepare, mask and publish dataset assets
  • Manage refresh cadence, quality and lineage in Dataplex
  • Run the intake pipeline for new datasets

Your first week

  1. 1Day 1: find your domain's datasets in the Datasets catalogue and check each Data Health Score and refresh cadence.
  2. 2Day 2: read the data classification and sensitivity tag guide, and check every one of your datasets is tagged correctly.
  3. 3Day 3: publish or refresh one dataset through Publish. Expect the PII scan and the metadata check to run first.
  4. 4Day 4: look at the Work queue for access requests on your data and practise the approve-with-conditions path.
  5. 5Day 5: review lineage for your most-used dataset, and the models that depend on it.

Screens you will use most

Datasets Publish My assets Work queue

Who to ask

The Data Governance Committee secretariat for classification disputes; Group Data & AI for Dataplex issues.

Read next

Was this helpful?