AI Governance

The model risk and AI governance framework the platform runs on, and why a regulated bank sharing AI across business units requires it.

AI governance is how the Bank decides what AI may be built and used, under what conditions, and who is answerable for the outcome.

The rules, the roles, and who is answerableGOVERNruns through all threeKnow what it is, and who it reachesMap01Test it against the risks identifiedMeasure02Act on what the tests find, and keep actingManage03GovernThe rules, the roles, and who is answerableModel risk policy and risk appetiteThree lines of defence, decision rights per roleA named owner on every modelMapKnow what it is, and who it reachesIntended and out-of-scope useModel tier, data classification and lineageWhich customers the decision affectsMeasureTest it against the risks identifiedAccuracy per segment, not just overallFairness tests at validation and retrainPSI, CSI and decay, measured continuouslyManageAct on what the tests find, and keep actingGates that block, not warnPurpose-bound accessRetrain, suspend, retire

Govern is not the first step. It runs through the other three, which is why governance cannot be a document somebody signs at the end.

What it has to cover

People, process, technology and data. Each fails differently, and a gap in any one is enough.

People

7

roles with distinct decision rights

Who decides, and can they actually decide?

  • Decision rights defined per role, not per person
  • Three lines of defence: developers, Model Risk Management, Internal Audit
  • Validators with the evidence, the time and the authority to say no
  • A named business owner on every model before it is registered

Without it: Validation becomes a signature. Someone approves forty models a quarter and nothing is really challenged.

Process

27

governed sub-processes

Where in the life of a model does governance apply?

  • Nine stages, from onboarding a team to suspending or retiring a model
  • Every stage has named sub-processes with a trigger and an artefact
  • Agents run what is checkable, people decide what is consequential
  • Gates that block a release, not warnings that get dismissed

Without it: Governance happens once, at the end, when every expensive choice has already been made.

Technology

9

governance agents in the path

Where do the controls actually live?

  • Controls run in the platform, not in a policy document
  • Sandbox isolated with no egress; production compute in Google Cloud Malaysia
  • Vertex AI model registry, agent manifests and a declared tool surface
  • An immutable audit record written by the system, not by hand

Without it: The rules exist but nothing enforces them, so compliance depends on who remembered.

Data

100%

catalogue assets carry a classification

What is it trained on, who may use it, and where can it go?

  • Classification set at source in Dataplex and inherited by every derived asset
  • PII and customer-identifier scanning on every load and every GenAI prompt
  • Purpose-bound, time-bound entitlements rather than standing access
  • Customer data kept in Google Cloud Malaysia; held-out validation samples owned by Model Risk Management
  • Lineage from a live endpoint back to the data it learned from

Without it: A model is governed while the dataset underneath it is not, so the restriction is lost the moment it is trained on.

Each pillar fails differently, and a gap in any one is enough. Rules nobody owns are unenforced, owners with no tooling leave no evidence, and all three are undone by a dataset that was never classified.

The AI lifecycle, and the agents inside it

Nine stages. Pick one to see its sub-processes, which agents run them, and where a person still decides.

Build & validate

A model enters the registry only after intake triage, automated assurance and independent validation by Model Risk Management.

4 unattended2 with a person
  • Intake triage, classification and model tiering

    Agent runs itTrust

    Triggered by: A model or dataset is submitted with its manifest

    Intake Triage Agent

    Leaves: Classification, model tier, validation path and routing record

  • Automated assurance

    Agent runs itTrust

    Triggered by: Intake routes a conforming submission

    Assurance Agent

    Leaves: Check report: performance, fairness, explainability, security scan, PII in prompts, licence scan

  • Validation pack assembly

    Agent prepares, person decidesTrust

    Triggered by: A check report is attached to the case

    Validation Case AgentModel validator reads the pack before writing the validation opinion

    Leaves: Validation pack with recommendation, basis and precedent

  • Independent validation and model approval

    Person decidesTrust

    Triggered by: A validation pack reaches the Model Risk queue

    Validation Case AgentModel validator, Model Risk Management signs; the business owner accepts the use

    Leaves: Validation opinion and approval with tier, conditions and revalidation date

  • Chargeback rate card

    Agent runs itValue exchange

    Triggered by: An asset is approved for reuse

    Chargeback Agent

    Leaves: Rate from the published rate card on the asset card, in MYR

  • Model registry versioning

    Agent runs itAssets

    Triggered by: A new version is registered

    Intake Triage Agent

    Leaves: Version history with digest, validation reference and consumer notice

27 governed sub-processes across 9 stages, with 9 agents in the path.

What the framework holds a system to

Eight standards. A system is measured against these, not against whether it was delivered on time.

Valid and reliable

It does what it claims, repeatably

Fair

Comparable customers get comparable credit outcomes

Accountable

A named business owner and validator own the model

Transparent

Its use is disclosed to the customer where it matters

Explainable

The reasons can be given to the customer and the RM

Privacy-preserving

PDPA purpose respected, banking secrecy kept

Secure and resilient

Holds under attack and under load

Safe

No unfair harm to a customer's money or access to credit

Obligations scale with consequence

The tier is decided by what the system affects, not by how it was built. It is what makes human review mandatory.

MINIMAL

No effect on a customer outcome; low materiality (typically Tier 3)

Requires: Model card, classification, audit trail, proportionate review

Human review: Not required

LIMITED

Informs a decision about a customer, which a person still makes (typically Tier 2)

Requires: Independent validation, fairness tests, accuracy per segment, bounded access

Human review: At the decision

HIGH

Decides or materially shapes credit, pricing or a customer outcome (typically Tier 1)

Requires: Full independent validation, explainability, champion–challenger, complaint route

Human review: Mandatory, and enforced

How a model gets validated

A submitted test report can be written rather than run. Each level removes the model developer from the evidence.

Too restrictive

Nothing gets validated in time. Business units build in spreadsheets again and the platform has no reason to exist.

The balance we are aiming at

Effort has to match materiality. A summariser over internal circulars and a model that declines an SME loan cannot carry the same burden of proof, and treating them alike fails in both directions at once.

Too lenient

The first failure reaches customers and the regulator, and the platform loses the trust it needs to be adopted at all.

What is being validated, and how far does it reach

Does the output affect a customer's money, credit, pricing or access to a product?

Yes. So how far does it reach?

Data classification does not change the level. It changes how the testing is done.

L3

Independently validated

Developers cannot tune to a sample they cannot see.

Developer submits
The model and its documentation. Not the evaluation set
Platform does
Scores it on an out-of-time sample Model Risk Management holds back. Discrimination, calibration and stability reported per segment, not only overall
Human in the loop
A named validator signs the validation opinion before the model is used
Revalidation
Every 2 years, or on each material change (typically Tier 2)

L3 and L4 depend on held-out, out-of-time samples that Model Risk Management owns and the developer has never seen. Building and refreshing those per portfolio is a standing duty of the second line, not of the developer.

If the data is above Internal, at any level

  • Validation runs inside a restricted project in Google Cloud Malaysia. The model does not leave, and neither does the sample
  • A PDPA purpose check is completed before the validation opinion can be signed
  • The data owner approves the validation use of customer data, separately from the developer
  • Results are reported per segment, but the validation records stay Strictly Confidential

Why a Group AI platform requires it

Four reasons that sharing AI across a regulated bank creates, each with what it has already cost elsewhere.

1

A shared platform concentrates model risk by design

1 : 6

one flaw, every consumer

The value of a Group platform is that one model serves many business units. The same property means one flaw reaches every consumer at once, and without lineage the resulting incidents look unrelated.

Six business units on six separate scorecards produce six independent errors. Six business units reusing one model produce the same error six times.

2

Models get used beyond the context they were validated for

A model is validated for one portfolio, one product and one decision, then found in the catalogue by a team with a different portfolio. Nothing in the model signals that the second use was never validated.

A retail credit scorecard validated on salaried applicants, applied to self-employed SME owners whose income looks nothing like the development sample.

3

Failures are systematic, and surface late

2019

a credit-limit algorithm under investigation

Model errors are not randomly distributed. They concentrate in particular segments, they read as objective because the output is a score, and they surface through complaints that the most affected customers are least likely to file.

In 2019 a card issuer's credit-limit algorithm was publicly accused of giving women lower limits than their husbands. The New York regulator found no unlawful discrimination, but the issuer could not explain individual limits to the customers who asked.

4

Credit decisions have to be defensible

Under BNM's expectations and the Bank's own policy, the Bank must be able to state who approved the model, on what evidence, and why a customer was declined. A log reconstructed a year later is not a decision record.

A declined SME loan that cannot be traced to a validated model version and a named approver is not a decision the Bank can defend to the customer or to an examiner.