ReviewHigh riskComparison recommended

AI Playbook for Post-Deployment AI Performance Monitoring

A retail bank deployed an AI-powered loan pricing model 14 months ago. The model has not been re-validated since deployment. Market conditions, interest rates, and customer demographics have shifted materially. The model risk officer is concerned about model drift but has no monitoring framework in place.

When to use this playbook

  • Use this playbook when the decision looks like the situation above: A retail bank deployed an AI-powered loan pricing model 14 months ago.
  • It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Post-Deployment AI Performance Monitoring".
  • Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.

What you'll need

  • Model validation report from 14 months ago
  • Current model output data (pricing decisions, last 6 months)
  • Market condition changes since deployment
  • Current portfolio demographics vs. deployment-time demographics
  • OCC SR 11-7 ongoing monitoring requirements

Attachments: Documents (Documents)

The Prompt

You are a model risk officer designing a post-deployment AI performance monitoring framework for a retail bank loan pricing model. I am attaching:

Work only from the attached source files. If a conclusion is not supported, say so.

Produce:
1. Design the model drift detection metrics: what statistical tests (PSI, KS test, Gini coefficient) indicate material performance degradation?
2. Identify the monitoring frequency: how often should each drift metric be calculated and what are the threshold triggers?
3. Assess whether the current 14-month gap already constitutes a model risk violation under SR 11-7.
4. Design the model re-validation trigger criteria: what monitored metric breach level requires immediate re-validation vs. enhanced monitoring vs. model retirement?
5. Build the monitoring dashboard specification: what the model risk officer reviews monthly, what escalates to the CRO quarterly, and what triggers a board notification.

Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.

What to expect

  • Model drift detection metric design
  • Monitoring frequency and threshold triggers
  • SR 11-7 compliance gap assessment
  • Re-validation trigger criteria
  • Monitoring dashboard specification with escalation levels

Review before you act

  • Validate this output against source files before relying on it: Design the model drift detection metrics: what statistical tests (PSI, KS test, Gini coefficient) indicate material performance degradation?.
  • Validate this output against source files before relying on it: Identify the monitoring frequency: how often should each drift metric be calculated and what are the threshold triggers?.
  • Validate this output against source files before relying on it: Assess whether the current 14-month gap already constitutes a model risk violation under SR 11-7.
  • Validate this output against source files before relying on it: Design the model re-validation trigger criteria: what monitored metric breach level requires immediate re-validation vs. enhanced monitoring vs. model retirement?.
  • Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
  • Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
  • Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.

Why compare models on this

For Post-Deployment AI Performance Monitoring, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface model drift detection metric design; monitoring frequency and threshold triggers; sr 11-7 compliance gap assessment; re-validation trigger criteria. Those are comparison artifacts — they only exist if more than one model runs. Reconciliation protocols exist because models disagree. The playbook's job is to make disagreement inspectable, not to hide it behind a single blended answer.

AI Governance LayerLifecycle and AccountabilityReviewHighDocuments

See governed multi-model AI on your own prompt

Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.