AI Playbook for Explainability for Regulatory Exam
A bank is under examination by the FDIC. The examiner has asked for an explanation of how the bank's AI-powered small business lending model makes decisions, why specific applicants were denied, and what fair lending testing has been conducted. The model is a gradient boosting ensemble from a third-party vendor.
When to use this playbook
- Use this playbook when the decision looks like the situation above: A bank is under examination by the FDIC.
- It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Explainability for Regulatory Exam".
- Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.
What you'll need
- Model vendor documentation (technical and business summary)
- SHAP value output for 50 sampled decisions (approved and denied)
- Disparate impact testing results (past 12 months)
- Adverse action notice template currently in use
- FDIC AI guidance and examination procedures
Attachments: Documents (Documents)
The Prompt
You are an AI governance specialist preparing a regulatory exam response for an FDIC examination of a bank's AI lending model. I am attaching: Work only from the attached source files. If a conclusion is not supported, say so. Produce: 1. Translate the SHAP values into plain-language explanations of the top 5 factors driving denials—language the examiner can understand without a data science background. 2. Assess whether the adverse action notices currently used accurately reflect the model's actual denial reasons or are generic and non-compliant. 3. Summarize the disparate impact testing results in exam-ready format: methodology, findings, and any remediation taken. 4. Identify the model documentation gaps: what the FDIC expects under SR 11-7 that the vendor has not provided. 5. Prepare the exam response narrative: what the model does, how it was validated, what ongoing monitoring is in place, and how the bank ensures fair lending compliance. Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.
What to expect
- Plain-language factor explanations from SHAP values
- Adverse action notice compliance assessment
- Disparate impact testing summary in exam format
- SR 11-7 documentation gap list
- Exam response narrative draft
Review before you act
- Validate this output against source files before relying on it: Translate the SHAP values into plain-language explanations of the top 5 factors driving denials—language the examiner can understand without a data science background.
- Validate this output against source files before relying on it: Assess whether the adverse action notices currently used accurately reflect the model's actual denial reasons or are generic and non-compliant.
- Validate this output against source files before relying on it: Summarize the disparate impact testing results in exam-ready format: methodology, findings, and any remediation taken.
- Validate this output against source files before relying on it: Identify the model documentation gaps: what the FDIC expects under SR 11-7 that the vendor has not provided.
- Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
- Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
- Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.
Why compare models on this
For Explainability for Regulatory Exam, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface plain-language factor explanations from shap values; adverse action notice compliance assessment; disparate impact testing summary in exam format; sr 11-7 documentation gap list. Those are comparison artifacts — they only exist if more than one model runs. Risk-tier assignments and 'high-risk system' calls vary with how a model reads a use-case description. Comparison exposes those classification fights before they reach an exam.
See governed multi-model AI on your own prompt
Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.

