AI Model Bias Audit & EU AI Act Risk Classification Playbook
Your AI governance team must produce a conformity assessment for a recidivism prediction model being considered for use in pretrial release decisions. The model was developed by a third-party vendor, trained on 10 years of historical justice data, and is classified as high-risk under the EU AI Act. Internal validation found a 23% disparity in false positive rates between racial groups. The assessment is due to your agency's AI review board in 21 days.
When to use this playbook
- Use this playbook when the decision looks like the situation above: Your AI governance team must produce a conformity assessment for a recidivism prediction model being considered for use in pretrial release decisions.
- It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Model Bias Audit & EU AI Act Risk Classification".
- Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.
What you'll need
- Source documents specified in the workflow
The Prompt
You are an AI governance auditor conducting a conformity assessment and bias audit for a high-risk AI system under EU AI Act and NIST AI RMF requirements. I am attaching: - Vendor model card and technical documentation - Internal validation report with disaggregated performance metrics - Training data provenance documentation - Intended use case description and deployment plan - Applicable regulatory requirements checklist Work only from the attached source files. If a conclusion is not supported, say so. Produce: 1. Assess the model card and technical documentation for completeness against EU AI Act Annex IV requirements — identify all required documentation elements that are missing, insufficient, or inconsistent with the validation findings. 2. Analyze the disaggregated performance metrics — characterize the 23% false positive rate disparity, assess its statistical significance, and evaluate whether it constitutes prohibited discrimination under applicable law. 3. Review the training data provenance documentation for issues that may have introduced or amplified bias — specifically, historical data reflecting past discriminatory practices, sampling imbalances, and label reliability. 4. Assess the deployment plan against NIST AI RMF GOVERN and MAP functions — identify gaps in human oversight, appeals mechanisms, and monitoring requirements for high-risk use in a justice context. 5. Produce a conformity assessment report with a risk classification determination, a list of non-conformities requiring remediation before deployment, and recommended conditions for any limited deployment authorization. Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.
What to expect
- Multi-model consensus on EU AI Act risk classification and conformity status
- Bias characterization report with statistical significance assessment
- Training data provenance red flag register
- NIST AI RMF gap analysis
- Draft conformity assessment report with model-agreement score on risk determination
Review before you act
- Validate this output against source files before relying on it: Assess the model card and technical documentation for completeness against EU AI Act Annex IV requirements — identify all required documentation elements that are missing, insufficient, or inconsistent with the validation findings.
- Validate this output against source files before relying on it: Analyze the disaggregated performance metrics — characterize the 23% false positive rate disparity, assess its statistical significance, and evaluate whether it constitutes prohibited discrimination under applicable law.
- Validate this output against source files before relying on it: Review the training data provenance documentation for issues that may have introduced or amplified bias — specifically, historical data reflecting past discriminatory practices, sampling imbalances, and label reliability.
- Validate this output against source files before relying on it: Assess the deployment plan against NIST AI RMF GOVERN and MAP functions — identify gaps in human oversight, appeals mechanisms, and monitoring requirements for high-risk use in a justice context.
- Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
- Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
- Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.
Why compare models on this
For Model Bias Audit & EU AI Act Risk Classification, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface multi-model consensus on eu ai act risk classification and conformity status; bias characterization report with statistical significance assessment; training data provenance red flag register; nist ai rmf gap analysis. Those are comparison artifacts — they only exist if more than one model runs. Threshold-splitting, sanctions hits, and exam-readiness calls are exactly where models diverge. Record the split and the human resolution.
See governed multi-model AI on your own prompt
Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.

