Risk AssessmentCritical riskComparison recommended

AI HMDA Data Integrity Audit Playbook

A bank's HMDA LAR is due in 90 days. An internal data quality review found that 340 records have missing or implausible data fields. The bank was assessed a $1.2M civil money penalty for HMDA violations 3 years ago and cannot afford another examination finding.

When to use this playbook

  • Use this playbook when the decision looks like the situation above: A bank's HMDA LAR is due in 90 days.
  • It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "HMDA Data Integrity Audit".
  • Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.

What you'll need

  • Preliminary HMDA LAR (all originations and applications, current year)
  • Prior year LAR for comparison
  • CFPB HMDA edit check specifications
  • The 340 flagged records with identified data issues
  • Prior examination findings

Attachments: Spreadsheets (Spreadsheets)

The Prompt

You are a compliance officer conducting a HMDA data integrity audit ahead of a filing deadline. I am attaching:

Work only from the attached source files. If a conclusion is not supported, say so.

Produce:
1. Run all CFPB required syntactical, validity, quality, and macro edits on the LAR and identify all records that would fail and require resubmission.
2. For the 340 flagged records, classify each error by severity: reporting error (wrong data), missing data, or systemic error (same field wrong across many records).
3. Identify whether any errors follow a pattern that could trigger a fair lending referral—errors concentrated in minority-applicant records are a red flag.
4. Prioritize the records that must be corrected before filing and those that can be submitted with a resubmission plan.
5. Tell me the resubmission procedures and what to document to demonstrate good-faith remediation to the CFPB.

Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.

What to expect

  • Full CFPB edit check results
  • Error classification by severity and type
  • Fair lending referral risk assessment
  • Correction priority list
  • Resubmission procedure and good-faith documentation plan

Review before you act

  • Validate this output against source files before relying on it: Run all CFPB required syntactical, validity, quality, and macro edits on the LAR and identify all records that would fail and require resubmission.
  • Validate this output against source files before relying on it: For the 340 flagged records, classify each error by severity: reporting error (wrong data), missing data, or systemic error (same field wrong across many records).
  • Validate this output against source files before relying on it: Identify whether any errors follow a pattern that could trigger a fair lending referral—errors concentrated in minority-applicant records are a red flag.
  • Validate this output against source files before relying on it: Prioritize the records that must be corrected before filing and those that can be submitted with a resubmission plan.
  • Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
  • Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
  • Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.

Why compare models on this

For HMDA Data Integrity Audit, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface full cfpb edit check results; error classification by severity and type; fair lending referral risk assessment; correction priority list. Those are comparison artifacts — they only exist if more than one model runs. Control specifications, geographic market definitions, and 'similarly situated' calls routinely diverge. Model disagreement is a signal to re-cut the file review, not to publish a single p-value.

Fair LendingRedlining and HMDA DataRisk AssessmentCriticalSpreadsheets

See governed multi-model AI on your own prompt

Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.