AI Medicare/Medicaid Billing Fraud Pattern Detection Playbook
CMS Program Integrity has referred a home health agency for review after automated edits flagged a 340% spike in high-complexity evaluation and management codes across a 6-month period. The agency's 12 clinicians are billing at the 99215 level for 94% of visits — compared to a regional peer average of 18%. You have claims data, medical records for a 50-claim sample, and the agency's internal billing guidelines.
When to use this playbook
- Use this playbook when the decision looks like the situation above: CMS Program Integrity has referred a home health agency for review after automated edits flagged a 340% spike in high-complexity evaluation and management codes across a 6-month period.
- It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Medicare/Medicaid Billing Fraud Pattern Detection".
- Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.
What you'll need
- Claims data extract (12-month, all providers at the agency) - Medical records for 50-claim sample - Regional peer benchmark report - Agency's internal billing and coding guidelines
Attachments: Documents (Documents)
The Prompt
You are a healthcare fraud analyst and clinical reviewer conducting a Medicare billing integrity review for CMS Program Integrity. I am attaching: - Claims data extract (12-month, all providers at the agency) - Medical records for 50-claim sample - Regional peer benchmark report - Agency's internal billing and coding guidelines Work only from the attached source files. If a conclusion is not supported, say so. Produce: 1. Analyze the claims data for upcoding patterns — calculate each clinician's E&M code distribution and compare against regional peer benchmarks, identifying statistical outliers by provider, code level, and service date. 2. For the 50-claim sample, cross-reference the billed code level against the medical record documentation — identify all claims where the documentation does not support the billed complexity level under CMS 1995 or 1997 guidelines. 3. Assess whether the billing pattern is consistent with documentation-driven upcoding (inadequate records written to justify higher codes) or with clinical-driven upcoding (records appear legitimate but coding is inflated). 4. Review the agency's internal billing guidelines for language that explicitly encourages high-level code assignment, targets, or incentive structures tied to billing levels. 5. Calculate the estimated improper payment amount for the full 12-month period based on the sample error rate and produce a referral package for the OIG. Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.
What to expect
- Multi-model consensus on upcoding risk classification per provider
- E&M distribution analysis vs. peer benchmarks with outlier flags
- Documentation-to-code alignment scoring for 50-claim sample
- Billing guideline red flag register
- Estimated improper payment extrapolation and OIG referral package draft
Review before you act
- Validate this output against source files before relying on it: Analyze the claims data for upcoding patterns — calculate each clinician's E&M code distribution and compare against regional peer benchmarks, identifying statistical outliers by provider, code level, and service date.
- Validate this output against source files before relying on it: For the 50-claim sample, cross-reference the billed code level against the medical record documentation — identify all claims where the documentation does not support the billed complexity level under CMS 1995 or 1997 guidelines.
- Validate this output against source files before relying on it: Assess whether the billing pattern is consistent with documentation-driven upcoding (inadequate records written to justify higher codes) or with clinical-driven upcoding (records appear legitimate but coding is inflated).
- Validate this output against source files before relying on it: Review the agency's internal billing guidelines for language that explicitly encourages high-level code assignment, targets, or incentive structures tied to billing levels.
- Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
- Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
- Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.
Why compare models on this
For Medicare/Medicaid Billing Fraud Pattern Detection, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface multi-model consensus on upcoding risk classification per provider; e&m distribution analysis vs. peer benchmarks with outlier flags; documentation-to-code alignment scoring for 50-claim sample; billing guideline red flag register. Those are comparison artifacts — they only exist if more than one model runs. Threshold-splitting, sanctions hits, and exam-readiness calls are exactly where models diverge. Record the split and the human resolution.
See governed multi-model AI on your own prompt
Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.

