RecommendationCritical riskComparison recommended

AI Playbook for Sepsis Risk Stratification Model Validation

A hospital has been using a commercial sepsis prediction model for 14 months. The vendor claims 87% sensitivity and 71% specificity. An internal data science team's validation study found sensitivity of 71% and specificity of 58% on the hospital's own patient population. The CMO needs a recommendation: keep, modify, or replace.

When to use this playbook

  • Use this playbook when the decision looks like the situation above: A hospital has been using a commercial sepsis prediction model for 14 months.
  • It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Sepsis Risk Stratification Model Validation".
  • Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.

What you'll need

  • Internal validation study results (AUC, sensitivity, specificity, PPV, NPV)
  • Vendor-claimed performance metrics and validation methodology
  • Hospital patient population demographics and case mix compared to vendor validation population
  • 14-month alert log with outcomes (confirmed sepsis, no sepsis)
  • Financial data: vendor contract cost, alert response workflow cost, estimated mortality impact

Attachments: Documents (Documents)

The Prompt

You are a clinical informatics specialist validating a commercial sepsis prediction model for a hospital CMO decision. I am attaching:

Work only from the attached source files. If a conclusion is not supported, say so.

Produce:
1. Assess the performance gap between vendor claims (87% sensitivity) and local validation (71% sensitivity): what explains the difference and is this clinically acceptable?
2. Calculate the clinical impact of the performance gap: at 71% sensitivity, how many sepsis cases are being missed per year that the 87% model would have caught?
3. Identify whether the performance gap is driven by patient population differences (case mix, demographics) or model degradation over time.
4. Model the cost-benefit of three options: keep the current model, recalibrate the thresholds, or replace with an alternative.
5. Tell me the CMO recommendation: which option, what is the projected mortality impact, and what is the business case.

Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.

What to expect

  • Performance gap analysis with root cause
  • Missed case calculation at current sensitivity
  • Population vs. degradation cause assessment
  • Cost-benefit model for three options
  • CMO recommendation with mortality and business case

Review before you act

  • Validate this output against source files before relying on it: Assess the performance gap between vendor claims (87% sensitivity) and local validation (71% sensitivity): what explains the difference and is this clinically acceptable?.
  • Validate this output against source files before relying on it: Calculate the clinical impact of the performance gap: at 71% sensitivity, how many sepsis cases are being missed per year that the 87% model would have caught?.
  • Validate this output against source files before relying on it: Identify whether the performance gap is driven by patient population differences (case mix, demographics) or model degradation over time.
  • Validate this output against source files before relying on it: Model the cost-benefit of three options: keep the current model, recalibrate the thresholds, or replace with an alternative.
  • Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
  • Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
  • Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.

Why compare models on this

For Sepsis Risk Stratification Model Validation, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface performance gap analysis with root cause; missed case calculation at current sensitivity; population vs. degradation cause assessment; cost-benefit model for three options. Those are comparison artifacts — they only exist if more than one model runs. Models disagree on whether an alert is noise, whether a death was sepsis-attributable, and whether a risk model is calibrated. Those disagreements belong in a morbidity-and-mortality style review, not an auto-implemented rule.

HealthcareModels and DocumentationRecommendationCriticalDocuments

See governed multi-model AI on your own prompt

Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.