Risk AssessmentHigh riskComparison recommended

AI Past Performance Scoring Assessment Playbook

Your firm is submitting 4 past performance references for a Navy IT services recompete. The RFP requires references relevant in scope, size ($10M+), and complexity. Two of your references are over 3 years old. One reference involves a subcontract role. You are uncertain which references score highest under the Navy's evaluation criteria.

When to use this playbook

  • Use this playbook when the decision looks like the situation above: Your firm is submitting 4 past performance references for a Navy IT services recompete.
  • It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Past Performance Scoring Assessment".
  • Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.

What you'll need

  • 4 past performance reference summaries (contract value, scope, role, recency, customer contact)
  • RFP Section L and M past performance requirements and evaluation criteria
  • CPARS ratings for each reference (where available)
  • Teaming partner past performance references (2 additional options)
  • Navy's past performance confidence assessment methodology

Attachments: Documents (Documents)

The Prompt

You are a proposal manager selecting and optimizing past performance references for a Navy IT services proposal. I am attaching:

Work only from the attached source files. If a conclusion is not supported, say so.

Produce:
1. Score each of the 6 available references (4 firm + 2 teaming partner) against the Navy's evaluation criteria: recency, relevance (scope and size), and performance quality.
2. Identify the highest-risk reference and tell me whether to include or exclude it—a weak reference can lower the overall confidence rating.
3. Draft the past performance narrative for the top 2 references: how to frame scope, size, and outcome in language that directly maps to the Navy's evaluation language.
4. Assess whether the subcontract reference will receive full consideration or partial credit under the Navy's methodology.
5. Tell me the optimal 3-reference combination and the order in which to present them.

Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.

What to expect

  • Reference scoring matrix against Navy criteria
  • Risk assessment for weak reference
  • Past performance narrative drafts for top 2 references
  • Subcontract reference credit assessment
  • Optimal reference combination and presentation order

Review before you act

  • Validate this output against source files before relying on it: Score each of the 6 available references (4 firm + 2 teaming partner) against the Navy's evaluation criteria: recency, relevance (scope and size), and performance quality.
  • Validate this output against source files before relying on it: Identify the highest-risk reference and tell me whether to include or exclude it—a weak reference can lower the overall confidence rating.
  • Validate this output against source files before relying on it: Draft the past performance narrative for the top 2 references: how to frame scope, size, and outcome in language that directly maps to the Navy's evaluation language.
  • Validate this output against source files before relying on it: Assess whether the subcontract reference will receive full consideration or partial credit under the Navy's methodology.
  • Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
  • Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
  • Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.

Why compare models on this

For Past Performance Scoring Assessment, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface reference scoring matrix against navy criteria; risk assessment for weak reference; past performance narrative drafts for top 2 references; subcontract reference credit assessment. Those are comparison artifacts — they only exist if more than one model runs. Models split on whether a requirement is mandatory, how to score a differentiator, and protest likelihood. Those splits should be resolved before color-team review, not after submission.

Government RFPPerformance and TeamingRisk AssessmentHighDocuments

See governed multi-model AI on your own prompt

Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.