AI Model Security Risk Assessment Playbook
Your enterprise has deployed a customer-facing LLM chatbot that has access to a product database, customer account records, and order history. The security team wants a threat model before the system goes to 100,000 active users.
When to use this playbook
- Use this playbook when the decision looks like the situation above: Your enterprise has deployed a customer-facing LLM chatbot that has access to a product database, customer account records, and order history.
- It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Model Security Risk Assessment".
- Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.
What you'll need
- System architecture diagram (LLM, API integrations, data stores)
- Data classification for all connected data sources
- User authentication and authorization design
- Prompt engineering and system prompt (redacted for sensitive instructions)
- Vendor security documentation for the LLM provider
Attachments: Images (Images)
The Prompt
You are a cybersecurity architect conducting a threat model for an enterprise LLM chatbot deployment. I am attaching: Work only from the attached source files. If a conclusion is not supported, say so. Produce: 1. Identify the top attack surfaces: prompt injection (direct and indirect), data exfiltration via LLM responses, privilege escalation through the API layer, and model inversion risks. 2. Assess whether the current system prompt and guardrails are sufficient to prevent the LLM from leaking customer account data or order history to unauthorized users. 3. Identify the data flow risks: what customer data could be sent to the LLM provider's infrastructure and what are the contractual data handling obligations? 4. Recommend specific security controls: input validation, output filtering, rate limiting, anomaly detection for unusual query patterns. 5. Tell me what to monitor in production and what an LLM-specific incident response playbook should include. Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.
What to expect
- Threat model with ranked attack vectors
- System prompt and guardrail gap analysis
- Data flow risk assessment
- Recommended security controls by attack surface
- Production monitoring and IR playbook outline
Review before you act
- Validate this output against source files before relying on it: Identify the top attack surfaces: prompt injection (direct and indirect), data exfiltration via LLM responses, privilege escalation through the API layer, and model inversion risks.
- Validate this output against source files before relying on it: Assess whether the current system prompt and guardrails are sufficient to prevent the LLM from leaking customer account data or order history to unauthorized users.
- Validate this output against source files before relying on it: Identify the data flow risks: what customer data could be sent to the LLM provider's infrastructure and what are the contractual data handling obligations?.
- Validate this output against source files before relying on it: Recommend specific security controls: input validation, output filtering, rate limiting, anomaly detection for unusual query patterns.
- Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
- Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
- Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.
Why compare models on this
For Model Security Risk Assessment, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface threat model with ranked attack vectors; system prompt and guardrail gap analysis; data flow risk assessment; recommended security controls by attack surface. Those are comparison artifacts — they only exist if more than one model runs. Models disagree on blast radius, attribution confidence, and whether a vendor finding is theoretical or exploitable. Those disagreements mark where an analyst should slow down.
See governed multi-model AI on your own prompt
Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.

