AI Training Data Audit Playbook
A legal technology company trained a contract review AI on 2.4 million legal documents. A client raised concerns that their confidential contracts—submitted through the platform—were used as training data without authorization. The company's terms of service are ambiguous on this point.
When to use this playbook
- Use this playbook when the decision looks like the situation above: A legal technology company trained a contract review AI on 2.4 million legal documents.
- It is a fit when you have source files in hand and need a structured, reviewable analysis — not a generic chat answer about "Training Data Audit".
- Do not use it as a substitute for licensed, legal, clinical, or authorized official judgment in the domain.
What you'll need
- Training data manifest (document sources, dates, volumes)
- Current and historical terms of service (all versions)
- Client engagement agreement for the complaining client
- Data processing logs showing which client documents were used in model training
- Competing legal tech company data policies for comparison
Attachments: Documents (Documents)
The Prompt
You are an AI governance specialist conducting a training data audit for a legal technology company facing a client data misuse claim. I am attaching: Work only from the attached source files. If a conclusion is not supported, say so. Produce: 1. Determine whether the client's documents appear in the training data manifest and whether the ToS version in effect at the time of processing authorized training use. 2. Assess whether the training data use constitutes a breach of the client engagement agreement, an unauthorized use of confidential information, or a violation of attorney-client privilege protections. 3. Identify the exposure under applicable privacy law: CCPA, GDPR (if EU data is involved), and state professional responsibility rules. 4. Draft the client response: what the company will acknowledge, what remediation it will offer, and what policy change it will commit to. 5. Design the revised data governance policy: what client data can and cannot be used for training, how consent is obtained, and how opt-out is handled. Call out where independent models are likely to disagree, and list follow-up documents a reviewer should request.
What to expect
- Training data inclusion confirmation
- ToS authorization analysis
- Privacy law and professional responsibility exposure
- Client response draft
- Revised data governance policy framework
Review before you act
- Validate this output against source files before relying on it: Determine whether the client's documents appear in the training data manifest and whether the ToS version in effect at the time of processing authorized training use.
- Validate this output against source files before relying on it: Assess whether the training data use constitutes a breach of the client engagement agreement, an unauthorized use of confidential information, or a violation of attorney-client privilege protections.
- Validate this output against source files before relying on it: Identify the exposure under applicable privacy law: CCPA, GDPR (if EU data is involved), and state professional responsibility rules.
- Validate this output against source files before relying on it: Draft the client response: what the company will acknowledge, what remediation it will offer, and what policy change it will commit to.
- Confirm every cited figure, date, counterparty, or requirement against the attached originals — models compress and can drop a qualifier.
- Treat disagreement between models as a review item, especially on classification, materiality, and recommended next action.
- Do not authorize an operational, clinical, legal, credit, or enforcement action solely because the models agree.
Why compare models on this
For Training Data Audit, running the same attachments across independent models is useful because the hard part is classification and completeness, not fluency. The workflow is already designed to surface training data inclusion confirmation; tos authorization analysis; privacy law and professional responsibility exposure; client response draft. Those are comparison artifacts — they only exist if more than one model runs. Risk-tier assignments and 'high-risk system' calls vary with how a model reads a use-case description. Comparison exposes those classification fights before they reach an exam.
See governed multi-model AI on your own prompt
Compare GPT-5, Claude, and Gemini side by side, with human review and a decision record built in.

