AI Data Retention Risks
M&A Targets Uploaded for Summarisation. Retained Thirty Days. Zero Retention: Not Configured.
6 min read · 26 August 2026 · AI governance
A private equity firm's investment professionals had adopted an AI document assistant from a specialised vendor , a tool that could quickly summarise lengthy prospectuses, extract key terms from contracts, and generate investment thesis drafts. Adoption had been rapid and informal , the vendor's product had been approved through a security review focused on the vendor's infrastructure security and data handling practices, and individual professionals had begun using it for day-to-day work within weeks of approval. Three months after adoption, the firm's general counsel conducted a review of the AI tool's usage patterns after a similar firm had disclosed a data handling incident with an AI vendor. The review identified that several investment professionals had uploaded documents to the AI assistant that contained material non-public information , preliminary term sheets, draft acquisition agreements, and board presentations for transactions that were not yet public. The vendor's privacy policy disclosed that inputs were retained for thirty days by default for service quality and debugging purposes. The thirty-day retention policy had not been disclosed prominently during the approval process, and the firm had not purchased the enterprise tier that provided zero-retention data handling. The documents were not retained by the vendor , they were retained by the vendor's LLM provider, OpenAI, on whose API the product was built. The firm's securities compliance obligations extended to material non-public information retention and access controls. The retention arrangement had not been assessed against those obligations.
What are AI Data Retention Risks, Really?
AI data retention risks are the security, compliance, and legal risks arising from the retention of data that users input to AI systems , specifically the risk that sensitive, regulated, or legally protected information uploaded to an AI assistant, submitted as an AI query, or processed by an AI pipeline is retained on the AI vendor's or AI provider's infrastructure longer than the enterprise's data handling obligations permit, or in ways that are not covered by the enterprise's data governance framework.
The tiered retention architecture problem is the specific risk driver. Most commercial AI products , particularly those built on third-party LLM APIs , offer different data retention configurations depending on the subscription tier. Standard tiers may retain inputs for varying periods , hours, days, or thirty days , for service improvement, debugging, and quality assurance purposes. Enterprise tiers typically offer zero retention or configurable retention, but require explicit configuration and higher subscription costs. The enterprise that evaluates and approves the standard tier may have approved a tier whose default retention configuration is incompatible with the enterprise's data handling obligations.
The informal adoption problem amplifies the retention risk. AI tools that are approved at the organisational level are frequently adopted informally at the individual level , employees begin using approved tools for their full workflow without specific guidance about what data categories are appropriate for the tool. An AI assistant approved for analysing publicly available research may be used to summarise board presentations, draft communications about confidential negotiations, and process documents containing regulated personal data. The retention configuration that was acceptable for the approved use case may not be acceptable for the actual use cases that develop organically.
The LLM provider retention dimension is the fourth-party risk. When a vendor's AI product is built on a third-party LLM API, the retention configuration at the LLM provider level may differ from the retention configuration at the vendor level. The vendor may retain inputs for seven days. The LLM provider may retain inputs for thirty days. The enterprise's data handling obligations attach to the data regardless of where it is retained. The LLM provider's retention , which the enterprise may not have specifically assessed , may be the binding constraint.
Why this matters
AI data retention risks matter for TPRM because the informal and rapid adoption of AI tools in enterprise environments creates systematic retention risk for sensitive data categories , MNPI, personal data, legally privileged information, regulated health data , that individual employees upload to AI assistants without specific awareness of the retention configuration. The enterprise's data governance framework may not have been applied to the retention configuration that governs those uploads.
Where most teams get this wrong
The most consistent failure is approving AI tools based on security assessment without specifically evaluating the retention configuration of the subscription tier being procured , and without establishing data classification guidance for users about which data categories are appropriate for the approved tool.
- Retention configuration of approved tier not specifically assessed
- LLM provider retention not assessed , fourth-party retention gap
- Data classification guidance not provided to users , all data uploaded by default
- Enterprise tier zero-retention not configured or required
- MNPI and regulated data upload not specifically addressed in AI tool governance
What good looks like
Mature AI data retention programmes specifically assess and configure retention at both vendor and LLM provider levels, require enterprise tier or zero-retention configuration for any AI tool that may process sensitive data categories, and provide users with explicit data classification guidance for approved AI tools.
- Retention configuration specifically assessed , vendor and LLM provider level
- Enterprise tier with zero-retention configured for sensitive data processing
- Data classification guidance for users , which data categories are permitted in which AI tools
- LLM provider retention assessed as fourth-party risk
- Sensitive data category audit , MNPI, PHI, legally privileged data handling in AI tools
Tooling
Enterprise AI Data Governance , Microsoft Purview with Copilot data governance, Nightfall for DLP in AI tools
Data loss prevention tools designed for AI platforms monitor inputs to AI tools and enforce data classification policies , preventing uploads of data categories that violate the enterprise's AI data governance policy. For TPRM practitioners, asking whether the enterprise has DLP controls applied to AI tool inputs , enforcing data classification policies at the point of upload , provides a specific retention risk reduction question.
Governance challenges
The governance challenge with AI data retention is the user behaviour gap. Enterprise-level retention configuration policies cannot prevent individual users from uploading data categories that are incompatible with the configured tier if users are not aware of the restrictions. The governance resolution is technical enforcement , DLP controls that block or flag uploads of sensitive data categories to AI tools , combined with user training.
- Require zero-retention configuration for AI tools that may process sensitive data
- Assess LLM provider retention as part of AI tool approval
- Implement DLP controls for AI tool inputs
- Provide data classification guidance for approved AI tools
- Audit AI tool usage for sensitive data upload patterns
If you are a small team
For your most widely adopted AI tools , document assistants, writing tools, code assistants , ask three retention questions: what is the retention period for inputs at the vendor level; what is the retention period at the LLM provider level; and have you configured the zero-retention enterprise tier? Then cross-reference those answers against the most sensitive data categories your employees are likely to upload. If the retention period exceeds the acceptable retention window for any of those data categories, you have identified an AI data retention risk that the security review did not surface.
- Ask for retention period at vendor level and LLM provider level
- Ask whether zero-retention enterprise tier is configured
- Cross-reference retention against sensitive data categories employees are uploading
- Implement DLP controls for AI tool inputs
What to require
Ask directly:
"What is the data retention period for inputs in our current subscription tier , at your infrastructure level and at your LLM provider's level , and do you offer a zero-retention enterprise tier, and are you willing to commit to that tier in our data processing agreement?"
Expect as evidence
- Vendor-level retention period for current tier
- LLM provider retention period disclosure
- Zero-retention enterprise tier availability
- DPA commitment to zero or configurable retention
A vendor whose data handling has been assessed should be asked specifically about retention configuration for the subscribed tier. Security assessment covers controls. Retention configuration determines what data is kept and for how long. Both are required.
How to evidence it
- Retention configuration documentation for vendor and LLM provider
- Enterprise tier zero-retention commitment in DPA
- Data classification guidance for AI tool users
- DLP control implementation for AI tool inputs
Key Takeaway
M&A targets. Financial projections. Strategic plans. Uploaded to an AI document summariser. Retained for thirty days. Not by the vendor , by the LLM provider the vendor built on. The vendor's privacy policy disclosed thirty-day retention. The employee had not read it. The enterprise had not configured zero retention. The security review had not specifically evaluated the retention configuration. MNPI on LLM provider infrastructure that was never assessed. AI data retention risk lives in the tier configuration, the LLM provider's terms, and the data categories that informal adoption allows. Zero-retention enterprise tier. DLP controls for sensitive data categories. Data classification guidance for users. Those three controls prevent the thirty-day retention of what should never have been uploaded.
Speak to It™
The term you nodded along to, explained in ninety seconds, so you can speak to it professionally. It is how most readers find these articles.
Join the Association