Stanford Confirms It: All 6 Major AI Vendors Train on Your Data by Default

← All posts

Stanford researchers did something no one in the AI industry wanted done. They read the privacy policies.

Not the marketing pages. Not the trust center FAQs. The actual legal documents governing how Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI handle user data. Their findings, published as a peer-reviewed paper (King, Klyman et al., arXiv:2509.05382), confirm what every compliance officer suspected but could not prove:

All six vendors train on user chat data by default.

Every conversation. Every document you paste into a prompt. Every client name, case number, financial figure, and medical record that enters a cloud AI system becomes training material for the next version of the model. Not by accident. By design.

Stanford University research paper — User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies

What Stanford found, vendor by vendor

The paper analyzed each vendor across five dimensions. Here is the full picture.

Vendor Trains by Default Retention Opt-Out Human Reviewers Child Data
Amazon YES Indefinite Available No YES
Anthropic YES Limited Switched to opt-out No No
Google YES Limited Available YES YES
Meta YES Indefinite Limited No YES
Microsoft YES Limited Available No No
OpenAI YES Indefinite Available YES YES
Source: King, Klyman et al. — Stanford University, 2025

The column that matters most: Trains by Default. All six. No exceptions.

Three vendors retain your data indefinitely: Amazon, Meta, OpenAI. Two employ human reviewers who read your actual conversations: Google and OpenAI. Four train on children's data.

The two-tier system

The paper reveals a structural divide in how AI vendors treat customers.

Enterprise Tier
Fortune 500 / Large Enterprise
✓ Training excluded by contract
✓ Zero-retention clauses
✓ Dedicated infrastructure
✓ Custom BAAs negotiated
Requires: dedicated procurement, outside counsel review, enterprise-scale commitment
Standard Tier
SMBs / Professional Firms / Consumers
✗ Data trains the model by default
✗ Retention may be indefinite
✗ Shared infrastructure
✗ Human reviewers may read chats
This is where your 50-attorney firm, 15-physician practice, or CPA office sits.

The protections that enterprise customers negotiate are not available to smaller organizations by default. You get the same model. You pay a lower price. The difference is subsidized by your data.

What "opt-out" actually means

Step 1
You sign up
Default: training ON. You accept terms without reading them.
Step 2
6 months pass
Client data flows through the system. Every query trains the model.
Step 3
You discover the opt-out toggle
You flip the switch. Future data excluded.
Result
6 months of data already embedded
You cannot un-train a model. The opt-out does not undo past training.

Anthropic is instructive. They launched Claude with opt-in training. Then in September 2025, the default flipped to opt-out. Users who had not explicitly opted in were now opted in unless they found the new setting. The policy you read when you signed up may not be the policy governing your data today.

Why enterprise-tier plans do not solve this

Dimension Enterprise Cloud Plan On-Premise (Dhakma Core)
Training protection Contractual promise Physical impossibility
Data location Vendor's global infrastructure Your server room
Access control Vendor's employee policies Your IT team
Subpoena exposure Vendor jurisdiction + yours Your jurisdiction only
Audit capability Vendor's compliance reports Full hardware + network logs
Cost model High minimum commitment Fixed monthly cost
Vendor acquisition risk New owner, new terms Hardware you own

A contractual promise is not the same as physical possession. For firms where the standard is physical control, enterprise cloud plans are insufficient.

What this means for regulated industries

Industry Governing Rule What the Stanford Paper Proves Consequence
Legal ABA Model Rule 1.6(c)
"Reasonable efforts" to prevent disclosure
Default training + human review = unauthorized disclosure Bar complaints, malpractice exposure
Medical HIPAA
Business Associate Agreements, minimum necessary
Training on PHI violates minimum necessary standard $141 to $2.13M per violation category per year
Finance SEC Reg S-P, Rule 30(a)
Fiduciary duty, safeguarding customer records
Indefinite retention of client financial data SEC enforcement, fiduciary breach
Defense CMMC 2.0 / NIST 800-171 / ITAR
Controlled Unclassified Information handling
CUI in training sets, potential foreign national access via human review ITAR violation, unauthorized export

The pattern is the same across every vertical. The Stanford paper documents default data practices that are incompatible with the regulatory frameworks governing these industries. If your vendor trains on your data by default, and you cannot demonstrate otherwise to an auditor, you have a compliance problem.

The paper's own recommendation

The Stanford researchers propose five solutions. Number five is direct:

Recommendation #5: On-device processing
"Processing sensitive data on local hardware eliminates the training risk entirely. If data never reaches the vendor's servers, it cannot be used for training, retained indefinitely, or reviewed by human employees."

The attack surface is the network connection. Remove the connection, and the cloud training, retention, and human review risks documented in this paper cease to exist.

This is what Dhakma Core is

A sovereign AI appliance. Built to implement exactly what Stanford recommends.

0
Bytes Transmitted
Models run locally. Data processed and indexed on your hardware. Physical air-gap switch.
100%
Audit Custody
You control the hardware, network, and logs. Demonstrate custody to any regulator or client.
Fixed
Monthly Cost
No enterprise-tier premium to earn a "we will not train on your data" promise.
How it works
Your Documents
Client files, records, communications
Dhakma Core
On your hardware. In your building.
AI-Powered Output
Analysis, classification, research
✗ No cloud connection    ✗ No third-party access    ✗ No training on your data

The Stanford paper recommends on-device processing as a privacy safeguard. Dhakma Core implements it as a product. The gap between academic recommendation and operational capability is closed.

The exit is available now

Before this research, the argument for private AI was based on inference. Now it is based on documented evidence: all six vendors train on your data by default, and the policies ensure you cannot fully undo it.

The proof is published. Peer-reviewed. From Stanford.

The question is no longer whether cloud AI vendors train on your data. The question is how long your firm will continue to accept the risk.

Your intelligence. Sovereign. Your data never leaves the room. The cloud is absent by design.

See investment details →

Schedule a confidential consultation →

Stay in the loop

Private AI insights for decision-makers. No spam. Unsubscribe anytime.

Have questions? Get in Touch →