A Federal Judge Just Ordered OpenAI to Hand Over 20 Million Conversation Logs

← All posts

A US federal judge ordered OpenAI to hand over 20 million ChatGPT conversation logs.

Full prompts. Full responses. Everything users typed in.

When I tell this to companies, the most common reaction is: "Yes, but we ticked the box. Our data isn't used for training."

That box has nothing to do with this.

The box you ticked does not do what you think it does

"Do not train on my data" means OpenAI will not use your prompts to improve their models. That is all it means.

It does not mean your data is not stored. It does not mean your data cannot be subpoenaed. It does not mean a US court will not order your AI provider to hand over everything your employees ever typed in.

That is exactly what happened. In January 2026, US District Judge Sidney Stein affirmed an order compelling OpenAI to produce 20 million ChatGPT conversation logs in In re: OpenAI, Inc. Copyright Infringement Litigation (SDNY). Not training data. Conversation logs. The actual prompts your people type every day.

OpenAI argued that producing the full sample would invade user privacy. The court disagreed. Three safeguards were deemed sufficient: reducing the sample size, de-identifying personal information, and a protective order governing access.

De-identification. That is the privacy protection standing between your client's M&A strategy and a courtroom.

The court accepted this as sufficient. But de-identification is an automated process. Can a vendor's script truly understand the sensitive context inside a legal brief, a patient's diagnostic notes, or a confidential M&A term sheet? Redacting a name is trivial. Redacting the meaning is not. And you are betting your firm's liability on their algorithm's accuracy.

The scale of what is sitting on those servers

38%
Of Employee AI Inputs
Contain sensitive data. Client names. Financial details. HR decisions. Strategic plans.
43%
Share Without Employer Knowledge
Including internal documents (50%), financial data (42%), and client data (44%).
Sources: CybSafe Annual Cybersecurity Attitudes Report 2025-2026; National Cybersecurity Alliance

All of it sitting on servers you do not control, in a jurisdiction where your GDPR rights become irrelevant the moment a court order lands.

Why "we turned off training" is not a governance strategy

What the checkbox actually controls
What it does
Tells the vendor not to use your prompts and responses as model training data.
What it does not do
✗ Prevent storage of your conversations
✗ Shield data from subpoenas or court orders
✗ Remove data already collected
✗ Stop human reviewers from reading logs
✗ Guarantee deletion on any timeline
✗ Apply across jurisdictions

"We turned off training" is one checkbox in a settings page. It says nothing about where your data lives, who can access it, or what happens when a court comes knocking.

What this means for regulated firms

Scenario Cloud AI On-Premise AI
Court orders conversation logs Vendor complies. Your data is produced. There is no vendor to subpoena. Data is on your hardware.
Regulatory audit requests AI data handling You produce a checkbox screenshot and hope. You produce hardware logs, network diagrams, air-gap verification.
Client asks where their data is processed "On Microsoft/OpenAI infrastructure. Somewhere." "On that server. In that room. Behind that door."
Vendor changes terms of service Your data governance changes with it. Your terms. Your hardware. Nothing changes.

For law firms: your opponent in litigation could gain access to your work product for a different client, simply because both matters were analyzed using the same cloud AI tool. The training opt-out creates no privilege protection against a discovery order.

For medical practices: a malpractice lawsuit could subpoena your AI vendor for every prompt a physician entered about a patient. Diagnostic reasoning, clinical notes, treatment considerations. None of it in your official medical record. All of it on someone else's server.

For financial firms: a regulatory audit from FINRA or the SEC could demand all AI conversations related to a specific trade or client, bypassing your compliance archives entirely and exposing informal strategy discussions your team assumed were private.

The only governance strategy that works

The only thing that protects you is making sure sensitive data never reaches a third-party server in the first place.

The principle
If your data never leaves your building, it cannot be subpoenaed from a vendor. It cannot be produced in discovery. It cannot appear in a court filing. There is nothing to find.

This is what on-premise AI infrastructure is built for. Not because cloud AI is bad at inference. It is excellent at inference. But inference is not the risk. Storage is the risk. Jurisdiction is the risk. The gap between what a checkbox promises and what a court order compels is the risk.

Dhakma Core processes everything on your hardware, in your building, behind your firewall. The physical air-gap switch disconnects the internet at the hardware level. Your conversations exist on your server and nowhere else. When a court orders a vendor to produce logs, there are no logs to produce. Your data never left.

0
Third-Party Servers
Your prompts, responses, and documents never leave your hardware.
Your
Jurisdiction Only
No vendor in another jurisdiction holding your data hostage to their legal obligations.
Full
Audit Trail
You control the logs. You control access. You demonstrate custody to any regulator.

The court order is the proof

Before this ruling, the argument for on-premise AI was based on hypothetical risk. Now it is based on judicial precedent. A federal court has ordered a major AI vendor to produce user conversation logs, and the privacy protections your employees relied on did not prevent it.

The question is not whether your AI vendor will be compelled to hand over data. The question is whether your sensitive data will be on those servers when it happens.

Your intelligence. Sovereign. Your data never leaves the room. The cloud is absent by design.

See investment details →

Schedule a confidential consultation →

Stay in the loop

Private AI insights for decision-makers. No spam. Unsubscribe anytime.

Have questions? Get in Touch →