Is Your Business Data Training ChatGPT? Here's How to Stop It
- Jul 27
- 4 min read
Someone on your team pastes a client quote into ChatGPT to tighten the wording. Someone else asks Gemini to summarize a service agreement. Your bookkeeper drops a payroll file into Claude to chase down a variance. Nobody did anything reckless. And yet, depending on which account they used, all three documents may have just been added to a training corpus, with a five-year retention window in one case.
This is not a doomsday scenario. It is the default behaviour of the consumer versions of these tools. The good news: fixing it takes about twenty minutes per platform, and most of the work is one structural decision rather than a hidden toggle.
The rule that explains everything: the account type decides
There is an expensive misconception going around: that paying for a subscription protects your data. It does not. Paying twenty dollars a month for a personal Plus, Pro or Max account changes nothing about training. What changes everything is the legal nature of the account.

In the right-hand column, the protection is contractual and on by default. OpenAI is explicit: it does not train on inputs or outputs from business products or the API. Anthropic specifically carves out Claude for Work, the API, Bedrock, Vertex AI and Claude for Education from its consumer training policy. Google applies separate terms to Workspace accounts.
In the left-hand column, you are the product until proven otherwise.
The setting protects your data. The contract protects you.
The practical conclusion is simple, and it is the one recommendation here I make without hedging: if your company uses generative AI for anything beyond curiosity, move to business accounts. A Business plan costs roughly double a personal one. It is the cheapest insurance policy in your tech stack.
The settings, platform by platform
That said, migration takes time, and your people almost certainly already have personal accounts. Here is what to do today.
ChatGPT (OpenAI)
In settings, open Data Controls and turn off the option that allows your content to improve the model for everyone. OpenAI's privacy portal at privacy.openai.com also offers a formal "do not train on my content" request that covers your whole account.
Two traps are worth knowing about:
Temporary Chat writes nothing to history and does not feed training. Make it the default reflex for any conversation touching a client file.
Even after opting out, giving a thumbs up or thumbs down makes that entire conversation eligible for training again. Tell your team: the reflex of rating an answer cancels the protection on that exchange.
Claude (Anthropic)
Go to your privacy settings at claude.ai/settings/data-privacy-controls and turn the model improvement setting off. The issue here is not only training, it is retention: with the setting on, Anthropic keeps conversations and coding sessions for five years. With it off, you drop back to thirty days.
Five years is longer than most of your client contracts will live. It is also well past what you could reasonably defend to a regulator, or to a client who asks where their information went.
Gemini (Google)
The setting is called Gemini Apps Activity, or Keep Activity depending on the interface. It is on by default, with an eighteen-month auto-delete window. Turn it off in your Google Account settings.
Here too there is a detail that stings: conversations already sampled for human review are kept for up to three years, disconnected from your account, and that retention survives deleting your chats. Turning the setting off today is not retroactive. What has left has left.

Why regulators care, not just your IT person
The moment personal information about a client, an employee or a supplier goes into an AI tool, you are disclosing personal information to a third party, usually one located outside your jurisdiction. Under Quebec's Law 25, and in similar terms under GDPR and other modern privacy regimes, that triggers three concrete obligations:
Run a privacy impact assessment before the cross-border transfer.
Govern the disclosure with a written agreement.
Be able to demonstrate that the level of protection is adequate.
A personal ChatGPT account whose conversations feed model training fails that test. There is no agreement, no control over secondary use, and no reliable deletion mechanism. An Enterprise account with a data processing addendum passes.
The weak link is not the platform, it is the missing rule
In the engagements I run with small and mid-size companies, the problem is almost never technical. It is that no clear instruction exists, so every employee improvises. The cautious ones use nothing and lose time. The enthusiastic ones paste everything, including what they should not have.

The fourth point deserves extra weight. When a workflow runs through an API inside n8n or Make, it automatically inherits the API terms, the ones that exclude training by default. Manual copy-paste into a web interface is simultaneously the riskiest and the least efficient path. Moving from manual use to automated use improves your compliance and your throughput at the same time, which is rare.
What to take away
Your data is not being stolen. It is being handed over, by default, because nobody changed a setting or picked the right kind of account. The fix is within reach this week: check the three settings above across every account in your organization, decide to move to business plans, and write your one-page rule.
If you want someone to inventory what is actually flowing out of your company, document your cross-border transfers, and replace the copy-paste habit with clean automated workflows, that is exactly the work Neurotek AI does. Clear intention, structure, clean execution. No magic, just systems that hold.



Comments