Skip to content
Technology Munch

Privacy  / Reference

What AI Tools Do With Your Text

Four things can happen to a prompt, and which ones apply depends on settings most people never open. The specific questions to ask of any tool.

Text you put into an AI tool leaves your device. What happens next varies enormously between products and tiers, and it is almost always described accurately in documents nobody reads.

The four things that can happen

Processing. The text is sent to a model and a response comes back. This always happens and is the point.

Retention. The conversation is stored. Duration varies from days to indefinitely. Storage is normal — it is what lets you see your history — and the duration and access controls are what matter.

Human review. Samples are read by people, for quality, safety and abuse detection. This is standard across the industry and it is the part that surprises people most.

Training. Your text is used to improve future models. This is the one with the most durable consequences and the one most often controlled by a setting.

The tier pattern

Consumer free tiers commonly default to training on inputs, with an opt-out available in settings if you know to look.

Consumer paid tiers vary. Some default to no training, some do not.

Business and enterprise tiers generally commit contractually to no training, shorter retention, and administrative controls. This is frequently the main reason to pay for them.

API access usually defaults to no training, which is a meaningful difference from the consumer product of the same company.

Do not assume the tiers behave the same. The same brand can have opposite defaults across products.

The questions to ask of any tool

Is training on my input the default, and can I turn it off?

Does turning it off apply retroactively or only to future conversations? Usually only future.

How long is data retained, including after I delete a conversation? Deletion from the interface and deletion from backups are different.

Who can access it? Employees, contractors, sub-processors.

Where is it processed? Relevant for regulated data and cross-border rules.

What happens if the company is acquired or fails? Data is an asset in both cases.

For a tool that is not from a company you can identify, assume the worst on every question.

What people put in that they should not

The gap between policy and practice here is large.

Client and customer information. Names, contracts, case details. For anyone under professional obligations — legal, medical, accounting — this is a confidentiality question, not a preference.

Unpublished work. Manuscripts, research, code, business plans.

Credentials. API keys, passwords, tokens. These appear in pasted configuration files constantly.

Personal health information, their own or others'.

Internal financial data.

Other people's personal information. Pasting an email thread shares the other participants' data, and they did not agree to it.

A useful test: would you be comfortable if a contractor at that company read this? That is a realistic description of what human review means.

The features that add exposure

Memory. Facts stored across conversations. Review what is stored occasionally; it accumulates more than you expect.

Connectors and integrations. A tool granted access to your email, files or calendar is processing far more than what you type. The permission scope is what to examine.

Browser extensions. Many request permission to read page content on every site. Some are straightforwardly data collection.

Team and workspace sharing. Conversations visible to colleagues or administrators.

Anything agentic, which reads and acts on material you did not review.

Practical arrangement

One paid tool with training disabled for anything sensitive, chosen from a company you can identify.

Free tiers for everything non-sensitive, which is most of what most people do.

Local models for anything genuinely confidential, where the text never leaves the machine.

Check the settings when you sign up, once, and again after any major product update. Defaults change.

Never paste credentials, in any tier.

That covers the realistic risk without requiring anyone to become a privacy specialist.

Setting up one tool properly

Rather than auditing everything, configure one tool to be the place sensitive work goes.

Choose a paid tier from a company you can identify, with a published data policy.

Disable training on inputs, and confirm it applied.

Set the shortest retention the product allows.

Turn off memory unless you specifically want it, and review what it already stored.

Do not connect it to your email, files or calendar unless you have a reason and have read what the connector accesses.

Do not install its browser extension on your main profile.

Then use it for anything confidential and nothing else, and use free tiers for the rest.

One configured tool plus a clear sorting rule handles the realistic risk without requiring anyone to read a dozen privacy policies.