← all articles

What a chatbot keeps after you close the tab

People type things into a chatbot that they would never type into a search box. A search box feels like a public place. A chat window feels like a conversation, so the guard comes down: medical symptoms, salary numbers, the actual text of an argument with someone, a draft resignation letter with the company name still in it.

So let’s go through what actually happens to that text. The answer has more moving parts than most people expect, and the moving parts are where the surprises live.

The conversation is never just the words

When you send a message to a hosted assistant, the text goes to the provider’s servers, gets processed, and gets stored against your account. Stored is the important word. The conversation sits in a history you can scroll back through, which means it exists somewhere as data, attached to an identity, with a timestamp on it.

Alongside the text there’s the usual envelope: your account identifier, the IP address you connected from, rough device and browser details, when you sent it, how long the session ran. None of that is unique to chatbots. It’s what every web service records. It does mean a conversation is never just the words in it.

Retention and training are two different questions

Two questions get mashed together constantly. One is whether your text is retained. The other is whether it’s used to train future models. Those are different, and the answer is often different for each.

Retention is close to universal. A service that shows you your chat history is storing your chat history, by definition. Training is the part that varies, and it’s usually a setting. On consumer tiers, several major providers default to using conversations for model improvement and give you a toggle to turn it off. The toggle is real, and it’s often buried two levels into a data controls menu that nobody opens.

Business and enterprise tiers usually flip this. Paid API access and workplace deployments typically come with a contractual commitment not to train on customer content, because no company would sign otherwise. That’s the clearest split in the whole landscape: the same model, reached through a different door, comes with a different promise attached.

Someone might actually read it

Providers sample conversations for safety and quality work, which means a small number of real conversations get read by real people. This is disclosed, it’s a normal part of running these systems, and it’s also the part that most surprises people when they hear it. The sampling is small. It isn’t zero.

Temporary chat modes: better, not gone

Temporary or incognito chat modes deserve a careful look, because the naming oversells them. A temporary chat generally won’t appear in your history and won’t be used for training. It doesn’t usually mean the text vanishes on send. Providers commonly hold those conversations for a window, often around thirty days, for abuse monitoring, then delete them. That’s a meaningful improvement over the default. It isn’t the same as never having existed.

Voice adds a second recording

Voice deserves its own line. When you talk to an assistant, the audio gets transcribed, and depending on the service both the transcript and the audio itself may be retained. Audio is a different category of data from text: it carries whoever else was in the room, background sound that places you somewhere, and a voiceprint that belongs specifically to you. Controls usually sit in the same data page as everything else, sometimes as a separate switch for voice recordings. It’s worth checking if you use voice mode much, because people talk more loosely than they type.

Memory turns one conversation into a profile

Memory features are the newer wrinkle. An assistant with memory carries facts across conversations, so something you mentioned months ago can surface in a chat today. That’s genuinely useful, and it changes the shape of the risk, because the sensitive thing is no longer confined to one conversation you could delete. It’s been promoted into a profile. Most implementations let you view what’s been remembered and remove individual entries, and that page is worth reading once, because the model’s idea of what mattered about you is often not what you would have picked.

Connectors and the permission you didn’t mean to grant

Connectors are where this gets properly interesting. An assistant plugged into your email, your files, or your calendar can read them to answer questions. The permission you grant at setup is frequently much broader than the task you had in mind. You connect a drive so it can summarise one document, and the grant covers the drive. Read what the consent screen actually says rather than what the marketing said, and check the connected apps list in the underlying account rather than in the chatbot, since that’s where the grant actually lives and where you can revoke it.

Most assistants let you turn a conversation into a shareable link. That link is public to whoever holds it, which is the entire point of the feature. What caught people out in 2025 was that some shared pages became findable through ordinary web search, because a public URL with nothing restricting it is exactly what a crawler exists to pick up. Providers reacted and changed the defaults, but the lesson outlives the specific incident: a share link is publishing. If the conversation has your employer’s name in it or your own medical history, publishing isn’t what you meant to do. Send a screenshot of the part you actually wanted to show someone instead.

The AI feature you’re using might not be the provider’s at all

Third party apps are the bigger blind spot. A lot of the AI features you meet in daily life aren’t the model provider’s own product at all. A writing app, a browser extension, a note tool, a support widget on some company’s website, all calling somebody’s API in the background. In that arrangement, the app developer sits between you and the model. They see your text, they decide what to log and for how long, and their privacy policy is the one that governs the whole thing. The provider’s enterprise commitment not to train on customer content may well apply to the developer’s account, which tells you precisely nothing about what the developer themselves keeps on their own servers. When an AI feature shows up inside a tool you already use, the question is whose infrastructure it runs on and what that party’s policy says.

Deletion has a schedule, and training doesn’t reverse

Deletion is the part people most misunderstand. Deleting a conversation removes it from your view and starts a removal process on the provider’s side, usually completing within some stated number of days. Backups take longer to age out. And if the content was already used in a training run, deleting the source conversation doesn’t reach into a model that has already been trained. That’s not a loophole anyone is hiding. It’s just how the technology works, and it’s the reason the training toggle matters more than the delete button.

Courts can extend the schedule

There’s a legal dimension too, and 2025 made it concrete. In the copyright litigation between the New York Times and OpenAI, a court ordered the preservation of output log data that would otherwise have been deleted on the normal schedule. I’m not going to characterise the merits of that case. The mechanical point stands on its own: a company’s stated retention period is what happens absent a legal reason to keep things longer, and a court can supply that reason. Every service in every industry works this way. Chatbots are just newer at it.

Work chat is a work record

Work is its own category. If your employer has deployed an assistant, the workplace admin usually has visibility into usage, and the conversations may sit inside the same compliance and retention regime as your email. That regime exists for legitimate reasons, and it also means a chat window at work is a work record. People forget this constantly, because the interface looks like the consumer one they use at home.

What to actually do

None of this calls for anything dramatic.

Find the data controls page on whatever assistant you use and look at the training toggle. Decide deliberately rather than by default. While you’re there, look at the memory page and clear anything you didn’t intend to hand over permanently.

For anything genuinely sensitive, use the temporary mode. Roughly thirty days of abuse-monitoring retention is a much smaller exposure than an indefinite history entry that also feeds training.

Strip identifiers before you paste. This is the highest-value habit on the list and it costs nothing. An assistant can help you with a contract dispute without knowing the counterparty’s name. It can review a medical letter with the names and reference numbers replaced. The quality of the answer almost never depends on the details that make the text identifying.

Check what you’ve connected, in the account that owns the data rather than in the assistant, and revoke what you’re not using. And treat a work assistant as a work system, because that’s what it is.

If you want a single rule instead of a checklist: before pasting something, ask whether you’d be comfortable with that text sitting in a support ticket at that company for a year. That isn’t the literal fate of it. It’s just a useful level of exposure to picture, and it lines up with the real mechanics: a hosted service, a stored record, a small chance a human reads it, a schedule that eventually clears it out.

Running your own model

A note on local models, since it always comes up. Running a model on your own hardware genuinely removes the provider from the equation, and for some people that’s the right answer. It also means worse output for most tasks, real setup effort, and a machine that can handle it. It’s better to use the hosted tools thoughtfully than to bounce off a local setup and go back to pasting everything into the default with training left on.

The plain version

A chatbot is a hosted service that stores what you send it, has settings governing what it does with that, samples some of it for review, honours deletion on a schedule, and answers to courts like everyone else. We already know how to think about services like that. The conversational interface is what throws people off, because it borrows the emotional register of a private conversation while running on the plumbing of a web application.

The answer isn’t to use these tools less. The answer is to notice which window you’re typing into and adjust what goes in it, the same way you already do between a work email and a message to a friend.

No setting here makes you anonymous to a provider you have an account with. What the settings do is shrink the retained pile and stop it feeding the next model. That’s worth having. It just isn’t the same thing as privacy.

For more breakdowns like this one, the guides and tool tests live at The Privacy Wire.

from the team
Want a real mobile IP, not a datacenter VPN endpoint?

Shared VPN exit nodes get flagged and blocked. Singapore Mobile Proxy runs real 4G/5G mobile IPs that give you a residential-grade address carriers still trust.

see how it works →
read on
More from The Privacy Wire

VPN and tool reviews, realistic opsec guides, and privacy news for people who want to protect their data.

browse all articles →