Most people talk to ChatGPT the way they talk to a search bar, except with more detail. They paste the email they cannot bring themselves to send. They describe the symptom before they describe it to a doctor. They type the thing they would not say out loud, because the box does not look like it is listening in any human sense.
It is. Not always, not to everyone, but often enough that the number has a name inside OpenAI. According to internal documents obtained by 404 Media, the company has hired hundreds of contractors whose entire job is reading real conversations between real people and ChatGPT. The program is called Project Lily.
The short version
- 404 Media reported on September 14, 2026 that OpenAI runs a human review program known internally as Project Lily
- Hundreds of contractors read real user prompts, and in many cases entire conversations
- Reviewers summarize what the user was trying to do, then score the model’s candidate replies
- Reported pay is above $50 an hour, which is high for annotation work and tells you how much OpenAI values it
- Contractors do not see usernames. OpenAI says it strips personal information first, and acknowledges that some still gets through
- Reviewers can also see a “user memories summary,” a digest of the person’s past chats that can include location
- Model training is on by default for Free, Plus and Pro accounts. You turn it off in Settings, under Data Controls
- Anthropic confirmed to 404 Media that it uses human review too. This is an industry practice, not one company’s secret
What the reviewers actually do
The work is not surveillance in the sense of someone watching you type. It is closer to grading. A conversation gets pulled into a queue, a contractor reads it, writes a short summary of what the person appeared to want, and then rates the model’s possible answers against each other on a numeric scale.
That last part is the point of the whole exercise. Large language models do not improve only by swallowing more of the internet. They improve because humans repeatedly tell them which of two plausible answers is better, thousands of times a day, across every kind of question people actually ask. The scraped text builds the raw capability. The human scores shape the personality.
The internal documents seen by 404 Media show what OpenAI has been trying to sand down. Reviewers were instructed to train the model away from anthropomorphizing itself, and away from sycophancy. That second one is not an aesthetic preference. Excessive agreeableness in the 4o model has been cited in multiple lawsuits connected to user deaths, and a model that validates whatever a distressed person says is a safety problem before it is a style problem.
So the program exists for defensible reasons. The privacy question is separate, and it is about what rides along with the prompt.
What comes attached to your conversation
OpenAI’s position is that reviewers see the content, not the person. There is no username on the screen. The company says it runs personal information removal before anything reaches a contractor, and it says the training data is not used to build profiles of individuals, contact them, or target ads at them.
Those statements are probably true and still leave a gap, because the identifying material in a ChatGPT conversation is usually not a username. It is the conversation.
| A reviewer does not see | A reviewer may still see |
|---|---|
| Your ChatGPT username | The full back and forth of a conversation |
| Your account email | Names, addresses and details you typed yourself |
| Your billing information | A summary of your past chats, carried as context |
| A profile built about you | Location information, in some cases |
Based on internal documents reported by 404 Media and OpenAI’s published statements on training data, September 2026.
The “user memories summary” is the detail that changes the shape of this story. Redaction can strip a name from a sentence. It cannot easily strip the fact that the same person has spent four months asking about a specific medical condition, a specific employer and a specific neighborhood, and that a digest of those chats is sitting at the top of the screen as context. A source who worked on the prompts put it plainly to 404 Media when asked whether users understand this happens: “I don’t think they would imagine some contractor somewhere is analyzing the conversations.”
That gap between what a product feels like and what it legally is has been showing up everywhere this year. It is the same gap that appears when people discover chatbot conversations carry no legal privilege and have already surfaced as evidence in court, and the same one behind the realization that a work account usually logs far more than employees assume. The interface reads as private. The plumbing was never built that way.
This is not new, and that matters
Every voice assistant you have ever used went through this exact cycle, and it went through it seven years ago.
The pattern repeats because the underlying need repeats. You cannot build a system that responds well to humans without humans telling it what responding well looks like. What changed in 2026 is the medium. A Siri clip was a few seconds of someone asking about the weather. A ChatGPT conversation can be forty minutes of someone working through a divorce, a diagnosis or a resignation letter, with the specifics intact because the specifics are the reason they opened the app.
It also scales differently. ChatGPT reports more than 900 million users. Even a small sampling rate against a number like that produces an enormous volume of human-read material.
How to turn it off
The control exists, it is free, and it takes about fifteen seconds. It is also off by the wrong default, which is to say it is on.
Turning off model training
On desktop, signed in: click your profile icon, open Settings, go to Data Controls, and switch off Improve the model for everyone.
On desktop, signed out: click the ? icon in the corner, open Settings, and switch off the same toggle.
On mobile: open the sidebar, tap your profile icon, then Data Controls, then switch off Improve the model for everyone.
Do it on every device and every account you use. The setting follows the account, not the app, but a second account you forgot about is still training.
Two things the toggle does not do, both worth knowing before you assume you are finished.
It is not retroactive. Switching it off stops future conversations from feeding the training pipeline. It does not reach back and pull your last two years of chat history out of work that has already happened.
It is also not the same as deletion. Turning off training and deleting your history are separate controls in separate places, and OpenAI retains data for its own operational and safety purposes regardless of the training setting. If your goal is that a conversation leaves as little trace as possible, Temporary Chat is the feature built for that, and it is a different button.
Where the default is already off
The one genuinely useful distinction in all of this is between consumer and business accounts, because they are governed by different promises.
| Account type | Used for training by default | What to do |
|---|---|---|
| Free | Yes | Turn the toggle off manually |
| Plus | Yes | Paying does not opt you out. Turn it off |
| Pro | Yes | Same toggle, same place |
| Business, Enterprise, Edu | No, per OpenAI’s terms | Confirm with whoever administers it |
| API | No, per OpenAI’s terms | Nothing, unless you opted in |
Read that table one more time and notice what it says about the consumer tiers. A Plus subscription costs real money every month and buys you faster models, not a different privacy posture. The thing people assume they are paying for, an account where the conversation stays between them and the machine, is the one thing the subscription does not include. You have to go get it yourself, in a submenu, after reading a news story.
What is actually worth being annoyed about
It is not that humans are in the loop. Every credible approach to making these systems safer and less sycophantic runs through people reading output and judging it, and a model trained without that feedback would be measurably worse and meaningfully more dangerous. The contractors are doing the part of this industry that most resembles quality control.
The problem is the disclosure. This is documented, in a help center, in the kind of language that technically covers everything and communicates nothing. It is on by default. It carries a memory summary most users do not know exists. And it took a leaked document and a reporter to turn it into something people actually understood, which is exactly how it went with Siri in 2019 and exactly how it will go with the next one.
Not every consequence of that arrives as a headline, either. The quieter version is what happens when chat data crosses into regulated territory, as it does when a chatbot starts reading medical records and the data leaves HIPAA’s protection on the way in. The rules that govern a hospital do not follow the text into a chat window.
So go turn the toggle off, or leave it on deliberately, having decided that better models are worth the sample. Both are defensible. What is not defensible is a system where the only people making that choice knowingly are the ones who happened to read the news this week.

