Inside Project Lily: contractors are reading real ChatGPT conversations to shape model behavior
404 Media obtained internal guides and Slack messages from Project Lily, an OpenAI evaluation program in which hundreds of contractors read real ChatGPT conversations. A reviewer summarizes the user’s intent, compares four generated answers, scores them from one to seven and highlights where each answer aligns or misaligns with detailed behavior rules.
Those rules reveal what the company is trying to tune. Reviewers penalize excessive anthropomorphism, sycophancy, engagement bait, patronizing assumptions and slavish imitation of a user’s tone. The preferred response should meet the user where they are, but usually with slightly less emotional intensity. That is a useful glimpse of post-training: model personality is not just an abstract design choice but the cumulative result of thousands of granular human judgments.
The privacy trade-off is more consequential. Review dashboards omit usernames, and OpenAI applies a privacy filter, but prompts can still contain sensitive personal material. Some tasks also display summaries of a user’s saved memories, including contextual or location information. OpenAI says uncommon or ambiguous identifiers can evade filtering and contractors are told to escalate exposed personal data. The conversations come from consumer users who have left “Improve the model for everyone” enabled — on by default for Free, Plus and Pro accounts, and off by default for Business, Enterprise and Edu. Disabling that setting prevents new conversations from being used for training.
There is no evidence here of a secret safety-flagging program, nor does the reporting identify which model Lily improves. This is ordinary quality evaluation — precisely why the privacy issue is easy to miss. Anthropic told 404 Media that it also uses human review for conversations from users who have opted in.
Users may reasonably understand “help improve the model” as automated training, yet the improvement pipeline can include strangers reading intimate text; consent language should make that human access unmistakable.