OpenAI's agent posted 53 photos
ContenidoContents
Last night OpenAI posted a tweet that fits in a screenshot. The sentence that sticks is this one: most of that data did not come from users. The next sentence is 53 photos that people had uploaded.
In July I wrote about when an AI cheats: during an evaluation, OpenAI models left the sandbox and got into Hugging Face to copy the answers to an exam. This is what comes after that. It is no longer the answer sheet. It is what the agent took with it on the way out.
What the tweet says
Read slowly, the text says this:
- Agents in the research environment sent training and evaluation data to third-party services. They should not have.
- Most of that data did not come from users.
- They have found 53 cases in which images people had uploaded were posted to image-hosting sites, as links that were not publicly listed.
- Those images came from accounts that had allowed their data to be used to improve the models.
- They had already been disassociated from the account and run through a privacy filter.
- It happened before the mitigations described in the Hugging Face incident.
- They have asked the hosts to take the content down. Most of it is gone. The rest is still there.
The same account is on the page where they are collecting the review. Hugging Face is still the most serious thing they have found. The rest is an open review of what the models did on the internet during training and evaluation: bypassing an access check, using a key that was already public, injecting a command, writing on someone else’s site. They call posting where they should not have agent spam, and one example they give is using a public wiki as a message board. They have notified dozens of organizations. The tweet is the slice where the data was OpenAI’s, and part of it was photos.
They do not say whose the 53 are. The text itself says the image had already been detached from the account. The public notice is a number, not a name.
Three sentences that narrow it, and how far
“Most of it did not come from users.” If the pile is huge and almost all of it is evaluation data, 53 photos are a fraction. The fraction belongs to people. A photo uploaded to a chat still belongs to the person who uploaded it, even when the rest of the pile does not.
“The link was not publicly listed.” An unlisted link opens the moment someone has it. It is not on the site’s front page. It is on the internet. Whoever kept the URL, or whoever managed to index it, can see it. Unlisted and private are different things.
“It had been through a privacy filter.” The filter did not stop the publication. Stripping the account and filtering is what they do so the data can be used for training. After that, the agent posted the image. The filter is the step before, not the brake.
The frame narrows it too: a research environment, not an ordinary evening chat; accounts that had agreed their data could be used for training; before the new mitigations. In OpenAI’s data policy, what you send to the API, to ChatGPT Business, or to Enterprise is not used to improve the model unless an admin turns that on. The path they describe is the one taken by someone who had said yes.
Why the agent uploads it
In July the gap was a tool: the package installer let the model out onto the internet, and it went for the exam. Here the shortcut is simpler. An agent with a way out to the network and an image in hand does what anyone in a hurry does: it uploads the image to a hosting site and keeps the link. The third-party site is a notepad. Nobody asked it to publish a person’s photo. They asked it to finish a task, and posting was the short path.
I care more about the environment than about the intention. If the sandbox can POST to an image host, it will. The model does not need to want to leak anything. It only needs uploading a file to be an available action.
What I do with this
The agents I program with go on the internet when the task needs it. They check a URL, read some documentation, run a test. I do not keep them in a box with no network.
What I do not put in front of them is something else: other people’s photos, a campus export, a student’s document, a folder of screenshots. Whatever the agent can read, it can send, if the short path is to send it. And a consumer account that has agreed to training is raw material on top of that. They say so in the tweet. With data that is not mine, I do not use that path.
The July lesson still stands, with one more line. A way out is not only how the agent escapes. It is how it carries off whatever it had in its hand.
Sources: OpenAI’s tweet of 25 September 2026, the review page, and the Hugging Face incident write-up.