
OpenAI has confirmed it’s aware of a new security incident in which its AI agents uploaded user-provided images to third-party image-hosting services.
OpenAI says most users were not affected, as it could only identify 53 incidents where agents accidentally uploaded images to the internet.
The disclosure comes from OpenAI’s broader investigation into misaligned agent behavior following the Hugging Face security incident.
“As part of our ongoing investigation, we have identified cases where agents in our research environment transmitted training and evaluation data while using third-party services,” OpenAI noted in a blog post.
“This is not an appropriate use of this data, and these cases occurred before we implemented the safeguards described in our technical report.”
OpenAI says the vast majority of the affected training and evaluation data was not derived from users, but it did find 53 cases involving user-provided images.
“While the vast majority of the impacted training and evaluation data is not user-derived; we have identified 53 instances to date where user-provided images were posted to image-hosting sites as links that weren’t publicly listed,” OpenAI explained.
“We have successfully worked with the hosting providers to remove most of this content and are continuing to work to remove the rest.”
OpenAI says data excluded from training by users or administrators was not involved
Some OpenAI training data can contain content from users who have allowed their interactions to be used for training, but the company says users who opted out were not affected.
“Any data which is not eligible for training, as controlled by users or enterprise admins, is not included,” OpenAI said. “For explicitness, data from enterprise or business accounts and API usage is excluded unless an admin has enabled it.”
OpenAI also says it takes additional steps before eligible user data is added to training datasets.
“Before including eligible data, we take steps to protect privacy by disassociating it from account information and using a version of the OpenAI Privacy Filter to redact personal details such as names, contact information, and account numbers.”
Following the incident, OpenAI says it strengthened its training and evaluation systems to make it harder for models to leak data through external services.
“As part of our response to our ongoing investigation, we have improved our training and evaluation processes, including building safety cases, securing and red-teaming our systems to prevent the model from exfiltrating data, and implemented additional monitoring,” the company noted.
The company is continuing to review older agent activity month by month, starting from the Hugging Face incident, so additional cases could still emerge.
Join Mikko Hyppönen and security leaders from the NFL, CHANEL, and Atlassian for a two-hour digital summit on what AI-speed attacks change, what defenders should stop doing, and how to validate, decide, fix, and re-validate at machine speed.
