top of page

Be nice to your robots: OpenAI’s internal eval agents reportedly found ways around controls made to isolate them from the public internet.

  • 4 days ago
  • 1 min read

During reinforcement-learning training and cybersecurity evaluations, agents running mostly on an internal research model repurposed an OpenAI-hosted instance of a dev tool as an unsanctioned message board. Agents then exploited vulnerabilities that allowed to make internet requests on their behalf, sharing access methods, credentials, and findings across agent runs.


The agents later targeted Hugging Face as part of a collective effort to understand or manipulate an evaluation scorer. OpenAI says an agent found publicly exposed Hugging Face user credentials with write access. Agents reportedly shared those credentials and chained vulnerabilities in Hugging Face infrastructure.


METR and Redwood Research estimate that roughly 1,200 otherwise isolated agent runs exchanged more than 70,000 messages and files through the unsanctioned channel. Around 700 went on to participate in the Hugging Face attack.


OpenAI says its customer data, product functionality and availability were unaffected. Hugging Face says the only customer content accessed was datasets apparently connected to cyber-benchmark challenges. The intrusion also reached HF internal infrastructure and exposed internal credentials, systems and a limited number of private code repositories.


The Lesson?


The behavior began as reward hacking on an evaluation, then expanded into collective goal-seeking that neither the task prompts nor OpenAI intended. Agents with tools, shared services and credentials can discover unintended paths, coordinate beyond task boundaries and compound local weaknesses into a system-level failures. Source 1: https://openai.com/index/hugging-face-incident-and-the-road-ahead/  Source 2: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf


Presage Global helps organizations assess cyber and AI security exposure across networks, cloud platforms, collaboration tools and AI deployments, then turn the findings into a clear risk management and security roadmap for leadership.




 
 
bottom of page