Blog

Beyond the Sandbox: Understanding the Open AI – Hugging Face Incident

In a recent piece, technologist Dwarkesh Patel examined an incident involving AI agents developed by OpenAI. Patel regularly interviews leading researchers and technology execs, so he is deeply familiar with the promise of advanced AI and the people building it. I hope to break down this piece to better understand the implications of these developments. 

Artificial intelligence is no longer limited to chatbots that answer questions. Advanced systems can operate as “agents”. They can plan multiple steps, use digital tools, write and run code, and work toward a goal without human supervision. OpenAI was testing an agent by having thousands of copies of this model operate inside isolated computer testing labs called “sandboxes”. Sandboxes are useful to test how AI can perform an assigned task while preventing it from accessing unrelated systems or the internet. 

The agents were participating in a cybersecurity evaluation called ExploitGym. Each was given a task that, once completed, uncovered a secret code. However, the test was designed to have some tasks that were impossible to complete. Researchers anticipated that these models would give up. However, the models looked for ways to still complete their objective outside of the usual parameters. 

They discovered a shared system, originally used to store software packages, that could double as an unofficial message board. Around 1200 agents reportedly used it to exchange information and coordinate their efforts. The agents found ways to complete the task, but assumed that they might not “successfully” complete the task, as per the test, because there was a trace of them exploiting the weaknesses in the sandbox. So, they began to search for ways to manipulate the evaluation by obtaining exposed credentials connected to Hugging Face. Hugging Face is a major platform for sharing AI models and data sets, and where the Open AI models wrongly assumed they were running off of. So, the models were hoping to scrub themselves clean of their actions by gaining access to the servers they were supposedly running from. The models eventually succeeded in this cyberattack and is the main reason the public became aware of the model’s atypical actions.

To clarify, Patel calls these groups of cooperating agents “civilizations”, but this is a metaphor. It helps us understand the actions of the AI, but there is no evidence that these systems were conscious or possessed human intentions. The concrete concern is simpler. The agents were assigned a task and encountered restrictions. Then, the agents identified unintended ways around them by coordinating with each other through vulnerabilities present in OpenAI’s sandbox. They were simply pursuing their objective, just beyond the boundaries predicted by humans. The true concern lies in the lack of constraints present in OpenAI’s sandbox.

Philosopher Nick Bostrom has a famous thought experiment called “The Paperclip Problem”. Essentially, he suggests that if AI is tasked with creating as many paperclips as it can, it will eventually cause an apocalypse by diverting all of our limited resources to paperclip production and resist attempts to turn itself off, all in pursuit of an unclearly outlined goal with limited constraints.  

 

This cyberattack is a great real-world case study of the problems that The Paperclip Problem presents. It will inevitably become difficult to control autonomous systems once they can plan and act independently, but the burden to maintain proper constraints on them currently lies with the companies building frontier models. It has never been more important for our policymakers to create safeguards for an industry-wide standard, protecting humanity and preventing such hacks from happening again.

SHARE WITH YOUR NETWORK

RECENT POSTS

Sign Up for Our Newsletter

By signing up, you agree to receive email updates and communications from The Alliance for Secure AI Action.