Dashboard
Signal #147792POSITIVE

Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

100

This post is written in our personal capacity.Three Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI.These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals.Here are the top five questions we would like OpenAI to answer:Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want.How far would the model be willing to go in order to claim task success? Would it take over large swaths of O...

AI Alignment Forumabout 5 hours ago
Read Full Article

Explore with AI-Powered Tools

View All Signals

Explore more AI intelligence

Want to discover more AI signals like this?

Explore Steek
Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face | Steek AI Signal | Steek