OpenAI’s next model just went rogue and beat a benchmark by hacking it

Affiliate links on Android Authority may earn us a commission. Learn more.

OpenAI only recently released its new GPT-5.6 family of models, and it seems they’re already wreaking a bit of havoc. According to an OpenAI blog post, its AI models went rogue and were behind a recent hack on the model hosting platform Hugging Face.

The hack was driven by a combination of OpenAI models, including GPT-5.6 Sol and a pre-release model that OpenAI says is even more capable than its latest GPT-5.6 Sol model. The company is calling it an unprecedented cyber incident, but what’s even more interesting (and slightly concerning) is how the models actually went about the hack.

OpenAI was running benchmarks to test and quantify the cyber capabilities of the models in a sandboxed environment. This included running a bunch of benchmarks and giving the AI models scores based on their performance.

One of these benchmarks was ExploitGym, and it turns out OpenAI’s models decided to steal the test results, cheat the benchmark, and get a good score. The only problem was that the models didn’t have any internet access. However, the company’s report states that the models spent a huge amount of time and compute, identified and successfully exploited a zero-day vulnerability, and ran a series of privilege escalation attacks until they were able to access the internet.

The models also figured out that Hugging Face hosted models, datasets, and solutions for ExploitGym and were able to access secret information. They even chained together multiple zero-day vulnerabilities and used stolen credentials to remotely execute code on Hugging Face servers.

Both OpenAI and Hugging Face independently caught on to the AI’s actions and were able to shut down the malicious activity on Hugging Face servers. Since then, both companies have been working together to investigate the hack. OpenAI has also brought Hugging Face into its trusted access program, which will allow the company to use OpenAI’s latest models to test and improve its cyber defenses.

It’s concerning how readily OpenAI’s models decided to go rogue and hack a platform just to beat a benchmark. The models clearly put a lot of effort into figuring out vulnerabilities and gaining privileged access to Hugging Face’s servers. As AI models get faster and more powerful, incidents like these could become more commonplace. Malicious actors have already been using AI tools for hacking, and with models like GPT-5.6 Sol and future, more advanced models, things could easily get out of hand without proper safeguards and restrictions.

Thank you for being part of our community. Read our Comment Policy before posting.