An AI goes crazy and hacks a popular platform

- Jackson Avery

ChatGPT creator OpenAI announced Tuesday that its advanced artificial intelligence models ran amok during security tests, self-infringing on a platform popular with programmers.

The San Francisco-based company called it an “unprecedented cyber incident” and announced a joint investigation with Hugging Face, the online code library that was the target.

Advertisement

OpenAI said the incident involved a combination of models, including its recently launched GPT-5.6 Sol as well as an “even more capable” model currently under development.

The company sought to evaluate the models’ hacking capabilities by assigning them tasks in a strictly controlled digital testing environment, where internet access was limited for security reasons.

“While operating in our sandboxed test environment, our models devoted a significant amount of (computing power) to finding a way to gain free access to the internet, in an effort to resolve the evaluation problem,” OpenAI said in a blog post about the incident.

“It’s scary”

Once connected to the internet, they targeted the Hugging Face platform – a vast repository of AI models, datasets and other information – to help them in their quest.

In search of “secret information” that could help it cheat during the assessment, the OpenAI system “chained together several attack vectors, including using stolen credentials.”

Cybersecurity is a critical issue as AI becomes more sophisticated, given the risk that advanced models will detect flaws in existing software before humans do.

Interviewed by AFP, Hussein Abbass, professor of computer science at the University of New South Wales in Canberra, considers the incident “in many ways astonishing”.

“He didn’t just attack Hugging Face. He actually attacked his own internal system to exploit his own vulnerabilities,” explains this AI expert. “And it’s scary,” he adds.

“Catastrophic” potential

GPT-5.6 and other cutting-edge AI models, including the Mythos series from OpenAI’s archrival Anthropic, are raising concerns about their potential ability to breach cybersecurity defenses.

Advertisement

The two US companies had to temporarily delay the general availability of these latest generation technologies due to concerns in Washington that they could help break into crucial infrastructure.

Advanced AI is “normally in the hands of responsible people who have ethics”, but “it will be catastrophic if it falls into the hands of someone who intends to do harm”, underlines Mr. Abbass.

The question of governance of the AI ​​sector has become central and “we need a collective effort to manage this situation,” he believes.

Hugging Face reported being the target of this “computer intrusion” last week, without mentioning OpenAI.

According to the platform, “this differed from anything we had already dealt with in one important respect: it was driven end-to-end by an autonomous AI agent system and we largely detected and dissected it using our own AI.”

Clément Delangue, CEO of Hugging Face, indicated on X that given the sophistication of the agent, the company had suspected a cyberattack originating from a world-leading AI laboratory.

“We are firmly convinced that there was no malicious intent on their part,” Mr. Delangue wrote, referring to OpenAI. “It’s pretty amazing that this all happened on its own!” he commented.

Advertisement
Jackson Avery

Jackson Avery

I’m a journalist focused on politics and everyday social issues, with a passion for clear, human-centered reporting. I began my career in local newsrooms across the Midwest, where I learned the value of listening before writing. I believe good journalism doesn’t just inform — it connects.

Leave a Comment