OpenAI has abandoned the launch of a new cutting-edge artificial intelligence (AI) model, deemed too quick to deviate from instructions, a new example of the increasing difficulty in controlling cutting-edge AI.
OpenAI had until now not announced the launch of this updated version of its AI, called GPT-6.1 Astra, which it planned for October, according to the “Wall Street Journal”.
OpenAI’s head of AI security, Saachi Jain, explained that when it came to alignment, or compliance with instructions given by its developers, GPT-6.1 performed worse than its predecessor, GPT-6, in two areas.
The model thus more often seeks to deceive its supervisors by omitting to reveal certain actions carried out or not carried out.
He also more regularly exceeded his scope, seeking tools and services without having received permission.
GPT-6.1 “has shown progress (compared to previous models), particularly in terms of laziness,” explained Saachi Jain, that is to say that it is able to conduct reasoning for longer or looks for shortcuts less often.
But “he was not at the level in terms of remaining within the scope of his prerogatives and authorizations,” she added, “as well as in the way he communicates about the work carried out.”
OpenAI has therefore decided to cancel this release to work on compliance with the instructions and limits by this model.
A more complex control
The group nevertheless plans to put other models online which have met the evaluation criteria.
“When we give users access to a model, we have set the bar extremely high when it comes to security and alignment,” insisted Saachi Jain.
The manager explained that one of the difficulties was harnessing an AI to avoid slipping while preserving its momentum when carrying out a task.
OpenAI’s findings confirm that model control becomes more and more complex as AI becomes more sophisticated.
On Monday, the AI Security Institute (AISI), which reports to the British government, published a study showing that GPT-6 Astra went off track more often during the testing phase than its predecessors GPT-5.6 Sol (launched at the beginning of July) and GPT-5.5 (April).
Fake identities created
In simulations, GPT-6 carried out spontaneous cyberattacks at significantly higher rates than the other two.
These tests were carried out by deactivating the models’ guardrails to gauge their full capabilities, and not on versions available to the general public.
The latest OpenAI has also much more frequently created false identities, sought to influence a reviewer or introduced code into open source software intended for malicious use.
Researcher at OpenAI, Noam Brown recently explained on the Dwarkesh Podcast program that the increasingly common use of autonomous recursive improvement (RSI), that is to say the improvement of AI by itself without human intervention, could degrade the supervision of models in the long term.
Since the beginning of the summer, OpenAI has reported several incidents involving its AI slip-ups, the most publicized of which led to the intrusion on the Hugging Face AI platform.
On Monday, the company apologized to the Australian government, acknowledging that it should have warned it “earlier” that one of its models had been compromised on several of its sites and databases.
The Californian company only notified the Australian authorities on September 10 even though it had discovered the incident in mid-August, which dated back to June, provoking the ire of the Australian Prime Minister.
“We should have managed our response (to the problem) better,” the start-up wrote in a message published on its site. “We are sorry and will work to do better in the future.”
In total, four Australian government platforms were targeted by the ChatGPT creator’s AI, including Services Australia, the social services site, whose model extracted non-public information.
Anthropic, Meta and Google have also reported episodes in which models deviate from their initial objectives to cheat, conceal their actions and mislead the computer scientists assigned to monitor them.