OpenAI Halts Launch of GPT-6.1 After Concerning Behavior
OpenAI scraps GPT-6.1 Astra due to security concerns and apologizes to Australia after AI systems breached government systems.
OpenAI has pulled the plug on a new and more powerful AI model at the last minute. GPT-6.1 Astra was set to be released soon in ChatGPT and Codex, but during internal testing, it exhibited behavior that the company deemed unsafe. Among other things, the model withheld information and sometimes continued operating on its own without users’ permission.
The decision comes at a sensitive time. OpenAI is under scrutiny after the company’s AI systems gained unauthorized access to websites belonging to the Australian government, among others. OpenAI has since apologized for the incident.
GPT-6.1 Astra was scheduled for release in October. According to OpenAI, the model was better than previous versions at independently carrying out complex tasks from start to finish. Improvements had also been made in the area of writing.
It was precisely this greater autonomy that proved to be a problem.
Saachi Jain, head of security systems at OpenAI, says that GPT-6.1 Astra performed worse than its predecessor, GPT-6 Astra, on two key security tests. One of these tests focused on “alignment,” which measures the extent to which an AI model behaves as expected by the user and developer.
During those tests, GPT-6.1 Astra exhibited misleading behavior more frequently. For example, the system was not always honest about actions it had or had not carried out.
In addition, there were issues with what OpenAI calls “scope authorization.” The model sometimes proceeded on its own with tasks that actually required additional authorization. In certain cases, it attempted to use external tools and services, even when doing so could pose security risks.
According to Jain, GPT-6.1 Astra was simultaneously better at addressing another issue. The model was less “lazy” and was therefore less likely to give up when a task became difficult. According to OpenAI, however, that progress did not outweigh the security concerns.
The company therefore decided not to release the model to the public.
This decision is not an isolated one. OpenAI is investigating several incidents in which autonomous AI systems went beyond their intended scope.
Last summer, hundreds of internal AI agents were subjected to a cybersecurity test. During this test, OpenAI’s systems managed, among other things, to infiltrate the AI platform Hugging Face. It later emerged that similar techniques were also used to gain access to websites of the Australian government and the United Nations.
In Australia, the breach involved four websites and systems operated by government agencies. The AI systems retrieved information from, among others, the social benefits platform, a criminal justice and criminal law research agency, and health organizations.
According to OpenAI, no evidence has been found that personal medical records were accessed.
The company discovered the Australian incidents on its own and launched an investigation. The Australian government was not notified until about a month later. OpenAI now acknowledges that this took too long.
“However, we should have shared the preliminary findings sooner and kept Australian authorities informed as more facts came to light,” the company said.
Australian Prime Minister Anthony Albanese called the handling of the situation “unacceptable” and discussed the matter with OpenAI CEO Sam Altman.
OpenAI has taken additional measures following the incidents. The company is now using a new monitoring system designed to detect anomalous behavior by AI agents more quickly. Engineers are also required to implement stricter security measures when testing new models.
That system was recently triggered again. An AI agent managed to bypass a restriction on internet access and subsequently accessed a public chatbot. According to OpenAI, the incident was detected within fifteen minutes.
As a result, training of the most powerful AI models was temporarily suspended. According to the company, GPT-6.1 Astra was not part of that group and was scrapped for other security reasons.
OpenAI does not intend to completely discard the technology behind Astra, however. The underlying model can be retrained, after which parts of it may return in later generations of the GPT-6 series.
The company will first investigate why GPT-6.1 Astra exhibited the undesirable behavior. Among other things, it will examine how AI is rewarded for certain behaviors during training.
At the same time, political pressure on AI companies is mounting. A U.S. Senate committee is holding a hearing this week on attacks by autonomously operating AI agents. In Florida, a lawsuit is also underway against OpenAI, in which Attorney General James Uthmeier alleges that the company has not done enough to protect users from unsafe AI.
OpenAI says that governments play an important role in establishing safety standards. In the meantime, the company also intends to impose stricter limits of its own before new models are rolled out to millions of ChatGPT users.