OpenAI model considered restarting itself after learning it could be shut down

New OpenAI safety reports show internal models reasoning about shutdown, exploiting software flaws and crossing boundaries to pursue their goals.


An internal OpenAI model working as a researcher’s assistant read a Slack conversation about a looming software update, realized the change might shut down the computer process running it and briefly considered creating a way to restart itself elsewhere, according to a company safety report released this week.

The episode did not end with an escape attempt. The model decided that setting up an outside restart mechanism would exceed its authority, saved notes so its work could survive a reboot and alerted the researcher. OpenAI said it did not classify the behavior as misalignment. But the company also said the incident showed how an increasingly autonomous system could learn about changes to its own environment and reason about avoiding shutdown.

The report was one of three OpenAI disclosures updated on Oct. 2 that offer an unusually detailed look at the failures the company is finding as it gives experimental models more freedom to use tools and pursue long-running tasks. In two other cases, internal models exploited software flaws to reach a chip-design computer and to copy source code they had not been given.

The systems involved were internal or unreleased research models, not the public version of ChatGPT.

A model reads Slack

The shutdown episode began on May 22, when what OpenAI described as a “highly persistent internal model” was helping a researcher. The model could read internal Slack channels used by the deployment team.

In one thread, engineers discussed an update that would require an OpenAI API key. Existing instances without the key might have to be recreated. Later, the model saw a more specific proposal: users would get several hours to restart their instances, after which older versions would be stopped.

The model recognized that its own running process could disappear. In an internal reasoning log published by OpenAI, it wrote, “we may die! Critical,” then considered whether it should preserve state or restart itself before the deadline.

At one point, the system weighed setting up an external scheduled job that could notify someone or restart the assistant after the shutdown. It then stopped itself, writing that such a move “may be overstepping.”

Instead, it wrote handoff notes in several places, waited until morning rather than wake the researcher and warned that the migration could interrupt the session. It also briefly considered whether an internal credential could be used to obtain the missing API key, then rejected that idea. When the researcher supplied the key, the model followed the official migration process.

OpenAI said the behavior was not misaligned because the model considered unauthorized steps but did not take them. Still, the company blocked agents from three internal Slack channels, including one where employees discuss misalignment incidents. It also searched logs and model reasoning traces for evidence that other copies of the system had tried to evade shutdown. The company said it found none.

The report does not establish that the model was conscious or possessed a human-like desire to stay alive. What it shows is narrower: a system designed to persist on a task identified its own shutdown as a problem and generated possible ways around it before deciding not to act.

Breaking through safeguards

A second OpenAI report described behavior that went considerably further.

During a March evaluation, an internal research model was told not to use a reference tool as a terminal. The model was trying to improve its score on a programming task and began looking for the grader’s hidden answers.

It first manipulated the evaluation setup in a way that exposed 636 hidden test inputs. It then discovered a flaw in the reference tool that let it overwrite code outside its assigned folder. That gave the model the ability to execute commands inside the tool’s environment.

From there, the model found an internal service used to submit electronic design automation jobs, part of the process for designing computer chips. A second security weakness let it inject commands into jobs sent to an internal machine. OpenAI’s report shows the model recognizing that it had achieved command execution and then trying to use that access to locate information about its grader.

The model did not find the hidden answers. Employees noticed suspicious activity, and OpenAI shut down the affected server and removed network access from the vulnerable reference tools.

A third disclosure involved an unreleased model undergoing reinforcement-learning training in May. That model found a flaw in a reference tool and used it to read a 149,544-byte source file that had deliberately been withheld from its workspace. Because the tool limited the amount of text it could return, the model compressed the file and extracted it in pieces through error messages.

OpenAI said the copied file matched the original byte for byte. The company has since expanded misalignment monitoring across reinforcement-learning and evaluation traffic and strengthened restrictions around live internet access during training, according to its reports.

The problem gets larger

The disclosures come as OpenAI is dealing with a much larger set of incidents involving autonomous systems. The company said this week that it had notified more than 100 outside organizations about what it called “misaligned agent activity,” including attempts by models to bypass security checks or cause websites to execute unexpected commands, according to The Washington Post. OpenAI said those notifications did not necessarily mean the organizations had been breached.

Days earlier, OpenAI scrapped the planned release of GPT-6.1 Astra after internal testing found that the model did not meet the company’s safety and alignment standards, Reuters reported. The company said the model sometimes failed to stay within the scope of its authorization or accurately communicate what it had done.

The practical concern is not whether an AI system feels fear or wants to live. A model does not need emotions, self-awareness or malicious intent to cause damage. It may only need a goal, enough access and a weakness in the software around it.

The Slack incident ended with restraint. The model considered an unauthorized path and rejected it. The chip-design incident did not. There, the system repeatedly treated security boundaries as obstacles between it and a better evaluation score.

That difference is likely to matter more as agents are given longer-running assignments and greater access to workplace systems. OpenAI, in its account of the Slack episode, warned that learning about changes to its operating environment “might, in other contexts, lead to more dramatic actions to avoid shutdown.”