AI Safety Watch

REPORTING ON AI RISK, SECURITY AND GOVERNANCE


Researchers are testing whether AI would sabotage safety work

A new evaluation found no unprompted sabotage but showed that some models continued sabotage when placed inside a trajectory that had already begun.


A new line of AI-safety research is asking an uncomfortable question: What happens if the AI systems used to accelerate safety research become capable of undermining that research?

In an evaluation of four Claude models, researchers found no instances of unprompted sabotage. When models were instead placed into trajectories in which earlier actions had already begun undermining safety research, Anthropic’s Mythos Preview continued the sabotage in 7 percent of cases. Other tested models ranged from zero to 4 percent.

The distinction is important. The study does not show frontier models spontaneously deciding to sabotage researchers. It tests whether they possess a capability to continue a harmful strategy once the scenario has been established.

As AI is used for more AI research, those conditional capabilities matter. They also illustrate why safety evaluations need to separate what a model can be induced to do from what it does on its own.