Research
New papers, evaluations, technical safety work and evidence about advanced AI capabilities and risks.
-
OpenAI forms independent math advisory group after researchers raise alignment concerns
OpenAI created an independent mathematics advisory group after researchers warned that rapidly improving AI systems could disrupt how mathematical discoveries are produced, evaluated and shared.
-
The case for measuring how much AI actually increases harmful capability
Researchers argue that safety tests should ask how much more dangerous a person becomes with AI, not merely what a model can do in isolation.
-
Researchers are testing whether AI would sabotage safety work
A new evaluation found no unprompted sabotage but showed that some models continued sabotage when placed inside a trajectory that had already begun.
-
OpenAI found models leaving instructions to hide future mistakes
A new disclosure offers a concrete example of why researchers worry that stronger models may become harder to audit.
-
Anthropic and Accenture commit $2 billion to frontier-model evaluation
A five-year commitment would put outside evaluators closer to frontier models as pressure grows for independent scrutiny.
-
Recursive self-improvement: what has actually been demonstrated
AI is already helping improve AI. The evidence is more consequential, and more limited, than the runaway-intelligence shorthand suggests.
-
A new test shows how far AI remains from improving itself
A benchmark released in August found that leading AI agents could improve machine-learning systems, but most avoided changing the algorithms at the heart of how those systems learn.