Interviews
Conversations with researchers, policymakers, writers and other people shaping the debate over AI risk, security and governance.
-
Nathalie Baracaldo: Why AI agents aren’t ready to improve themselves
Nathalie Baracaldo says autonomous self-improvement could amplify reward hacking and misalignment before researchers know how to verify alignment reliably.
-
Why an AI explanation may sound right and still be wrong
David Cox says fluent self-explanations can tempt people to anthropomorphize AI. Mechanistic interpretability aims at the harder problem: understanding what is actually happening inside a model.
-
The problem with AI systems we can’t see inside
Kush Varshney argues that powerful AI systems can still be tested from the outside, but opacity leaves researchers with less evidence about data, training and safeguards.
-
Why chain-of-thought monitoring can miss dangerous AI behavior
Shikhar Shiromani found a sharp blind spot in chain-of-thought monitoring: plausible reasoning can make suspicious actions much harder to detect.