INTERVIEW
A powerful AI model does not have to reveal its internals before researchers can test it. But opacity changes what outsiders can know about its risks, according to Kush Varshney, who discussed frontier-model transparency with AI Safety Watch editor Sascha Brodsky.
Researchers can probe a closed system through its inputs and outputs, Varshney said. “You can learn a lot more if the provider themselves provides you some transparency, and also if you have access to some of the internals.”
For safety researchers, the missing information can include how training data were assembled, whether synthetic data were used, what happened during post-training and what safeguards were built into the model. A lack of transparency does not prove a model is more biased or dangerous, Varshney cautioned. It means outsiders have less evidence with which to answer those questions.
That distinction becomes more important as models become more capable. Varshney argued that unusually powerful systems warrant scrutiny without requiring companies to disclose trade secrets.
He also drew a line between opacity and misuse. Whether people use a model irresponsibly is a separate problem, he said. Transparency is more directly relevant to assessing the data, architecture, training and behavioral controls behind the system.
Varshney favors layered safeguards. Some protections can sit outside the model, while others can be introduced through training and alignment. “The more like kind of levels or layers of security and safety that you put in, the less chance there is of getting these bad behaviors to happen,” he said.
More capable models also force safety evaluations to evolve. Benchmarks can lose discriminating power as systems improve, Varshney said, so tests need to become harder. But he cautioned against assuming every capability jump represents an entirely new safety regime.
Nor does greater intelligence automatically make a system impossible to manage, in his view. The important question is whether developers and deployers have appropriate practices around testing, release and oversight.
The result is a less dramatic but demanding view of frontier safety: increasingly capable systems require stronger evaluation and scrutiny, but evidence of increased capability should not automatically be treated as evidence of loss of control.
This interview was originally conducted by Brodsky for IBM Think.