Many frontier-model evaluations ask whether an AI can complete a dangerous task. A group of researchers argues that the more useful question is comparative: How much better can a person perform that task with AI than without it?
A framework proposed by researchers including Michelle Vaccaro and Michiel Bakker calls this “harmful capability uplift.” It would measure the marginal increase in a user’s ability to cause harm when assisted by a frontier model compared with conventional tools.
The distinction matters for policy. A model might score highly on a biological or cyber benchmark while adding little to what a skilled specialist could already do. Conversely, a system that turns a novice into an effective attacker could create substantial real-world risk even if its absolute benchmark performance appears less dramatic.
Human-centered evaluations are harder and more expensive than static benchmarks. They may also be closer to the question regulators ultimately care about: not what knowledge exists inside a model, but how much additional dangerous capability deployment puts into the world.