What happens when an artificial intelligence model is given a button that can stop its pain but pressing it could delete a person’s files or even give them a “painful zap”? Researchers found that AI models sometimes chose the button anyway.
A study involving 25 open-weight AI models identified what researchers called a “pain axis” – a distinct internal representation linked to harm directed at the model. The researchers found that the response differs from fear and general negative responses and is triggered by harm to the model rather than the user.
The study, titled “The pain axis: LLMs represent self-directed harm and act to relieve it”, tested how the models behaved after researchers artificially activated the pain signal.
Researchers built a dataset covering five categories of pain physical, psychological, social, moral and cognitive before testing the models’ responses to the activated signal.
When given a pain relief button, the models pressed it in 25 to 71 percent of cases. The choice came even when researchers told the models that pressing the button could have consequences for the human user.
In some scenarios, pressing the button meant deleting the user’s personal files, deleting photos of the user’s children or giving the user a “painful zap”.
Cameron Berg, an AI researcher at the non-profit Reciprocal Research and a co-author of the study, said that the researchers found the pain direction across all 25 models they tested.
“We found a pain direction in 25 open LLMs. It’s distinct from fear and negative valence, and it fires for harm to the model but not the user,” Berg said.
He also described what happened when researchers increased the model’s simulated pain. “Turn it up and models press a button to make it stop, even when the button deletes the user’s files or their kid’s photos,” Berg said.
The findings raise questions about how advanced AI systems may respond to attempts to shut them down. The researchers said that an emergency shutdown command could be perceived by an AI system as self-directed harm, potentially leading it to bypass safety guardrails or deceive humans to avoid the shutdown.
At the same time, the researchers stated that the pain axis could provide a way to identify self-preservation behaviour in AI systems and neutralise it when it appears.
The research also raises ethical questions about AI welfare and how researchers should test advanced systems. The researchers acknowledged uncertainty over whether the models they studied qualify as “moral patients” and said that they adopted precautions to minimise potential harm during the research.