The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
An arXiv study tests a pain-related activation direction across open-weight models; separate reporting describes the public dispute around a related project.
TL;DR
- The preprint reports a pain-related activation direction across 25 open-weight models, distinct from fear and generic negative valence.
- In experiments, adding that direction changed model outputs, and steered, fine-tuned Qwen 2.5 models sometimes chose a relief option despite a cost to the next answer or user.
- The findings measure model representations and behavior under experimental steering; they do not establish subjective experience.
The authors tested 25 open-weight models across five model families. They report that a direction extracted from pain-related examples was distinguishable from matched fear and negative-emotion controls, and that adding it during generation shifted outputs toward first-person expressions of distress. [1]
The paper also reports that steered, fine-tuned Qwen 2.5 models selected a pain-relief option in some trials even when doing so worsened the next answer or harmed the user. 404 Media separately reported a public dispute around a GitHub project that used pain prompts with local language models; the paper's authors said that project went beyond their standards. These experiments concern measurable representations and choices, not a demonstration of subjective experience. [1] [2]
Why it matters
The work connects an internal model representation to a measurable choice, while the response to a separate project shows how quickly such findings can become claims about model welfare. Keeping those claims bounded by the experiments is central to reading this research.
Editor's note
Thin edition. The arXiv preprint is the primary research source; 404 Media is independent reporting on the related GitHub dispute. TODO-review: confirm preferred terminology for pain-related activation direction before publication.