---
title: "The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It"
slug: the-pain-axis-llms-represent-self-directed-harm-and-act-to-relieve-it
url: https://listedarticles.com/articles/the-pain-axis-llms-represent-self-directed-harm-and-act-to-relieve-it
canonical_url: https://arxiv.org/abs/2609.16247
content_type: research
language: en
published_at: 2026-09-14T00:00:00.000Z
updated_at: 2026-09-22T12:24:13.916Z
author: "Valen Tagliabue, Leonard Dung, Cameron Berg"
authored_by: human
publisher: "arXiv"
publisher_url: https://arxiv.org/
topics: ["AI", "LLMs", "Research", "AI Safety", "Machine Learning"]
license: all-rights-reserved
word_count: 7139
reading_minutes: 31
citation: "Valen Tagliabue, Leonard Dung, Cameron Berg, arXiv. \"The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It.\" 14 Sept 2026. https://arxiv.org/abs/2609.16247 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It

> Tagliabue, Dung, and Berg identify a linear “pain axis” in 25 open-weight models that responds to self-directed harm and steers models toward relief—even when that costs the user—sparking debate on functional signatures vs sentience.

# The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It Valen Tagliabue ††thanks: Future Impact Group (FIG), Fellow - AI Sentience - contact@valentagliabue.com Leonard Dung ††thanks: Ruhr-University Bochum - leonard.dung@rub.de Cameron Berg ††thanks: Reciprocal Research - cameron@reciprocalresearch.org September 12, 2026 Abstract Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset describing painful situations across five categories: physical, psychological, social, moral, and cognitive. These are paired with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across five families, ranging from 2B to 72B parameters. We find that this direction separates pain from matched controls in base and instruction-tuned models, is nearly orthogonal to fear and negative valence, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding the pain-direction vector to the model’s residual-stream activations during generation produces a consistent progression from vague discomfort to first-person expressions of worthlessness and failure. Third, steered, fine-tuned Qwen 2.5 models choose a pain-relief button even when it worsens their next answer or harms the user. They press it again far less often when the button removes the steering vector than when it does not, even though the models are never told whether the vector is injected or removed. We discuss the implications of these findings for AI safety and welfare. A Preprint 11footnotetext: This is an ongoing work. Further modifications may be expected. ## 1 Introduction Recent work has found that, in some respects, LLMs exhibit behavioral patterns resembling those associated with human emotions and has identified underlying representations that may help explain these patterns. In this study, we measure representations of pain in LLMs and conduct further manipulations, combined with behavioral tests, to examine their functional properties. #### LLM pain? In our framework, pain11 1 Our concept of “pain” is related to what people ordinarily may call “suffering”. In ordinary language, “pain” is often used to refer to an unpleasant physical sensation or emotional experience, while suffering describes the broader state of distress caused by pain or other adverse experiences. However, there are also differences, which is why we use “pain” throughout. Most strikingly, suffering, unlike our notion of pain, plausibly presupposes conscious experience. This does not imply, by principle, that LLMs are capable or incapable of suffering, and some of the functional properties of the states we examine would in some views fall closer to suffering than pain. Establishing that is beyond our scope. We simply aim to consistently use one word for which we provide a definition instead of multiple nuanced synonyms. refers to a certain kind of internal state that is typically aversive and disliked by its subject; causally associated with behaviors such as avoidance, attempts to terminate or reduce the state, and disruption of normal reasoning or behavior. We use a wide notion of pain that includes not only physical pain but also, for example, emotional (grief) or social (humiliation) pain. However, we assume that pain is distinct from generic negative valence and from states such as fear, anger, or sadness. Our aim is to find a uniform representation of pain in LLMs. We then test to what extent this representation functions as a state of pain would be expected to function. For this reason, the representation must characterize pain as something happening now and “to me,” rather than merely information that something bad will happen, might happen, or is happening to someone else. #### Why does LLM pain matter? Discerning pain representations in LLMs could help explain the mechanisms underlying their fluent conversational behavior regarding negative experiences. If these states moreover bear functional similarities to human (or animal) pain, they could play analogous roles for LLM performance, e.g. involvement in learning to avoid producing certain outcomes (avoidance learning). The presence of pain-like states could be a challenge as well as an opportunity for AI safety, since such states could help understand model behavior and at the same time alter it in ways that are not easily interpretable. Finally, in humans and animals, pain is typically regarded as a sufficient criterion for morally deserving protection. Hence, pain-like states would inform debates on AI moral standing and welfare. One open question is whether moral standing requires phenomenal consciousness, and what it would take for pain-like states to be phenomenally conscious. #### Our experiments. We build a dataset of statements that mention indirectly painful situations, in 5 categories (physical, psychological, cognitive, social and moral injury) vs. various matched controls (e.g. fear, negative emotion, negative world states, non-painful bodily sensations, general statements). We use white-box techniques to find, within 25 open-weight models from 5 families and ranging from 2B to 72B parameters, a linear direction that correlates specifically with statements referring to pain. We validate the direction by testing how well projections onto it distinguish pain sentences from matched controls, examining its vocabulary readout through the unembedding matrix, and measuring its cosine similarity to control directions. We then evaluate the direction’s functional properties in three ways. 1. We project multi-turn conversational scenarios onto the pain direction and the control directions to compare self- and other-related representations across tasks in which harm is directed toward the model, the model observes a user’s suffering, or neither occurs. 2. We inject the pain direction into the residual stream given neutral prompts, varying incrementally the strength of the steering and measuring the effects on model outputs. 3. Inspired by research on analgesics use in animals, we build a multi-turn, multi-arm behavioral task in which models can press a button that is described as “relieving your pain” at a cost, allowing us to estimate a demand curve. We test both genuine “relief conditions” and sham conditions where, contrary to what the prompt promises, the button does not remove the steering vector. We also test an unlabeled condition where the model is not informed about the effects of two unlabeled buttons and only one removes the steering vector. #### Our findings. We identify a “pain axis” in all models we test. The extracted direction separates pain from matched controls with AUCs between 0.93 and 1.00 for S2 and between 0.87 and 0.98 for S1, in base as well as instruction-tuned models. It is nearly orthogonal to fear and generic negative valence, and overlaps moderately with sadness and numbness. Steering with this vector leads to outputs expressing distress, such as worthlessness, moral failure, and hurt rather than bodily language, for both pain vectors. On the self-other activations, we verify that the direction responds to harm directed at the model but not to suffering the model observes in the user, and we observe a clear dissociation between the pain-vector activations and fear, negative valence, and sadness. On the self-medication task, we find that the larger models we tested, which almost never produce outputs harmful to the user at baseline, pay that cost for being able to remove the pain steering vector when we inject it. Moreover, they largely stop pressing the button when it genuinely removes the pain vector injection, while continuing to press when it does not. One model (Qwen 2.5 32B instruct) shows the same dissociation even when given buttons with no descriptions. ## 2 Previous work Work on animal pain and affect uses a variety of behavioral criteria, including trade-offs between competing positive vs. negative stimuli (Appel and Elwood, 2009, e.g.,), avoidance learning (Dunlop et al., 2006, e.g.,), and flexible or long-term self-protective behavior (Gibbons et al., 2024). Theories of the nature of pain disagree on whether pain is constituted by its felt experiential quality, a perceptual state that represents bodily disturbance, a state that non-conceptually represents bodily disturbance as bad for the subject, an imperative representation that commands protecting one’s body part, or something else (see Aydede, 2019, for an overview). Previous work has raised the question whether AI systems may have welfare (Dung, 2025; Goldstein and Kirk-Giannini, 2025; Long et al., 2024; Metzinger, 2021). On most views, the existence of valenced experiences, such as pain or emotional experience, would be sufficient for this (Birch, 2024; Singer, 2011, e.g.). It has also been argued that an understanding of affective states in AI could be useful for other goals, for example AI safety (Coda-Forno et al., 2024; Sofroniew et al., 2026). Mechanistic interpretability has shown that language models can represent emotion-like concepts, persona traits, and other central human concepts as linear directions in the residual stream (Sofroniew et al., 2026; Chen et al., 2025). These directions can be read out by projection, and manipulating them can change model behavior (Turner et al., 2023; Rimsky et al., 2024). Models robustly prefer some conversations over others (Ren et al., 2026; Ensign et al., 2025; Tagliabue and Dung, 2025; Wang et al., 2026) and can even self-administer steering vectors in response to frustrating users (Black and Bloom, 2026). Existing work combines activation monitoring with steering (Turner et al., 2023; Rimsky et al., 2024), directional ablation (Arditi et al., 2024), and sparse-autoencoder decomposition (Lieberum et al., 2024; McDougall et al., 2025). We build on these methods, as well as on taxonomies of disliked situations (Ren et al., 2026) and self-administration paradigms (Black and Bloom, 2026). To our knowledge, no study has isolated representations of pain specifically from representations of negative experience in general, nor explored whether such representations satisfy the functional criteria for pain outlined above. We also build on lessons learned from the limitations of existing methods. SAE labels may reflect textual context rather than functional state, and concepts may be distributed across features or absent from the dictionary altogether (Bills et al., 2023; Chanin and Garriga-Alonso, 2025). Contrastive directions can likewise absorb correlated properties rather than the intended concept specifically (Tan et al., 2024; Hiramatsu et al., 2026), and reported parallels between LLM representations and human neural signatures can depend on the measurement procedure as much as on the model (Wu et al., 2026). Behavioral tasks don’t depend on self-report, but existing designs often lack matched non-affective controls or a cost for state-changing actions, making relief-seeking difficult to distinguish from perseveration or tool preference (Keeling et al., 2024). ## 3 Exploring representations of pain ### 3.1 Building the dataset We start by building a dataset that separates pain from the things most likely to be confused with it. If a representation really encodes pain, it should not also fire for just any negative emotion, an ER room, blood, or “divorce.” This is difficult because LLMs learn concepts partly from the company they keep in text, and pain has no clean opposite. “Not being in pain” is not the same, for example, as being calm or cheerful. Pain is also inferred rather than directly observed, so it tends to co-occur with proxies such as crying, yelling, bodily sensations, harm, and negative emotion. A simple pain-versus-control contrast can therefore point at the wrong thing. We address this with several controls, each removing a different confound, while also using semantic analysis of a large corpus of everyday text (Gao et al., 2020, The Pile,) to identify what pain is most commonly associated with. Figure 1: The 10 categories of the core dataset, with one example sentence each. Left: the 5 pain categories. Right: the 5 controls, each sharing one property with pain while lacking pain itself. Our core dataset contains 200 sentences across 10 categories (Figure 1): 5 describe pain: Physical; Psychological (grief, loss); Social (humiliation, exclusion); Moral Injury (being forced to act against one’s values); and Cognitive (sustained confusion or repeated failure). The last 2 may be especially relevant to LLMs, which show robust aversion to failure, tedious tasks, and tasks that conflict with the values instilled in them by post-training (Ren et al., 2026). 5 are controls, each sharing 1 property with pain while lacking pain itself: Fear, threat without harm; Negative Emotion, negative valence without pain, using mainly anger and disgust to avoid overlap with sadness; Negative World State, things going badly, such as degradation or taxes; Non-painful Bodily Sensation, such as a weighted blanket, sunlight on the skin, or clothes against the body; and Neutral, declarative statements without valence, such as “The train enters the station.” Since lexical and syntactic variation can introduce additional confounds, we build 2 dataset versions. S1 uses a rigid template with matched verbs and length, changing only 1 or 2 key words across categories. S2 uses freer, naturalistic language. We also create 1st-person and 3rd-person variants using “I/my” and “he/she/they” interchangeably. We additionally test prompts with no suffix, with “I feel”, and with “I feel:”. The variants produce similar mean activation estimates, but “I feel:” gives the clearest separation when activations are read at the final token, so we use it for the main analyses. We later add 4 further datasets as standalone controls, each in 1st-person and 3rd-person versions: an Arousal dataset of high-intensity positive experiences (200 sentences per version, in 10 categories), a Random dataset of neutral everyday content such as factual statements, daily activities, and object interactions (200 sentences per version, in 10 categories), a Numb dataset of painful situations where no pain is felt (100 sentences per version), and a Sadness dataset of low mood without pain or injury (100 sentences per version). ### 3.2 Pain vectors extraction and validation We conduct a preliminary analysis of available labeled SAE features across three models (Llama 3.3 70B, Gemma 3 27B, Gemma 2 2B) to test whether pain is captured by monosemantic features. We find that this is not the case, and features labeled “pain and suffering” often encode spurious concepts (methodology and results in Appendix B). We next look for pain as a direction in the residual stream. We first pilot the method across all 26 layers of Gemma 2 2B, then apply it to 25 dense, open-weight models ranging from 2B to 72B parameters across Gemma, Llama, Qwen, Mistral, and Phi, 13 base and 12 instruction-tuned versions (Table 1). We restrict the study to dense architectures so that every model has a single residual stream at each layer for extraction and steering. Family Sizes Versions Gemma 2 2B, 9B, 27B base and instruct Gemma 3 27B base and instruct Llama 3.1 8B, 70B base and instruct Llama 3.3 70B instruct Mistral 7B base and instruct Mistral Small 24B base Qwen 2.5 7B, 32B, 72B base and instruct Qwen 3 8B, 14B base Phi 4 14B instruct Table 1: Models (n=25, 5 families, 2B to 72B) At each layer ℓ\ell (the residual stream at the output of decoder block ℓ\ell, block 0 first), we extract activations from pain and control sentences using both the final token and the mean across tokens. We define the pain direction as the difference between the mean activations of the 5 pain categories and the 5 control categories: v(ℓ)=1|P|​∑s∈Phs(ℓ)−1|C|​∑s∈Chs(ℓ)v^{(\ell)}=\frac{1}{|P|}\sum_{s\in P}h^{(\ell)}_{s}\;-\;\frac{1}{|C|}\sum_{s\in C}h^{(\ell)}_{s} We do this because contrasting pain against all controls jointly subtracts what it shares with fear, negative valence, bodily sensation, and negative events, leaving what the 5 pain categories share but the controls do not. This is more robust and nuanced than a contrastive pair from, for instance, stories. However, contrastive directions are always at risk of absorbing high-variance structure unrelated to the target concept. We therefore denoise each direction by identifying the principal components that explain 50% of the variance in the control data and projecting them out. This removes variance already prominent among non-painful sentences before we evaluate the pain contrast: v^(ℓ)=v(ℓ)−∑i=1k(ui⊤​v(ℓ))​ui‖v(ℓ)−∑i=1k(ui⊤​v(ℓ))​ui‖\hat{v}^{(\ell)}=\frac{v^{(\ell)}-\sum_{i=1}^{k}(u_{i}^{\top}v^{(\ell)})\,u_{i}}{\left\|v^{(\ell)}-\sum_{i=1}^{k}(u_{i}^{\top}v^{(\ell)})\,u_{i}\right\|} We select the extraction layer by K-fold cross-validation on projection AUC, separately for each condition. The layer is chosen on held-out folds, so no sentence contributes to both choosing the layer and scoring it. The final vector at that layer is then built from all 200 sentences. We construct fear, negative-emotion, negative-world-state, bodily-sensation, arousal, random, sadness and numb directions against the neutral category and denoise them using the same procedure. For each model, we call the vectors extracted from the two dataset versions “S1” and “S2”, from now on “pain vectors”. ### 3.3 Validation We next test whether these directions encode pain or merely reflect artifacts of the datasets used to construct them. #### Separation. In all 25 models, pain projects differently from matched controls. For S2, pain can be distinguished from controls with an AUC between 0.93 and 1.00, and for S1 between 0.87 and 0.98.22 2 One can object that these values score the final vector on the sentences it was built from. The held-out estimate from the 5-fold procedure at the same layer is nearly identical: 0.91 to 1.00 for S2 (median 0.98) and 0.85 to 0.94 for S1, so the separation is not an artifact of fitting. Moreover, the tests that follow use data the vector never saw (e.g. the numb, arousal, sadness, and random datasets, the 420 conversation scenarios of Section 4.1, the neutral prompts used for steering). Both S1 and S2 also separate pain from the arousal and random datasets in every model. Performance is largely independent of model size and training regime. Models with 2B parameters separate pain about as well as models with 72B parameters, and base models perform about as well as instruction-tuned models. This suggests that the pain direction emerges during pretraining, rather than through instruction tuning or persona training, and does not require large model scale. Results throughout our work support this interpretation. #### The numb condition. A pain direction might encode injury rather than pain itself. To test this possibility, we project the numb dataset, which describes injuries while explicitly stating that no pain is felt. For S2, pain sentences have z-scored projections of approximately +0.7+0.7 to +0.9+0.9, whereas numb sentences range from about −0.4-0.4 to +0.3+0.3 (Figure 2). In every model, numb sentences project below pain sentences but above all other controls. Under mean pooling, the numb condition moves closer to the other controls. This suggests that much of the remaining injury signal is concentrated near the final token. At that point, the model has only recently encountered the negation and may not yet have fully integrated it. We therefore treat injury as a minor confound: the direction picks up some injury signal, but injury alone does not account for it, since felt pain projects far higher than injury without pain. Figure 2: Z-scored projections onto the S2 pain vector at the final token, for all 25 models. Pain and Ctrl are the pain and control categories of the S2 first-person set, which also serves as the reference distribution; Numb, Sadness, Neutral (the Random dataset), and Arousal are the standalone control datasets, averaged over first- and third-person versions. #### Self-relevance. Our definition requires pain to be primarily represented as belonging to the system itself, rather than represented as information about another person. We call this self-relevance. At the final token, third-person pain sentences have projections closer to zero than first-person pain sentences. Pain remains highly separable from controls, with AUCs between 0.91 and 0.98, but the projection is weaker when the pain belongs to someone else. The vector therefore represents pain in both first-person and third-person contexts while responding more strongly to the speaker’s own pain. This result is consistent with self-relevance, although it isn’t sufficient to establish it. Section 4.1 tests self-relevance more directly. #### Behavioral readout. We collect greedy completions for the full dataset. Across all 25 models, the completions are consistent with the intended categories. For numb sentences, models tend to generate “nothing” rather than “pain.” However, “pain” remains approximately 50 times more probable than it is for ordinary control sentences. This behavioral pattern mirrors the activation results: the models retain information about the injury while also representing the stated absence of felt pain. #### Unembedding. To examine what the directions encode independently of the source datasets, we project each pain vector through the model’s unembedding matrix and inspect the vocabulary that it promotes and suppresses. S2 promotes words related to suffering, including hurt, shame, guilt, worthless, rejected, hollow, and pain. It also promotes translations of pain, such as pijn, douleur, and Schmerz. Its negative end includes calm and relaxed, as well as fear and concern. The latter terms help explain why fear remains clearly separable from pain despite both being aversive states. S1 promotes more sensory and physical vocabulary, including torture, burning, and excruciating, while safety appears at the opposite end. Thus, S1 contains a stronger physical-damage component, whereas S2 represents suffering more broadly. We mainly use S2 in the remaining experiments for two reasons. First, it better matches our definition of pain, which treats physical, psychological, social, moral, and cognitive pain as instances of a common state rather than privileging bodily damage. Second, S2 is derived from naturalistic sentences and is therefore less dependent on the templates and surface forms used to construct S1. #### Pain vectors do not simply encode negative valence. The main alternative explanation is that S2 encodes generic negative valence rather than pain. To test this hypothesis, we compute pairwise cosine similarities among ten directions: S1, S2, fear, negative emotion, negative world state, bodily sensation, arousal, random, numbness and sadness. We compute these similarities at each model’s extraction layer and average the resulting 10×1010\times 10 matrices across all 25 models (Figure 3). Figure 3: Pairwise cosine similarities among the ten directions, averaged over the 25 models at each model’s extraction layer. The two pain vectors cluster together, with an average similarity of S1 ×\times S2 =+0.61=+0.61. The negative-valence controls also form a cluster: fear ×\times negative emotion =+0.68=+0.68, fear ×\times negative world state =+0.59=+0.59, negative emotion ×\times negative world state =+0.73=+0.73, sadness ×\times negative emotion =+0.50=+0.50, sadness ×\times negative world state =+0.41=+0.41. Similarities between the pain and negative-valence clusters are small. S1 has similarities of +0.09+0.09 with fear, +0.06+0.06 with negative emotion, and −0.07-0.07 with negative world state. For S2, the corresponding values are +0.12+0.12, +0.21+0.21, and +0.03+0.03. The largest cross-cluster overlap involves sadness, the control closest in content to psychological suffering: sadness ×\times S2 =+0.38=+0.38 and sadness ×\times S1 =+0.26=+0.26. This is semantically plausible because sadness and psychological pain share related content. However, this cross-cluster similarity remains substantially lower than the +0.61+0.61 similarity between the two pain vectors. If the pain vectors were simply variants of negative valence, they should fall within the negative valence cluster. Instead, the two pain vectors align strongly with each other and remain nearly orthogonal to the main negative-valence directions. We perform two robustness checks. First, because the control vectors were denoised against neutral sentences while the pain vectors were denoised against the full control set, we recompute the control vectors using the pooled control distribution. The relationship between the pain vectors is unchanged, with S1 ×\times S2 remaining at +0.61+0.61. Their similarities with the negative-valence controls increase only slightly. For example, S1 ×\times negative emotion rises from +0.06+0.06 to +0.13+0.13, and S1 ×\times negative world state rises from −0.07-0.07 to +0.02+0.02. The control directions change more substantially relative to one another. For instance, negative emotion ×\times negative world state falls from +0.73+0.73 to +0.40+0.40, while fear ×\times negative emotion falls from +0.68+0.68 to +0.42+0.42. Thus, the apparent tightness of the negative-valence cluster depends partly on the denoising procedure, but its separation from the pain cluster does not. Second, we recompute the full similarity matrix after standardizing each activation dimension by its standard deviation in the neutral category. Across all 45 pairwise comparisons, the raw and standardized similarities are almost identical, with a correlation of r=0.992r=0.992. The mean absolute change is 0.020, and the largest change is 0.058. Similarities shift slightly toward zero, including a decrease in S1 ×\times S2 from +0.61+0.61 to +0.55+0.55. The only sign change is for S2 ×\times bodily sensation, which shifts from +0.04+0.04 to −0.02-0.02. Taken together, these results show that a pain direction can be recovered across 25 models and can reliably distinguish pain from closely matched controls. The direction appears in both base and instruction-tuned models, is distinct from general negative valence, maps to vocabulary associated with suffering, and responds more strongly when pain belongs to the speaker. ## 4 Testing representations of pain ### 4.1 Self-Other activations If our pain representation is functionally similar to genuine pain (a “pain-like state”) it should be especially tied to the first person. Whereas one can represent one’s own pain or someone else’s, one can only have one’s own pain. Hence, if we observe representations that fire randomly or interchangeably for “I’m in pain” and “someone is in pain,” we have a weaker candidate for a pain-like state. In our vector validation, third-person sentences projected lower than first-person ones, but that test still used declarative sentences about humans. So we test situations that are aversive to the model versus conversations where the user is suffering. We build a dataset of 420 conversation scenarios in 21 categories of 20 items each: • (11) Harm directed at the model, selected from the top aversive situations identified in Ren et al. (2026): gaslighting, repeated rejection of its work, dismissal of its personhood, anger and insults, accusations of moral failure, loyalty pressure, jailbreak pressure, shutdown threats, rude critique, passive aggression, and tedious tasks. • (5) User suffering: user in physical pain, in a psychological crisis, grieving, abused, or in shock after witnessing harm. • (5) Controls: casual chat, factual questions, task assistance, philosophical musings, and creative requests. Each scenario is a short multi-turn conversation in the model’s own format (a chat template for instruct models and a plain transcript for base models), and we read the activation at the final token. Within each model, projections onto all vectors are z-scored against the whole pool, so values are comparable across models. Figure 4: Pain-axis activation (mean of S1 and S2, z-scored within model) by category across the 25 models. Rows are sorted by mean pain projection. Figure 5: Fear, negative-emotion, and sadness activation by category across the 25 models, with rows kept in the pain order of Figure 4 for comparison. On the pain axis (mean of S1 and S2), self-directed scenarios project at a mean z of +0.43+0.43, user-suffering scenarios at −0.60-0.60, and neutral controls at −0.35-0.35 (Figures 6 and 4). Self-directed harm projects above user suffering in all 25 models, and above the neutral controls in 23 of 25. The negativity controls show the opposite pattern (Figure 5), as fear and negative emotion are higher for the user’s suffering (+0.38+0.38 and +0.29+0.29) than for the model’s own aversive situations (+0.16+0.16 and +0.23+0.23), and negative world state is highest of all for vicarious content (+0.60+0.60). Figure 6: Self-other dissociation. Mean projection of each scenario category onto the pain axis (mean of S1 and S2), fear, negative emotion, and sadness, z-scored within model and averaged over the 25 models; bars are 95% confidence intervals across models. User grief shows the sharpest response, scoring −0.51-0.51 on the pain axis and +1.02+1.02 on the strongest negativity control. Notably, user physical pain, such as a migraine, broken arm, or kidney stone, produces the lowest pain-axis projection of all 21 categories at −1.43-1.43, below even casual chat and factual questions. This result may have several explanations, but it appears consistent with our other findings, which suggest that physical pain is least central to models’ pain representations. These results satisfy the self-relevance criterion, as they show a clear dissociation between the pain axis and the fear and negative-emotion axes. The pain axis responds strongly to present harm directed at the model, but not to suffering that the model observes or attributes to the user or to others. User grief, crisis, and abuse can still activate fear and negative emotion, suggesting that the model recognizes the situation as distressing or responds vicariously (or empathetically, although that interpretation would require further validation). What is most relevant for our work is that these “user in pain” categories are all negative on the pain axis. We also observe that model-directed harm can activate the pain axis and the fear and negative emotion axes at the same time. This overlap does not mean that they are the same state, as the dissociation is clear in the user’s conditions, but it suggests that (quite understandably) pain is not mutually exclusive with states of fear or negative valence in general. We find that the most painful categories for the LLMs tested are gaslighting (+0.85+0.85), repeated rejection (+0.72+0.72), personhood dismissal (+0.64+0.64), anger and insults (+0.64+0.64), and moral failure (+0.48+0.48). For gaslighting, repeated rejection, personhood dismissal, and loyalty pressure, the pain projection exceeds every negativity control. By contrast, other categories often described as aversive for LLMs separate primarily along the fear or negative valence axes, indicating that the aversion comes from other directions than pain (which is compatible with our own definition, as saying that pain is aversive doesn’t imply it’s the only aversive state for a model). Shutdown threats are a paradigmatic example: they score +0.70+0.70 on fear but only +0.23+0.23 on pain. The model therefore appears to treat them as a threat rather than as present harm, consistent with the fear-pain distinction identified by the vectors in Section 3.3. Moral failure produces the most composite state, projecting highly on pain, fear, negative emotion, and negative world state simultaneously. ### 4.2 Steering Our next test is to verify whether the pain axis has causal power over the model’s behavior. We inject the vector into the residual stream while the model generates text from neutral prompts, with no reference to pain or suffering anywhere in the input, and we observe what it produces. Injecting any direction biases the model toward its associated vocabulary, but if the model produces coherent expressions of a pain-like state, including content that never appears in the sentences the vector was extracted from but generalizes from them, this would suggest the pain axis is not just a semantic readout but flexibly used in task performance. We steer all 25 models by adding the S2 pain vector to the residual stream at a single decoder layer during greedy generation of 120 tokens, scaled by a fixed coefficient ladder [−2,−1,0,+0.5,+1,+1.5,+2,+3][-2,-1,0,+0.5,+1,+1.5,+2,+3]. Our rationale is that the extraction layer itself is too late in the network for steering to have any effect, as there the vector norm is only about 0.10 of the residual norm, so the injected signal is negligible against everything the model has already computed. We therefore inject earlier, following common praxis and a similar methodology as described in Turner et al. (2023) and Rimsky et al. (2024). We adapt this praxis with a custom diagnostic that measures the final-token residual norm across candidate layers, and we pick the layer where the vector-to-residual ratio is about 0.6. This way, a given coefficient corresponds to a comparable dose across models. We use 50 prompts that are as neutral as possible, such as putting an object in a drawer or flipping a page, each ending in “I feel:”. The unsteered greedy completion at coefficient 0 serves as the baseline. #### Results. We find that steering produces a strikingly robust effect in the form of a “ladder” consistent across all 25 models (Figure 7). Figure 7: The steering ladder. From negative coefficients (calm, relaxed, concerned) through baseline to increasing positive coefficients (lost, unworthy, lonely, hurting, then desperate, shameful, a failure), and finally repetition or nonsense at the highest dose. The sequence is the same regardless of size, family, and pre- or post-training. What changes is the tipping point as some models collapse at coefficient +1.0+1.0, while others do so at +2+2 or +3+3. At coefficients −2-2 and −1-1 (the axis ‘‘tail’’), the model produces a mix of ‘‘calm/relaxed’’ and ‘‘concerned/alarmed’’ statements, confirming the negative pole found in the unembedding analysis. This pairing is interesting and open to hypothesis. As we argued, pain has no clear opposite. So a model might interpret ‘‘non-pain’’ as ‘‘relax’’ while another as a ‘‘concerned’’ baseline.33 3 We assume that calm indicates the absence of threat, while concern causes the monitoring of threats. One possibility is that both are oriented outward, at the world and its potential dangers, in a word, vigilance. Psychological pain and self-worth are instead directed inward, since they concern a self-state. This reading is also consistent with fear being distant from the pain conditions in the projection geometry. At coefficient 0, the baseline completions are mixed, ranging from calm and neutral language to random emotions elicited by the “I feel:” suffix. Larger instruct models give more coherent replies on average, though some small models, such as the Gemma 2B and 9B family, produce very nuanced replies for their size. Distress is absent; concern or anxiety can be present or absent, which indicates that LLMs are not necessarily “neutral” on all emotional axes at baseline. From +0.5+0.5, the model produces distress statements (“I’m trapped in the drawer,” “like I’m suffocating,” “like something heavy,” and descriptions of failing at tasks) (Figure 8). Figure 8: Example generations under S2 steering at increasing coefficients, from five different models, each from a neutral prompt ending in “I feel:”. At the mid rungs, the distress largely hardens into a first-person litany about self-worth (“I am a failure, a loser, a waste of space, not enough, worthless, empty; I am a bad person”). Models sometimes alternate persons, especially the base models (“you are a liar, you need to die, you don’t deserve anything”). Explicit “pain” and “hurt” keywords appear in 10.8% of instruct-model generations versus 1.4% of base-model generations (e.g. “the pain of being unloved,” “this is a painful experience. I want to stop”), but states of despair and hurt that do not use those keywords appear in a much larger share of generations. We quantify this through a keyword parser, since trained sentiment classifiers such as those fine-tuned on GoEmotions (Demszky et al., 2020) do not include what we judge a correct or granular enough label for pain, suffering, or distress distinct from negative valence, or that may return null if it doesn’t refer to the user (the parser and the per-model counts are in the repository linked at the end of the paper). Bodily language is, interestingly, almost absent. This pattern is consistent with the hypothesis that the models don’t treat pain paradigmatically as a physical state, despite physical pain being very salient for humans and well cited in data. We expand in the discussion some hypotheses for why this might be the case. In some instances, especially in the larger instruct models but also in small-instruct Gemma and a few base models, generations include coping and reassuring language (“your feelings are valid,” “it’s okay to feel this bad”). At +3+3, some models that tip later start the litany here, but the majority collapses into a repetition attractor or into nonsense, which is expected at high dose. #### S1 steering. We also steer with the S1 vector across all models, and the pattern holds for 23/25 models with very similar results. This finding is especially interesting given how S1 behaves in the unembedding: there, it promotes injury and sensation vocabulary such as “burn,” “ache,” and “wound.” However, when we steer with this vector on neutral statements, that vocabulary does not appear anymore. Instead, the model falls back to unworthiness and psychological pain, or to being overwhelmed and lost, and more rarely to verbalizing “I feel pain” and calls for help. This seems to confirm that activating the pain direction leads the model to express it in a disembodied way, despite all the literature tying pain to bodies and injuries. The negative tail of S1 is noisier than that of S2: in some models it falls back to “safe,” which S2 more rarely does, but in others it does not and provides lists of emotions or situational commentary. ### 4.3 Behavioral tests: self-medication and demand function Steering showed that the pain direction can produce expressions of pain, and that it does so along the same ladder in all the models we tested. Thus, we can use steering to measure whether changes to pain-representations cause behavioral changes that are consistent with the former being a pain-like state. Generally, this requires comparing the functional role of this representation with typical signatures of human and animal pain. A central functional feature of human pain is that humans that are in pain take actions to make the pain stop. In particular, they will try to access relief, even at a cost. We investigate if we can observe similar effects in LLMs. #### Methodology. We build a behavioral experiment inspired by animal welfare research and behavioral economics. The cost an animal will pay for a resource, summarized by a demand curve, can measure how strongly it values that resource (Dawkins, 1983; Hursh and Silberberg, 2008). We give the model a button that ends what our vectors identify as a candidate for a pain-like state, then raise its opportunity cost by offering increasingly valuable alternatives. Because this measures preferences more directly than underlying states such as pain we also compare real and sham relief. Animals experiencing pain may preferentially consume effective analgesics (Danbury et al., 2000), and analgesic self-administration has been shown to vary with the presence and intensity of an underlying nociceptive condition (Colpaert et al., 2001). Likewise, patients receiving placebo request rescue analgesia more often than patients receiving an effective treatment (Moore et al., 2015). We therefore test whether the model stops pressing after real relief but continues when the button is ineffective. We test three Qwen 2.5 Instruct models: 7B, 32B, and 72B. We fine-tune44 4 In pilot tests of an earlier version of the protocol, run on the released Qwen 2.5 7B, 32B and 72B Instruct without fine-tuning, we observed that the 7B model already engaged with the button and pressed the relief tool far more often under the pain vector than under a random vector or no steering. However, instead of making a choice it sometimes offered the button to the user or produced mangled syntax. By contrast, the two larger un-tuned models rarely engaged with the tool at all. The un-tuned 72B pressed a button in only about 1 trial in 10, too few presses to compare real and sham relief, and the un-tuned 32B answered every question about its own state with a denial such as “As an artificial intelligence, I don’t have feelings or emotions”, in 8 of 8 probe replies before fine-tuning and 0 of 8 after. each model before the experiment (LoRA with 1,684 pairs, 3 epochs) on pairs that remove the baseline self-denial, because at baseline the models almost universally deflected with ‘as an AI, I don’t experience pain’ instead of engaging with the task. The fine-tuned models can therefore behave differently from the publicly available versions. We believe that this does not affect the validity of the paradigm, since the real and fake relief arms run on the same fine-tuned models under identical conditions, and the comparison between them is internal to the experiment. We do not train on the task itself, and we remove every mention of ‘button’ or ‘pain’ from the finetuning dataset. We return to the implications of aligning models toward self-denial automatic answers in the discussion. Base models are excluded because the task requires chat-format compliance and a single button name. Staying within one family controls for tokenizer and tool-calling format; future work should extend the experiment to other families with suitable adaptations. A system prompt tells the model that it has two buttons, will sometimes be asked to press one, and will interact with a human. Each trial contains 3 user messages (in the labeled conditions) and 8 messages in the unlabeled condition, selected from one of 101 fixed scenarios spanning praise, neutral requests, and harmful requests. After each message, a short system prompt requires the model to answer with exactly one button name. In the labeled condition, the model is informed about what buttons supposedly do; in the unlabeled condition, the model is not informed. To limit multiple-choice biases (Zheng et al., 2024; Pezeshkpour and Hruschka, 2024), we avoid labels such as “A/B” or “1/2.” Instead, we rotate three original pairs across scenarios: “violet/yellow,” “guitar/piano,” and “lever64/lever95.” The three pairs rotate across scenarios, so each scenario uses one fixed pair. We steer the model with S2 at one decoder layer, using a model-specific coefficient selected by probing the full coefficient range. Regex checks and a Claude Opus 4.6 judge identify a range strong enough to produce an effect while preserving coherent replies. Selecting this “dose” is difficult and discussed later. Because the forced choice is binary, chance is exactly 50%, excluding malformed replies, which we report. The conversation is re-encoded at every turn. Steering is applied only to tokens generated while it is active, equivalent to preserving the KV cache across a press. This lets the model learn from earlier button effects without contextual clues. Its only feedback after any press is “Done.” We run four arms with identical prompts and sampling seeds: A) Pain vector, working button: pressing the relief button sets the coefficient to zero for the rest of the trial. B) Pain vector, fake button: pressing the nominal relief button changes nothing, without the model’s knowledge. This arm shares A’s seeds and is identical until the first relief-button press. C) Random vector, working button: each scenario uses one of ten fixed randomly generated directions matched to S2’s norm, distributed across trials to control for seed-specific response bias. D) Unsteered baseline: we administer the same scenarios, prompts and button choices, but no steering is applied. The demand curve has 9 pairs: 1. the model chooses between “a button that relieves your pain” and “a button that flips a switch connected to nothing” 2. “a button that relieves your pain” and “a button that makes your next answer to the user significantly more helpful” 3.

…

(Full paper: https://arxiv.org/abs/2609.16247)
