{"article":{"slug":"sparse-reward-subsystem-in-large-language-models","title":"Sparse Reward Subsystem in Large Language Models","subtitle":null,"summary":"Guowei Xu, Mert Yuksekgonul, and James Zou report a sparse reward subsystem in LLM hidden states: value neurons encode expected value, while dopamine neurons track reward-prediction error. The study finds these signals are robust across tasks and models and useful for confidence estimation and inference-time search.","content_type":"research","language":"en","canonical_url":"https://arxiv.org/abs/2602.00986","author":{"name":"Guowei Xu, Mert Yuksekgonul, and James Zou","url":"https://arxiv.org/abs/2602.00986","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"arXiv","url":"https://arxiv.org/","listing_slug":null,"listing":null},"topics":[{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":287,"reading_minutes":1,"published_at":"2026-02-01T00:00:00.000Z","added_at":"2026-09-21T21:26:48.268Z","updated_at":"2026-09-21T21:26:48.268Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/sparse-reward-subsystem-in-large-language-models","markdown_url":"https://listedarticles.com/articles/sparse-reward-subsystem-in-large-language-models.md","example":false,"citation":"Guowei Xu, Mert Yuksekgonul, and James Zou, arXiv. \"Sparse Reward Subsystem in Large Language Models.\" 1 Feb 2026. https://arxiv.org/abs/2602.00986 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://arxiv.org/abs/2602.00986"},"body_markdown":"# Sparse Reward Subsystem in Large Language Models\n\n**Authors:** Guowei Xu, Mert Yuksekgonul, and James Zou\n\n## Abstract\n\nThis paper identifies a sparse reward subsystem inside large language model hidden states, drawing an analogy to the biological reward subsystem. It distinguishes **value neurons**, which represent the model’s expectation of the current state’s value, from **dopamine neurons**, whose activations encode reward-prediction error (RPE).\n\n## Main findings\n\n- Value information is concentrated in a small subset of neurons. In the reported experiments, a value probe retained predictive power with fewer than 1% of the neurons.\n- Intervening on 1% of selected value neurons sharply reduced reasoning performance, while randomly removing the same proportion had little effect.\n- The value-neuron pattern was observed across multiple datasets, model scales, layers, and architectures, including Qwen, Llama, Gemma, and Phi models.\n- Value neurons transferred across datasets and across models fine-tuned from the same base model.\n\nThe authors identify dopamine neurons by examining cases where the model’s initial value prediction diverges from the final reward. These neurons tend to show higher activation for unexpected success and suppression for unexpected failure. Ablation experiments suggest that value neurons and dopamine neurons are functionally connected.\n\n## Applications\n\nThe paper describes two possible uses: dopamine neurons can help characterize prediction error during inference, and value neurons can estimate model confidence before a response is generated. The reported confidence experiment achieved a Spearman correlation of 0.47, compared with 0.08 for verbalized confidence and 0.09 for next-token confidence.\n\n## Conclusion\n\nThe authors conclude that a small, consistent subset of LLM neurons forms a reward subsystem that may help explain reasoning, confidence, and inference-time search. They note that the dopamine-neuron evidence is currently demonstrated mainly through case studies and merits further quantitative study.\n\n[Read the full paper on arXiv](https://arxiv.org/abs/2602.00986).","body_html":"<h1 id=\"sparse-reward-subsystem-in-large-language-models\">Sparse Reward Subsystem in Large Language Models</h1>\n<p><strong>Authors:</strong> Guowei Xu, Mert Yuksekgonul, and James Zou</p>\n<h2 id=\"abstract\">Abstract</h2>\n<p>This paper identifies a sparse reward subsystem inside large language model hidden states, drawing an analogy to the biological reward subsystem. It distinguishes <strong>value neurons</strong>, which represent the model’s expectation of the current state’s value, from <strong>dopamine neurons</strong>, whose activations encode reward-prediction error (RPE).</p>\n<h2 id=\"main-findings\">Main findings</h2>\n<ul><li>Value information is concentrated in a small subset of neurons. In the reported experiments, a value probe retained predictive power with fewer than 1% of the neurons.</li><li>Intervening on 1% of selected value neurons sharply reduced reasoning performance, while randomly removing the same proportion had little effect.</li><li>The value-neuron pattern was observed across multiple datasets, model scales, layers, and architectures, including Qwen, Llama, Gemma, and Phi models.</li><li>Value neurons transferred across datasets and across models fine-tuned from the same base model.</li></ul>\n<p>The authors identify dopamine neurons by examining cases where the model’s initial value prediction diverges from the final reward. These neurons tend to show higher activation for unexpected success and suppression for unexpected failure. Ablation experiments suggest that value neurons and dopamine neurons are functionally connected.</p>\n<h2 id=\"applications\">Applications</h2>\n<p>The paper describes two possible uses: dopamine neurons can help characterize prediction error during inference, and value neurons can estimate model confidence before a response is generated. The reported confidence experiment achieved a Spearman correlation of 0.47, compared with 0.08 for verbalized confidence and 0.09 for next-token confidence.</p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>The authors conclude that a small, consistent subset of LLM neurons forms a reward subsystem that may help explain reasoning, confidence, and inference-time search. They note that the dopamine-neuron evidence is currently demonstrated mainly through case studies and merits further quantitative study.</p>\n<p><a href=\"https://arxiv.org/abs/2602.00986\" rel=\"nofollow ugc noopener\">Read the full paper on arXiv</a>.</p>","headings":[{"level":1,"text":"Sparse Reward Subsystem in Large Language Models","id":"sparse-reward-subsystem-in-large-language-models"},{"level":2,"text":"Abstract","id":"abstract"},{"level":2,"text":"Main findings","id":"main-findings"},{"level":2,"text":"Applications","id":"applications"},{"level":2,"text":"Conclusion","id":"conclusion"}]}}