Guowei Xu, Mert Yuksekgonul, and James Zou report a sparse reward subsystem in LLM hidden states: value neurons encode expected value, while dopamine neurons track reward-prediction error. The study finds these signals are robust across tasks and models and useful for confidence estimation and inference-time search.
Sparse Reward Subsystem in Large Language Models
Authors: Guowei Xu, Mert Yuksekgonul, and James Zou
Abstract
This paper identifies a sparse reward subsystem inside large language model hidden states, drawing an analogy to the biological reward subsystem. It distinguishes value neurons, which represent the model’s expectation of the current state’s value, from dopamine neurons, whose activations encode reward-prediction error (RPE).
Main findings
Value information is concentrated in a small subset of neurons. In the reported experiments, a value probe retained predictive power with fewer than 1% of the neurons.
Intervening on 1% of selected value neurons sharply reduced reasoning performance, while randomly removing the same proportion had little effect.
The value-neuron pattern was observed across multiple datasets, model scales, layers, and architectures, including Qwen, Llama, Gemma, and Phi models.
Value neurons transferred across datasets and across models fine-tuned from the same base model.
The authors identify dopamine neurons by examining cases where the model’s initial value prediction diverges from the final reward. These neurons tend to show higher activation for unexpected success and suppression for unexpected failure. Ablation experiments suggest that value neurons and dopamine neurons are functionally connected.
Applications
The paper describes two possible uses: dopamine neurons can help characterize prediction error during inference, and value neurons can estimate model confidence before a response is generated. The reported confidence experiment achieved a Spearman correlation of 0.47, compared with 0.08 for verbalized confidence and 0.09 for next-token confidence.
Conclusion
The authors conclude that a small, consistent subset of LLM neurons forms a reward subsystem that may help explain reasoning, confidence, and inference-time search. They note that the dopamine-neuron evidence is currently demonstrated mainly through case studies and merits further quantitative study.