Sparse Reward Subsystem in Large Language Models

Authors: Guowei Xu, Mert Yuksekgonul, and James Zou

Abstract

This paper identifies a sparse reward subsystem inside large language model hidden states, drawing an analogy to the biological reward subsystem. It distinguishes value neurons, which represent the model’s expectation of the current state’s value, from dopamine neurons, whose activations encode reward-prediction error (RPE).

Main findings

  • Value information is concentrated in a small subset of neurons. In the reported experiments, a value probe retained predictive power with fewer than 1% of the neurons.
  • Intervening on 1% of selected value neurons sharply reduced reasoning performance, while randomly removing the same proportion had little effect.
  • The value-neuron pattern was observed across multiple datasets, model scales, layers, and architectures, including Qwen, Llama, Gemma, and Phi models.
  • Value neurons transferred across datasets and across models fine-tuned from the same base model.