Base Models Can Reason By Taking a Cue From Training Data
Tl;dr: Two opening tokens can bring a base model’s reasoning performance close to that of its RL-trained counterpart. These token cues come from associations learned during training, and RL makes effective cues more likely. Changing those associations can make even “Chicken” elicit reasoning.
This writeup gives a high-level overview of our paper and aims to build intuition for what it means for a base language model to reason. We walk through some interesting nuggets from our findings, illustrate them with interactive figures, and discuss implications raised by our results. Details can be found in the paper.
Note that these figures use the Olmo3-7B base model, but our results generalize to other model families.
#Background
There is a growing body of work questioning the extent to which base language models learn to reason through next-token prediction, and in particular, what post-training adds. [1] [2]
So, if this capability is already present (i.e., base models already "know how to reason"), then how can we easily access it? Different papers have offered different answers to this question:
- Steering in activation space. [2]
- Random perturbations in weight space. [3]
- Post-training with random rewards. [4]
- Alternative decoding strategies. [5] [6]
In this paper, we ask if there is an even simpler way. The answer is yes (just set the first few tokens!), and we share an explanation for why (it's in the data!).
#Token Cues Elicit Reasoning in Base Models
This work is motivated by our observation that a base model’s accuracy is strongly correlated with starting on certain tokens. This made us wonder, what if we force these opening tokens, i.e., token cues, and let the model generate the rest?
We find that prefilling token cues can bring base-model reasoning performance close to that of RL-trained models. For example, prefilling .\n\nOkay raises Olmo-3-7B’s MATH-500 accuracy from 41.5% to 76.9%, comparable to its RL-trained counterpart’s 74.9%. Yet, these effective openings are not necessarily the ones that the model is most likely to generate. You can explore the effect of different opening tokens below.
Loading explorer…
#RL Exploits Token Cues
If there exist token cues that can bring base models close to their RL-trained counterparts in reasoning performance, then what is RL doing?
We consider the RL-Zero setting [7], using reinforcement learning with verifiable rewards (RLVR) and the Group Relative Policy Optimization (GRPO) algorithm. The base model learns from rewards on its own responses rather than being taught with new demonstrations of how to reason. We find that after RL, the model becomes more likely to start with cues that already elicit reasoning in the base model, such as .\n\nOkay. You can explore this change below.
Loading explorer…
In fact, the largest changes in next-token predictions after RL occur at the first two token positions. After the first two tokens, the two models largely agree on what comes next. Below, you can explore how the base and RL models predict the next token given exactly the same preceding text.
#Token Cues Come From Training Data
Where do these token cues come from? These tokens don’t explicitly tell the base model to reason, yet they yield drastic changes in its response and performance. Still, it might not seem so surprising that “Okay” works, given how one might start thinking through a problem with “Okay, let’s see”. More surprisingly though, we find that simple training-data associations can make even an arbitrary token like “Chicken” elicit reasoning.
We construct a training-data counterfactual by replacing “Okay” with “Chicken” in the 100B-token mid-training mix, which contains many synthetic reasoning traces. We resume training from the same checkpoint on the base and edited data mixes. With the edited mix, “Chicken” elicits reasoning, raising its cued accuracy from 18.2% to 76.9%. Below, you can explore how this data edit changes the effect of prefilling .\n\nChicken.
“Chicken” can now elicit reasoning when solving a problem, but what about its original meaning? The model trained on the edited data mix still uses it to refer to a bird or food. For instance, it completes “A chicken is a” with “bird that belongs to the family of Galliformes,” and “She roasted the chicken” with “and made a salad.”
#Implications
What does it mean, then, for a language model to learn to “reason”? We tend to think of reasoning as a complex ability, so it seems counterintuitive that whether a model uses it can hinge on something as simple as a seemingly arbitrary token. Our findings invite us to consider what this could mean for how we train and evaluate language models, and how we interpret their behavior.
###Post-training
We and many others have wondered about the extent to which base models already know how to reason, and what post-training adds. When a model fails to solve a problem, it’s hard to tell whether it lacks the ability or whether that ability wasn’t elicited. Thus, while evaluating the effectiveness of post-training, we should separately consider what a model can do and what a model is likely to do.
This also raises a design question. If we can elicit a desired behavior from a base model through additional conditioning, when should we instead post-train to make that behavior more likely? What do we gain by sharpening the response distribution, and what other useful behaviors might become harder to elicit?
###Training Data
We often think about training data in terms of what it teaches a model to do. But how it shapes when a model uses what it has learned may be less front of mind. Token cues offer a concrete example. In our case, repeated openings in synthetic reasoning traces inadvertently reinforced a strong association between particular tokens and reasoning behavior. This suggests that design choices about training data can be incredibly consequential.
For example, do we want models to reason reliably across all sorts of wordings, and if so, should we increase data diversity to encourage that robustness? Or should we exploit these associations to reliably elicit particular behaviors with particular cues?
###The Meaning of Words
A, perhaps, more philosophical implication is about what words mean to language models. Phrases like “Let’s think step by step” or “Wait” seem to elicit reasoning in ways that fit our understanding of them: one asks for a step-by-step solution, while the other suggests pausing or reconsidering. It feels natural to explain their effects through those meanings. However, it is surprising that even a word like “Chicken,” with no apparent connection to reasoning, can elicit it through associations learned during training.
The reason why this might also be surprising is because “chicken” refers to a domesticated bird kept for its eggs or meat. While the model can still express this meaning in everyday contexts, it can also learn to associate the word with reasoning behavior, substantially improving its accuracy when cued with it. This leaves us wondering what, if anything, distinguishes learning correlations among words from understanding what they mean.
“You shall know a word by the company it keeps!”
— J. R. Firth (1957) [8]
References
[1]
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Yang Yue et al. · 2025
[2]
Base Models Know How to Reason, Thinking Models Learn When
Constantin Venhoff et al. · 2025
[3]
Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
Yulu Gan and Phillip Isola · 2026
[4]
Spurious Rewards: Rethinking Training Signals in RLVR
Rulin Shao et al. · 2025
[5]
Chain-of-Thought Reasoning Without Prompting
Xuezhi Wang and Denny Zhou · 2024
[6]
Reasoning with Sampling: Your Base Model Is Smarter Than You Think
Aayush Karan and Yilun Du · 2025
[7]
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · 2025
[8]
A Synopsis of Linguistic Theory, 1930–1955
J. R. Firth · 1957
BibTeX
@misc{wang2026basemodelsreasontaking,
title={Base Models Can Reason By Taking a Cue From Training Data},
author={Sophie L. Wang and Amil Dravid and Rulin Shao and Kevin Farhat and Sewon Min and Alexei A. Efros},
year={2026},
eprint={2610.06851},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2610.06851},
}