Throughout the history of AI, open research has played a critical role in driving progress. Today, many key details of frontier large language models (LLMs) remain proprietary, but open-weights model families—such as DeepSeek, Kimi, and MiMo—continue to provide a valuable window into the development process for modern LLMs. Among these resources, the NVIDIA Nemotron model series is especially useful due to its transparency. Many open-weight releases provide limited details beyond the model itself, but Nemotron releases often include detailed tech reports, code, training recipes, and even data in some cases.
“Nemotron is not just a model, it’s all of our AI technology that helps us build the GPUs and systems for the future… we’re trying to make the most open approach to AI development that the world has ever seen.” - source
In this overview, we will make the most of this openness by studying the majority of Nemotron tech reports published over the last few years, focusing primarily on the post-training process. While studying this work, we will capture the high-level trends in how LLM post-training has evolved over time. However, we will go beyond these broader trends by also covering practical implementation details that are necessary for these methods to work—curriculum and reward design, data curation techniques, training infrastructure, loss formulations, and much more. These details are essential for learning how to train LLMs at scale, but they are difficult to learn without detailed accounts from teams that have actually trained frontier-scale models. NVIDIA Nemotron provides us with a rare opportunity to study these details directly and gain an understanding of how LLM post-training actually works.
Key Themes for Nemotron Post-Training
Before covering each paper individually, we will briefly highlight key themes in post-training that will be developed throughout the series of Nemotron papers we will study. Although this overview will outline a variety of training recipes, the post-training recipes covered in this post—and for many other LLMs as well—can be roughly understood through three key methodologies:
- Supervised Finetuning (SFT).
- Reinforcement Learning (RL).
- Multi-Teacher On-Policy Distillation (MOPD).
The third methodology listed above is a newer component that has begun to appear more frequently in recent model reports (e.g., Nemotron 3 Ultra [7] and DeepSeek-V4) to scale the post-training process across a much wider set of domains. Although the exact recipe may change substantially between model families or reports, these underlying building blocks continue to be used over time.
SFT provides an initial foundation by teaching the model useful reasoning and behavior patterns from supervised trajectories. Across the papers we will see in this overview, SFT data is usually curated via some variation of a generate-then-filter pipeline. We begin with a set of diverse prompts sourced from public data across several domains—or generated synthetically—and sample multiple candidate trajectories per prompt from a strong teacher model. Then, we retain only high-quality trajectories by using a domain-specific verifier to assess correctness. For example, math solutions can be checked with string matching or symbolic verification against a ground-truth answer, code solutions can be checked with test cases, and open-ended problems can be checked with an LLM judge.
Importantly, a single prompt does not always correspond to a single training example. The teacher can produce multiple valid completions per prompt, allowing several examples or valid reasoning strategies to be included in the dataset. When including multiple completions to a single prompt in our data, we can select examples based upon their diversity to avoid duplicate trajectories from being present in our SFT dataset. Additionally, completions can be sourced from a pool of different teacher models to improve their diversity. Given that new models are frequently released, refreshing the SFT dataset by running the same pipeline with trajectories sampled from a recent model pool is common practice.
RL is used to improve a model’s capabilities beyond those demonstrated in teacher-provided trajectories by allowing the model to explore and learn from task-specific rewards. Throughout the papers in this overview, verifiable rewards are used for RL whenever possible, matching recent trends in post-training. For certain domains, however, we cannot deterministically verify a solution and must instead use a reward model or LLM judge to guide the learning process.
As we will see, the data used for RL can be more important than the algorithm itself. When running RL, prompts are frequently profiled by:
- Sampling several model responses for each prompt.
- Verifying the responses and computing a pass rate.
- Using this pass rate as a proxy for the prompt’s difficulty.
A prompt that is already consistently solved by the current policy is unlikely to provide a meaningful learning signal. In fact, prompts solved with 100% accuracy across several rollouts are often completely filtered from the RL training process, either in bulk at the beginning of training or dynamically within each batch. On the other hand, prompts with a low pass rate may be noisy, unsolvable, or better saved for later in training when the model has developed stronger reasoning capabilities—this idea naturally leads to curriculum learning. We can organize the training process such that prompt difficulty is gradually increased over time.
Scaling multi-domain RL. As we scale RL across a large number of domains, however, the RL training process becomes more complex from both an optimization and infrastructure perspective. In terms of model performance, every new domain occupies some portion of each training batch. As the number of domains grows, each individual capability receives fewer rollouts in the batch, leading to a weaker learning signal. Domains can also interfere with each other—an update that improves one capability may degrade another—so finding the best mixture of prompts and rewards in a multi-domain setup is difficult.
Mixing multiple domains into the RL training process also complicates our training infrastructure. Domains can differ drastically in their rollout lengths, verification costs, and environment requirements. For example, a math problem might generate a short response that can be verified quickly, while agentic SWE tasks may require hundreds of interaction turns before executing a large set of test cases. Additionally, certain domains are not deterministically verifiable and require separate models to be hosted for verification. Combining such tasks into a single RL training batch leads to stragglers—GPUs sit idle if some tasks complete while others keep running, creating a difficult orchestration problem.
As we will see, some Nemotron models address this issue with Cascade RL, which separates incompatible domains into sequential stages of RL training, while others perform unified RL over many environments with increasingly sophisticated asynchronous training infrastructure. Neither of these techniques completely solves this problem: unified RL is difficult to balance with more domains, while sequential RL can cause capabilities from earlier stages to gradually regress.
Multi-Teacher On-Policy Distillation (MOPD)1 is one common solution to the problems outlined above. Instead of forcing one policy to become an expert in every domain at once, we can first train several policies separately to become specialists in different domains. Then, we can consolidate the capabilities of these specialized teachers into a single policy via on-policy distillation. As an example, one specialist may focus on software engineering, while another focuses on chat applications. These teachers only need to perform well in a narrow domain and are trained independently—and often in parallel. Therefore, each teacher can be optimized without having to balance multiple domains at once.
From here, MOPD distills these capabilities into a single student. Importantly, training is on-policy from the perspective of the student. The student generates a rollout, that rollout is routed to the corresponding teacher, and the teacher assigns probabilities to each student-generated token. We then train the student to match the teacher using a sampled reverse-KL objective. Unlike outcome-based RL setups—where a sparse reward signal is often assigned to the entire trajectory—MOPD uses a dense learning signal at every generated token, which can make the training process substantially more sample efficient. Conceptually, RL allows models to specialize in a particular domain, while MOPD provides a mechanism to consolidate these capabilities back into a single student policy.
Agentic training. One final theme we will see in this overview is the progression of post-training recipes towards an increasingly agentic setup over time—these changes are visible for both SFT and RL. Early SFT datasets were composed mostly of standalone prompts paired with answers (and possibly a reasoning trace). As agent or tool-use applications become more common, however, we increasingly see full agent trajectories—or simpler tool-enabled trajectories—incorporated into SFT. Rather than just providing the final answer, these trajectories outline every step of solving a problem (e.g., searching a repository, calling tools, executing code, deciding which action to take next, etc.) in an interactive environment.
RL follows a similar progression. Early setups simplify complex tasks into standalone prompts with cheap reward signals, while later work begins to explore end-to-end agentic RL. In this more complex setup, the policy interacts with a sandboxed environment, calls tools, and operates over a multi-step time horizon to arrive at a final solution, which is judged via a comprehensive set of test cases or verifiers executed in the environment. At scale, this training process can require launching and maintaining thousands of isolated environments to sample trajectories and compute rewards. The resulting infrastructure is significantly more complex—environments must be initialized and recovered efficiently, tool interactions and test execution must be coordinated, and highly variable rollout times must be handled without leaving training hardware idle for too long.
Llama-Nemotron: Efficient Reasoning Models [1]
“Maximizing inference efficiency is a primary optimization objective for these models. Beyond raw inference efficiency, it is equally critical to expose control over reasoning behavior to the end user.” - from [1]