{"article":{"slug":"training-search-agents-with-grpo","title":"Training Search Agents with GRPO","subtitle":null,"summary":"Hands-on introduction to reinforcement learning by training a search agent with group-relative policy optimization (GRPO), with open rollouts, code, and reward-design lessons for LLM search.","content_type":"tutorial","language":"en","canonical_url":"https://jasperlu.com/blog/training-search-agents-grpo/","author":{"name":"Jasper Lu","url":"https://jasperlu.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Jasper Lu","url":"https://jasperlu.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":6446,"reading_minutes":28,"published_at":"2026-08-24T12:00:00.000Z","added_at":"2026-09-17T15:35:50.819Z","updated_at":"2026-09-17T15:35:50.819Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/training-search-agents-with-grpo","markdown_url":"https://listedarticles.com/articles/training-search-agents-with-grpo.md","example":false,"citation":"Jasper Lu, Jasper Lu. \"Training Search Agents with GRPO.\" 24 Aug 2026. https://jasperlu.com/blog/training-search-agents-grpo/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://jasperlu.com/blog/training-search-agents-grpo/"},"body_markdown":"This post is a hands-on introduction to reinforcement learning through training a search agent with group-relative policy optimization (GRPO). Search is a fun place to learn RL because it has so many levers, and each of them visibly changes how the model searches. It is also a domain where, in my experience, a well-designed reward function can influence how a model searches more effectively than system prompt changes or harness engineering.\n\nOur approach is based loosely on the Harness-1 paper from Jiang et al. We use their training dataset and a much-simplified version of their search harness.\n\nTraining is done using the Tinker API from Thinking Machines. Many thanks to the team for the training credits to support this work.\n\nThe dataset\n\nWe use the RL split of the Harness-1 dataset for all experiments. The dataset contains 3,453 queries over 37,113 SEC filings split into 2,115,106 chunks.\n\nEach query includes a preamble followed by a question. The agent must both retrieve evidence for facts in the preamble and the answer to the question. In the canonical version of this task, retrieving the answer is more important than evidence for facts. However, for simplicity we will treat both the same.\n\nTo keep initial ablations inexpensive, we start with just 256 training queries and 32 evaluation queries over a reduced corpus of 124,395 chunks. For our scaled runs, we expand to 1,024 training queries over 237,533 chunks.\n\nThe harness\n\nOur agent operates as a search subagent. Given a query, it is meant to retrieve all relevant docs. We keep the harness deliberately simple, only providing lexical search tools. The agent’s output is a curated set it builds up as it goes rather than a single submission at the end. We chose this because it works well for a fixed turn budget: since the set exists from the start, an episode that runs out of turns will will return something gradable.\n\n| Tool | What it does |\n\n|---|---|\n\n| bm25_search | Keyword search over every chunk. Returns an id, title, score and snippet per hit, plus a running count of chunks seen and curated so far. |\n\n| grep_corpus | A short Python regex scan across the corpus. Returns an id, title, and snippet per hit. |\n\n| read_document | The title and text of one chunk. Only ids the agent has already seen in a result list can be read. |\n\n| curate | Add chunk ids to the result set. |\n\n| drop_curated | Remove ids from the result set. |\n\n| finish | End the episode and return the set. |\n\nThe system prompt is similarly plain. We name the tools and give a general strategy for search. We don’t try to encourage any specific search characteristics (e.g. high recall or high precision) so that the agent can discover the right behaviors during training.\n\nRetrieval agent system prompt\n\nYou are a retrieval subagent in a multi-agent system. Your specific role is to identify and retrieve the most relevant documents from a large corpus to help another agent answer questions. You do NOT answer questions yourself - you only find and retrieve relevant documents.\n\nThe user message contains the query you need to find documents for.\n\n**Available Tools:**\n\n- bm25_search: Keyword search over the corpus\n\n- grep_corpus: Text pattern matching with a short Python regex\n\n- read_document: Read a specific document that looks promising but incomplete\n\n- curate: Add relevant documents to your curated result set\n\n- drop_curated: Remove curated documents that turned out to be irrelevant\n\n- finish: End the search and return your curated set\n\n**Your Process:**\n\n- Break down the query into its key concepts and information needs (list each one explicitly)\n\n- For each key concept, develop a specific search strategy that targets that concept\n\n- Consider what types of documents and evidence would be most helpful for answering this query\n\n- Plan several distinct, non-overlapping search strategies that approach the question from different angles\n\n- Then execute your searches using multiple parallel tool calls.\n\n**Your Thinking:**\n\n## After each round of searches, consider the following\n\n- **What do I know?**: List the key topics, themes, or aspects of the question that your curated documents address. What specific information do you have?\n\n- **What should I search for next?**: Systematically consider what search approaches, keywords, or document types you haven't yet tried that might yield valuable information.\n\n- **What should I curate or drop?**: Curate documents as soon as you have verified they are relevant; drop curated documents that later look redundant or off-topic.\n\n- **Do I have enough information?**: Given the question's complexity and requirements, do you have sufficient information to help answer it, or are there critical gaps?\n\n- Decide if additional searches are needed (and if so, ensure they use genuinely different approaches and do not duplicate prior searches)\n\n- Avoid getting stuck on a single search strategy - if one approach isn't yielding results, backtrack and try different approaches\n\n**Tactics to Consider:**\n\n- When queries fail, try different approaches or keywords to improve the results\n\n- Avoid duplicate or redundant searches\n\n- Execute multiple tool calls in parallel when possible\n\n- Focus on gathering as much relevant information as possible; it is useful to get multiple perspectives on the same topic to confirm the information you have found is correct\n\n- Follow explicit textual evidence rather than speculation\n\n**Output (IMPORTANT):**\n\n- YOU MUST use the curate tool. This is the ONLY way to return documents.\n\n- Every response must contain at least one tool call. Never reply with plain text, and never end a response while still planning a search - emit that search as a tool call instead.\n\n- As soon as you verify a document is relevant, call curate with its ID.\n\n- When your curated set covers the query's information needs, call finish. Your curated set is scored even if you run out of turns, but finishing cleanly is always better than timing out.\n\nWe use gpt-oss-20b for all training runs. We chose this model because it is cheap to sample and train, allowing us to do many ablations.\n\nQuick GRPO recap\n\n## We train with Dr. GRPO, a variant of GRPO. The loop is as follows\n\n- Take a query and run the agent on it N times using the current policy (in this setting, essentially the model’s weights). Each run is a rollout, and the N tries together form a group.\n\n- Score each rollout. This score is its reward.\n\n- Subtract the group’s mean reward from each rollout’s reward. The resulting difference is that rollout’s advantage.  In vanilla GRPO, we would also divide each advantage by the group’s standard deviation. Dr. GRPO drops this term to remove a bias: a query where the eight rollouts nearly agree has a small deviation. Dividing by standard deviation scales up the advantages, so the update can end up dominated by queries the policy is already consistent on.\n\n- Update the policy (model weights) proportionally to the rollout’s advantage, so that the tokens in higher-advantage rollouts become more likely and tokens in lower-advantage rollouts become less likely.\n\nWe would also typically use ratio clipping during the policy update to help stabilize training. However, to simplify training for the sake of pedagogy, we will train in a fully synchronous setting with zero training-inference mismatch (which Tinker provides by default) and only do one update per batch.\n\nWe won’t go deeper into the math behind here, but there are many other great resources out there for understanding GRPO.\n\nThe reward\n\nEach task here consists of a query, a corpus, a set of facts, and a set of supporting documents for each fact. The agent has to curate a set of documents it determines to be relevant to those facts, and success is evaluated on how well this curated set covers the target facts.\n\nA reward function is essentially a kind of evaluation scheme for a rollout. A few metrics common to search retrieval work naturally here.\n\nRecall measures the fraction of relevant documents our agent has curated. Since our search problem is fact-based, this will be the fraction of relevant facts the curated chunks cover. Precision measures the fraction of curated documents that are actually relevant.\n\nWe can’t use just recall or precision alone in our reward function. High recall is desirable in our problem space, but can be trivially maximized by curating every document in the corpus. This is useless for an end user! Why use a search agent at all? High precision can be achieved just as trivially by curating only a single correct chunk per query, which is useless in the other direction.\n\nWe need a metric that punishes both forms of degenerate behavior. F1, the harmonic mean of precision and recall, is the usual choice here:\n\nF1=P+R2PR\n\nThis metric is only high when both precision and recall are high. This is a totally reasonable metric to choose for our reward function. But thinking about our problem space again… it’s likely that a missing document here costs far more to an end user than an extra, useless one. So, we might want to bias the reward towards recall. With this in mind, we will use F-beta as our reward function. This is a weighted version of F1 with a coefficient for setting how much recall matters more than precision:\n\nFβ=β2P+R(1+β2)PR\n\nWe set β = 4 for our reward function, which weighs recall sixteen times more than precision. I chose this number based on my prior experience and notes from the Harness-1 paper rather than a formal ablation. Under this metric, which we call F4, a set with recall 1.0 and precision 0.25 scores 0.85; under F1 the same set would score 0.40.\n\nOne more thing we add to our reward function is a punishment for   not curating anything—an episode that curates nothing at all gets a flat -0.2. This helps with differentiating between rollouts where the model never gets to curate anything and ones where it only curates incorrect documents.\n\nPutting it together, the reward for a rollout with curated set C is\n\nr(C)={−0.2F4if ∣C∣=0otherwise\n\nInitial explorations\n\nThe core idea behind reinforcement learning is to let models explore and reward them for behaviors that lead to good results. This means that RL works specifically by reinforcing behaviors the model already produces. If the model either always reaches for the same actions or always fails, then there is nothing for it to learn from.\n\nThis is especially true in GRPO, where learning depends on variation in rewards within each group. If all rollouts receive the same reward, their advantages are zero and the group provides no gradient.\n\nThis is why it can be helpful to gauge the RL-ability of a task before training. Best-of-N evals are a good proxy for this. Just sample N runs over the same task and report the highest score. A good candidate for RL is a model that has noticeably higher best-of-N scores than best-of-1 — this gap is the headroom RL can close by just sharpening the existing policy.\n\nBelow, we sample gpt-oss-20b four times on each of the eval queries. All rollouts are viewable.\n\nThe gap between best-of-4 and single trial F1 is large: best-of-4 F1 is 0.323 vs 0.195 for a single try. The F1 score is also precision-dominated. The model is likely being over-conservative in its curation.\n\nClose inspection of the rollout supports this. On seven of the 32 queries, none of the four tries ever calls curate. The model stalls, exhausts all its turns, or runs out of context before curating a single document. Failures like these signal more of a harness problem than a search problem, and that is the kind of behavior reinforcement learning is good at fixing.\n\nLearning rate sweep\n\nLearning rate is a critical hyperparameter to get right. Just like in supervised fine-tuning, a bad learning rate can destabilize training. In reinforcement learning, a learning rate too high can cause the policy to collapse; one that is too small means the model never learns anything during training. And even when a run looks fine, I’ve found that further tuning the learning rate can often lead to incrementally better performance.\n\nThat said, if you’re an experienced practitioner and know directionally what the right hyperparameters are, you can skip this step for initial experiments. My opinion is that once training works stably, tuning the learning rate only starts to matter when you’re trying to squeeze out extra performance from your data.\n\nAll runs in this section and the next use the 256-query subset of our larger dataset, with 32 groups of 8 rollouts per step. Here are the results of a sweep over five different learning rates, each run for one epoch:\n\ntrain reward\n\ntrain reward | step | 5e-6 | 1e-5 | 5e-5 | 1e-4 | 5e-4 |\n\n|---|---|---|---|---|---|\n\n| 0 | -0.02 | -0.02 | -0.02 | -0.02 | -0.02 |\n\n| 1 | -0.03 | -0.03 | -0.02 | -0.02 | 0.01 |\n\n| 2 | -0.04 | -0.04 | -0.04 | -0.03 | -0.03 |\n\n| 3 | -0.02 | -0.03 | -0.02 | 0.00 | 0.05 |\n\n| 4 | 0.01 | -0.01 | 0.02 | 0.05 | 0.12 |\n\n| 5 | -0.01 | 0.02 | 0.03 | 0.09 | 0.16 |\n\n| 6 | -0.01 | 0.01 | 0.04 | 0.12 | 0.12 |\n\n| 7 | -0.04 | -0.02 | 0.02 | 0.15 | 0.01 |\n\n| 8 | -0.03 | -0.03 | 0.06 | 0.14 | 0.08 |\n\nvalid train rollouts\n\nvalid train rollouts | step | 5e-6 | 1e-5 | 5e-5 | 1e-4 | 5e-4 |\n\n|---|---|---|---|---|---|\n\n| 0 | 43% | 43% | 43% | 43% | 43% |\n\n| 1 | 44% | 39% | 42% | 44% | 46% |\n\n| 2 | 36% | 34% | 34% | 42% | 53% |\n\n| 3 | 40% | 39% | 39% | 45% | 73% |\n\n| 4 | 48% | 45% | 53% | 54% | 91% |\n\n| 5 | 45% | 47% | 51% | 66% | 100% |\n\n| 6 | 40% | 43% | 52% | 63% | 100% |\n\n| 7 | 37% | 41% | 55% | 79% | 75% |\n\n| 8 | 40% | 38% | 63% | 88% | 100% |\n\neval f1\n\neval f1 | step | 5e-6 | 1e-5 | 5e-5 | 1e-4 | 5e-4 |\n\n|---|---|---|---|---|---|\n\n| 0 | 0.19 | 0.19 | 0.19 | 0.19 | 0.19 |\n\n| 4 | 0.18 | 0.21 | 0.22 | 0.20 | 0.12 |\n\n| 8 | 0.16 | 0.14 | 0.16 | 0.19 | 0.03 |\n\nTraining results over one epoch with different learning rates. The valid rollouts chart shows the percentage of training rollouts that curate at least one doc.\n\nWe can see in here the effects of both a too large and a too small learning rate:\n\n- 1e-5 and5e-6 barely move train reward. We clearly need a larger gradient signal to meaningfully update the policy.\n\n- 5e-4 improves reward rapidly at first and then quickly destabilizes.\n\n- 1e-4 and5e-5 train steadily, with1e-4 resulting in a higher final train reward and a higher eval F1.\n\nSince lr=1e-4 performs the best in the sweep, we’ll use it for the rest of our experiments.\n\nIf you want a better idea of what the training collapse for lr=5e-4 looks like, expand the explorer below to see its eval rollouts at steps 0 and 8.\n\nBatch size ablation\n\nConventional wisdom when training with GRPO is to use large batch sizes, where batch size is groups per step x group size. To study the effects of batch size on training here, we fix group size at 8 and ablate over groups per step.\n\nThe charts below are plotted against prompts seen rather than total updates, since all runs have a different total number of steps.\n\ntrain reward\n\ntrain reward | step | 8 groups/step | 16 groups/step | 32 groups/step | 64 groups/step |\n\n|---|---|---|---|---|\n\n| 0 | -0.03 | -0.03 | -0.03 | -0.03 |\n\n| 8 | -0.01 |  |  |  |\n\n| 16 | 0.04 | 0.02 |  |  |\n\n| 24 | 0.04 |  |  |  |\n\n| 32 | -0.06 | -0.07 | -0.01 |  |\n\n| 40 | -0.06 |  |  |  |\n\n| 48 | -0.00 | -0.06 |  |  |\n\n| 56 | 0.04 |  |  |  |\n\n| 64 | 0.25 | 0.04 | -0.04 | -0.05 |\n\n| 72 | 0.03 |  |  |  |\n\n| 80 | 0.06 | -0.01 |  |  |\n\n| 88 | 0.11 |  |  |  |\n\n| 96 | 0.14 | 0.09 | -0.01 |  |\n\n| 104 | 0.03 |  |  |  |\n\n| 112 | 0.07 | 0.04 |  |  |\n\n| 120 | 0.13 |  |  |  |\n\n| 128 | 0.21 | 0.20 | 0.01 | -0.01 |\n\n| 136 | 0.24 |  |  |  |\n\n| 144 | 0.16 | 0.21 |  |  |\n\n| 152 | 0.17 |  |  |  |\n\n| 160 | 0.02 | 0.10 | 0.10 |  |\n\n| 168 | 0.10 |  |  |  |\n\n| 176 | 0.13 | 0.14 |  |  |\n\n| 184 | 0.19 |  |  |  |\n\n| 192 | 0.09 | 0.14 | 0.08 | 0.07 |\n\n| 200 | 0.16 |  |  |  |\n\n| 208 | 0.14 | 0.16 |  |  |\n\n| 216 | 0.07 |  |  |  |\n\n| 224 | 0.08 | 0.11 | 0.12 |  |\n\n| 232 | 0.13 |  |  |  |\n\n| 240 | 0.06 | 0.07 |  |  |\n\n| 248 | 0.05 |  |  |  |\n\n| 256 | 0.17 | 0.12 | 0.16 | 0.05 |\n\nvalid train rollouts\n\nvalid train rollouts | step | 8 groups/step | 16 groups/step | 32 groups/step | 64 groups/step |\n\n|---|---|---|---|---|\n\n| 0 | 41% | 41% | 41% | 41% |\n\n| 8 | 36% |  |  |  |\n\n| 16 | 55% | 45% |  |  |\n\n| 24 | 60% |  |  |  |\n\n| 32 | 39% | 43% | 44% |  |\n\n| 40 | 66% |  |  |  |\n\n| 48 | 72% | 48% |  |  |\n\n| 56 | 85% |  |  |  |\n\n| 64 | 98% | 61% | 42% | 40% |\n\n| 72 | 92% |  |  |  |\n\n| 80 | 97% | 62% |  |  |\n\n| 88 | 100% |  |  |  |\n\n| 96 | 98% | 72% | 45% |  |\n\n| 104 | 100% |  |  |  |\n\n| 112 | 100% | 75% |  |  |\n\n| 120 | 100% |  |  |  |\n\n| 128 | 100% | 90% | 54% | 40% |\n\n| 136 | 100% |  |  |  |\n\n| 144 | 100% | 94% |  |  |\n\n| 152 | 100% |  |  |  |\n\n| 160 | 100% | 92% | 66% |  |\n\n| 168 | 100% |  |  |  |\n\n| 176 | 100% | 95% |  |  |\n\n| 184 | 100% |  |  |  |\n\n| 192 | 100% | 94% | 63% | 51% |\n\n| 200 | 100% |  |  |  |\n\n| 208 | 100% | 98% |  |  |\n\n| 216 | 100% |  |  |  |\n\n| 224 | 98% | 95% | 79% |  |\n\n| 232 | 100% |  |  |  |\n\n| 240 | 100% | 97% |  |  |\n\n| 248 | 100% |  |  |  |\n\n| 256 | 100% | 95% | 88% | 56% |\n\neval f1\n\neval f1 | step | 8 groups/step | 16 groups/step | 32 groups/step | 64 groups/step |\n\n|---|---|---|---|---|\n\n| 0 | 0.19 | 0.19 | 0.19 | 0.19 |\n\n| 64 | 0.20 | 0.18 | 0.24 | 0.17 |\n\n| 128 | 0.13 | 0.20 | 0.20 | 0.21 |\n\n| 192 | 0.12 | 0.11 | 0.20 | 0.29 |\n\n| 256 | 0.14 | 0.18 | 0.19 | 0.24 |\n\nAt a glance, the eval results are in line with expectations: at the end of training, 64 groups/step has the highest eval F1, and each smaller batch size monotonically decreases in final eval F1. Larger batches also give smoother reward and eval curves.\n\nThere’s an interesting contradiction in the charts, though. Although 64 groups/step ends with the highest eval score, it has the lowest train reward at the end of training. Also, 8 groups/step has the highest train reward, but the lowest F1 score.\n\nThis has to do with our reward setup and training dynamics — the reward on a rollout that curates nothing is -0.2, which far outweighs the average reward on most valid rollouts. So, the fastest way to raise train reward early on is to just curate any document. Because 8 groups/step takes more updates than any other run, the model starts to produce valid rollouts sooner. The middle chart above shows this: as batch size decreases, the model reaches >75% valid train rollouts faster. The problem from there is that noisy updates (due to small batch size) keep the policy’s curation quality from ever meaningfully improving. What we’re seeing here is actually a form of reward hacking. The next section will explore this more.\n\nSince 64 groups/step gives the most stable training and has the highest eval score, we will use it for our full training runs.  We could technically increase groups/step further, but for the sake of budget we will stop here.\n\nScaling up training\n\nFor our next set of training runs, we will move to the larger dataset of 1,024 queries and 237,533 chunks. This is 4x the data samples and 2x the corpus size of the dataset we used for ablations. We train for 1.5 epochs, which at 64 groups/step is 24 total updates.\n\nFull training configuration\n\n| model |\n\n| model | gpt-oss-20b |\n\n| adapter | LoRA, rank 32 |\n\n| optimisation |\n\n| learning rate | 1e-4 |\n\n| optimizer | Adam, β₁ 0.9, β₂ 0.95, ε 1e-8 |\n\n| weight decay | 0 |\n\n| KL penalty to reference | 0.0 |\n\n| updates per batch | 1 |\n\n| sampling temperature | 1.0 |\n\n| harness |\n\n| max turns | 40 |\n\n| max new tokens per turn | 2,048 |\n\n| context length limit | 30,720 tokens |\n\n| curated set cap | 30 |\n\nWe first look at training metrics over the full run. In the spirit of openness, every training rollout can be viewed in the explorer below.\n\ntrain reward\n\ntrain reward | step | train reward |\n\n|---|---|\n\n| 0 | -0.06 |\n\n| 1 | 0.02 |\n\n| 2 | 0.01 |\n\n| 3 | 0.03 |\n\n| 4 | 0.04 |\n\n| 5 | 0.10 |\n\n| 6 | 0.17 |\n\n| 7 | 0.15 |\n\n| 8 | 0.14 |\n\n| 9 | 0.15 |\n\n| 10 | 0.13 |\n\n| 11 | 0.17 |\n\n| 12 | 0.15 |\n\n| 13 | 0.16 |\n\n| 14 | 0.18 |\n\n| 15 | 0.21 |\n\n| 16 | 0.25 |\n\n| 17 | 0.17 |\n\n| 18 | 0.22 |\n\n| 19 | 0.23 |\n\n| 20 | 0.21 |\n\n| 21 | 0.27 |\n\n| 22 | 0.24 |\n\n| 23 | 0.21 |\n\nvalid rollouts\n\nvalid rollouts | step | valid rollouts |\n\n|---|---|\n\n| 0 | 30% |\n\n| 1 | 42% |\n\n| 2 | 47% |\n\n| 3 | 64% |\n\n| 4 | 79% |\n\n| 5 | 92% |\n\n| 6 | 96% |\n\n| 7 | 98% |\n\n| 8 | 98% |\n\n| 9 | 98% |\n\n| 10 | 98% |\n\n| 11 | 97% |\n\n| 12 | 94% |\n\n| 13 | 95% |\n\n| 14 | 94% |\n\n| 15 | 90% |\n\n| 16 | 92% |\n\n| 17 | 92% |\n\n| 18 | 95% |\n\n| 19 | 96% |\n\n| 20 | 96% |\n\n| 21 | 94% |\n\n| 22 | 94% |\n\n| 23 | 96% |\n\nentropy (nats)\n\nentropy (nats) | step | entropy (nats) |\n\n|---|---|\n\n| 0 | 1.1 |\n\n| 1 | 1.0 |\n\n| 2 | 1.0 |\n\n| 3 | 1.0 |\n\n| 4 | 1.0 |\n\n| 5 | 0.9 |\n\n| 6 | 0.8 |\n\n| 7 | 0.8 |\n\n| 8 | 0.8 |\n\n| 9 | 0.8 |\n\n| 10 | 0.8 |\n\n| 11 | 0.8 |\n\n| 12 | 0.8 |\n\n| 13 | 0.8 |\n\n| 14 | 0.9 |\n\n| 15 | 0.9 |\n\n| 16 | 0.9 |\n\n| 17 | 0.9 |\n\n| 18 | 0.9 |\n\n| 19 | 0.9 |\n\n| 20 | 0.9 |\n\n| 21 | 0.9 |\n\n| 22 | 0.9 |\n\n| 23 | 0.8 |\n\nTrain reward moves the way we’d expect it to here: up rapidly at first, and then slowing down as training matures. The training shape also mirrors what we see in ablations: the reward first teaches the model to curate. After that, reward moves around as the model starts to learn how to curate well instead of just how to curate. You can think of the -0.2 floor on no curations as jump-starting curation behavior.\n\nEntropy is another useful metric to watch during reinforcement learning.  Entropy is measured in nats here. One nat is 1 / ln 2, or about 1.44 bits.  It measures how uncertain the policy’s output distribution is over the course of a rollout. High entropy means at several points in the rollout, the model is considering several possible actions (next tokens). As we’ve discussed earlier, this is good for RL because it means the model will explore more and hopefully stumble into behaviors that get rewarded. Low entropy isn’t inherently bad, but a rapid drop during training is a sign the model is becoming overconfident or committing to a single set of behaviors. This is known as entropy collapse and is particularly bad for GRPO — a policy that always takes the same actions produces groups with  no variance, and groups with no variance produce no gradient.\n\nDuring healthy training, entropy typically decreases gradually as reward improves. That is exactly what we see here: entropy starts at 1.1 nats, gradually goes to ~0.8 a quarter of the way through training, and then stays there.\n\nLet’s see how our policy performs on held-out data now.\n\neval f1\n\neval f1 | step | eval f1 |\n\n|---|---|\n\n| 0 | 0.19 |\n\n| 2 | 0.22 |\n\n| 4 | 0.17 |\n\n| 6 | 0.19 |\n\n| 8 | 0.21 |\n\n| 10 | 0.17 |\n\n| 12 | 0.16 |\n\n| 14 | 0.24 |\n\n| 16 | 0.23 |\n\n| 18 | 0.29 |\n\n| 20 | 0.31 |\n\n| 22 | 0.28 |\n\n| 24 | 0.33 |\n\neval precision\n\neval precision | step | eval precision |\n\n|---|---|\n\n| 0 | 0.30 |\n\n| 2 | 0.38 |\n\n| 4 | 0.30 |\n\n| 6 | 0.29 |\n\n| 8 | 0.32 |\n\n| 10 | 0.27 |\n\n| 12 | 0.17 |\n\n| 14 | 0.35 |\n\n| 16 | 0.30 |\n\n| 18 | 0.36 |\n\n| 20 | 0.40 |\n\n| 22 | 0.36 |\n\n| 24 | 0.40 |\n\neval recall\n\neval recall | step | eval recall |\n\n|---|---|\n\n| 0 | 0.16 |\n\n| 2 | 0.17 |\n\n| 4 | 0.13 |\n\n| 6 | 0.16 |\n\n| 8 | 0.17 |\n\n| 10 | 0.14 |\n\n| 12 | 0.18 |\n\n| 14 | 0.21 |\n\n| 16 | 0.20 |\n\n| 18 | 0.27 |\n\n| 20 | 0.28 |\n\n| 22 | 0.26 |\n\n| 24 | 0.32 |\n\nProgress is uneven early on as the model learns our harness, and eval performance only starts to improve around halfway through training. Mean eval F1 rises from 0.193 at the start to 0.332 at update 24, driven almost entirely by recall, which doubles from 0.16 to 0.32. Precision goes up as well, moving from 0.31 to 0.40. This is a good result because it means we are getting an all around improvement. That is, we aren’t trading off precision for recall in any way.\n\nThe model performance doesn’t come at the cost of test-time compute, either. As the following charts show, the turn count actually falls slightly over training and output tokens go down, which means our final model is finding more relevant docs with less searching.\n\nturns per episode\n\nturns per episode | step | turns per episode |\n\n|---|---|\n\n| 0 | 18 |\n\n| 1 | 18 |\n\n| 2 | 18 |\n\n| 3 | 16 |\n\n| 4 | 15 |\n\n| 5 | 13 |\n\n| 6 | 12 |\n\n| 7 | 12 |\n\n| 8 | 12 |\n\n| 9 | 12 |\n\n| 10 | 12 |\n\n| 11 | 12 |\n\n| 12 | 14 |\n\n| 13 | 14 |\n\n| 14 | 16 |\n\n| 15 | 15 |\n\n| 16 | 16 |\n\n| 17 | 16 |\n\n| 18 | 15 |\n\n| 19 | 15 |\n\n| 20 | 14 |\n\n| 21 | 14 |\n\n| 22 | 16 |\n\n| 23 | 17 |\n\noutput tokens per episode\n\noutput tokens per episode | step | output tokens per episode |\n\n|---|---|\n\n| 0 | 3.998k |\n\n| 1 | 3.47k |\n\n| 2 | 3.05k |\n\n| 3 | 2.621k |\n\n| 4 | 2.202k |\n\n| 5 | 1.754k |\n\n| 6 | 1.408k |\n\n| 7 | 1.438k |\n\n| 8 | 1.414k |\n\n| 9 | 1.497k |\n\n| 10 | 1.698k |\n\n| 11 | 1.709k |\n\n| 12 | 1.972k |\n\n| 13 | 1.954k |\n\n| 14 | 2.328k |\n\n| 15 | 2.461k |\n\n| 16 | 2.527k |\n\n| 17 | 2.898k |\n\n| 18 | 2.693k |\n\n| 19 | 2.654k |\n\n| 20 | 2.673k |\n\n| 21 | 2.649k |\n\n| 22 | 2.825k |\n\n| 23 | 2.74k |\n\nSo, we have a full training run that works and pushes eval scores noticeably higher. This is a great result! In all likelihood, we haven’t saturated our training setup, either—train reward is still steadily going up around update 24. We’d probably continue to see eval gains by scaling up the dataset or running for another epoch…\n\nBut that’s a boring next step to explore here. There are many other fun ways to squeeze out performance from our dataset.\n\nAside: What happened with 8 groups/step?\n\nIn our ablations, we saw that the 8 groups/step run quickly learned to curate, but this failed to translate over meaningfully into eval improvements. We extended that run over the larger dataset to see if more training steps would help. The charts below show the results of training over 8 groups/step vs 64 groups/step over the same first 384 queries of the larger dataset.\n\ntrain reward\n\ntrain reward | step | 64 groups/step | 8 groups/step |\n\n|---|---|---|\n\n| 0 | -0.06 | 0.06 |\n\n| 8 |  | 0.03 |\n\n| 16 |  | 0.06 |\n\n| 24 |  | 0.03 |\n\n| 32 |  | 0.13 |\n\n| 40 |  | 0.07 |\n\n| 48 |  | 0.23 |\n\n| 56 |  | 0.16 |\n\n| 64 | 0.02 | 0.15 |\n\n| 72 |  | 0.21 |\n\n| 80 |  | 0.30 |\n\n| 88 |  | 0.18 |\n\n| 96 |  | 0.27 |\n\n| 104 |  | 0.12 |\n\n| 112 |  | 0.09 |\n\n| 120 |  | 0.20 |\n\n| 128 | 0.01 | 0.11 |\n\n| 136 |  | 0.12 |\n\n| 144 |  | 0.13 |\n\n| 152 |  | 0.06 |\n\n| 160 |  | 0.10 |\n\n| 168 |  | 0.15 |\n\n| 176 |  | 0.08 |\n\n| 184 |  | 0.17 |\n\n| 192 | 0.03 | 0.09 |\n\n| 200 |  | 0.14 |\n\n| 208 |  | 0.09 |\n\n| 216 |  | 0.13 |\n\n| 224 |  | 0.09 |\n\n| 232 |  | 0.09 |\n\n| 240 |  | 0.21 |\n\n| 248 |  | 0.30 |\n\n| 256 | 0.04 | 0.10 |\n\n| 264 |  | 0.18 |\n\n| 272 |  | 0.23 |\n\n| 280 |  | 0.03 |\n\n| 288 |  | 0.10 |\n\n| 296 |  | 0.14 |\n\n| 304 |  | 0.25 |\n\n| 312 |  | 0.19 |\n\n| 320 | 0.10 | 0.19 |\n\n| 328 |  | 0.25 |\n\n| 336 |  | 0.36 |\n\n| 344 |  | 0.13 |\n\n| 352 |  | 0.15 |\n\n| 360 |  | 0.08 |\n\n| 368 |  | 0.21 |\n\n| 376 |  | 0.16 |\n\n| 384 | 0.17 |  |\n\ncurated documents\n\ncurated documents | step | 64 groups/step | 8 groups/step |\n\n|---|---|---|\n\n| 0 | 1 | 1 |\n\n| 8 |  | 1 |\n\n| 16 |  | 1 |\n\n| 24 |  | 1 |\n\n| 32 |  | 1 |\n\n| 40 |  | 1 |\n\n| 48 |  | 2 |\n\n| 56 |  | 3 |\n\n| 64 | 1 | 2 |\n\n| 72 |  | 3 |\n\n| 80 |  | 4 |\n\n| 88 |  | 4 |\n\n| 96 |  | 5 |\n\n| 104 |  | 5 |\n\n| 112 |  | 6 |\n\n| 120 |  | 7 |\n\n| 128 | 1 | 6 |\n\n| 136 |  | 9 |\n\n| 144 |  | 8 |\n\n| 152 |  | 7 |\n\n| 160 |  | 7 |\n\n| 168 |  | 7 |\n\n| 176 |  | 7 |\n\n| 184 |  | 7 |\n\n| 192 | 1 | 9 |\n\n| 200 |  | 12 |\n\n| 208 |  | 9 |\n\n| 216 |  | 9 |\n\n| 224 |  | 11 |\n\n| 232 |  | 11 |\n\n| 240 |  | 9 |\n\n| 248 |  | 13 |\n\n| 256 | 1 | 18 |\n\n| 264 |  | 13 |\n\n| 272 |  | 12 |\n\n| 280 |  | 17 |\n\n| 288 |  | 16 |\n\n| 296 |  | 17 |\n\n| 304 |  | 21 |\n\n| 312 |  | 21 |\n\n| 320 | 2 | 24 |\n\n| 328 |  | 22 |\n\n| 336 |  | 23 |\n\n| 344 |  | 24 |\n\n| 352 |  | 24 |\n\n| 360 |  | 24 |\n\n| 368 |  | 25 |\n\n| 376 |  | 25 |\n\n| 384 | 2 |  |\n\nentropy (nats)\n\nentropy (nats) | step | 64 groups/step | 8 groups/step |\n\n|---|---|---|\n\n| 0 | 1.1 | 1.1 |\n\n| 8 |  | 1.1 |\n\n| 16 |  | 1.0 |\n\n| 24 |  | 0.9 |\n\n| 32 |  | 0.9 |\n\n| 40 |  | 0.9 |\n\n| 48 |  | 0.8 |\n\n| 56 |  | 0.7 |\n\n| 64 | 1.0 | 0.8 |\n\n| 72 |  | 0.7 |\n\n| 80 |  | 0.7 |\n\n| 88 |  | 0.7 |\n\n| 96 |  | 0.7 |\n\n| 104 |  | 0.7 |\n\n| 112 |  | 0.7 |\n\n| 120 |  | 0.7 |\n\n| 128 | 1.0 | 0.8 |\n\n| 136 |  | 0.7 |\n\n| 144 |  | 0.8 |\n\n| 152 |  | 0.7 |\n\n| 160 |  | 0.7 |\n\n| 168 |  | 0.7 |\n\n| 176 |  | 0.7 |\n\n| 184 |  | 0.8 |\n\n| 192 | 1.0 | 0.7 |\n\n| 200 |  | 0.8 |\n\n| 208 |  | 0.7 |\n\n| 216 |  | 0.7 |\n\n| 224 |  | 0.7 |\n\n| 232 |  | 0.6 |\n\n| 240 |  | 0.6 |\n\n| 248 |  | 0.7 |\n\n| 256 | 1.0 | 0.5 |\n\n| 264 |  | 0.6 |\n\n| 272 |  | 0.6 |\n\n| 280 |  | 0.5 |\n\n| 288 |  | 0.5 |\n\n| 296 |  | 0.5 |\n\n| 304 |  | 0.4 |\n\n| 312 |  | 0.5 |\n\n| 320 | 0.9 | 0.4 |\n\n| 328 |  | 0.4 |\n\n| 336 |  | 0.4 |\n\n| 344 |  | 0.4 |\n\n| 352 |  | 0.3 |\n\n| 360 |  | 0.3 |\n\n| 368 |  | 0.3 |\n\n| 376 |  | 0.3 |\n\n| 384 | 0.8 |  |\n\neval f1\n\neval f1 | step | 64 groups/step | 8 groups/step |\n\n|---|---|---|\n\n| 0 | 0.19 | 0.19 |\n\n| 128 | 0.25 | 0.12 |\n\n| 256 | 0.15 | 0.16 |\n\n| 384 | 0.17 | 0.16 |\n\neval precision\n\neval precision | step | 64 groups/step | 8 groups/step |\n\n|---|---|---|\n\n| 0 | 0.30 | 0.30 |\n\n| 128 | 0.38 | 0.11 |\n\n| 256 | 0.30 | 0.14 |\n\n| 384 | 0.29 | 0.13 |\n\neval recall\n\neval recall | step | 64 groups/step | 8 groups/step |\n\n|---|---|---|\n\n| 0 | 0.16 | 0.16 |\n\n| 128 | 0.17 | 0.14 |\n\n| 256 | 0.13 | 0.23 |\n\n| 384 | 0.16 | 0.29 |\n\nThe 8 groups/step run reaches higher training reward, but it gets there by learning to just curate everything — although eval recall rises from 0.16 to 0.29, precision drops from 0.31 to 0.13. Entropy also falls from 1.1 to 0.34 nats over the course of training. This is a clear sign of entropy collapse. It’s likely that learning will outright stop if we were to continue training.\n\nPart of this is our reward choice. We chose β = 4 to push the model towards recall over precision, and F4 weights recall 16x. Small-batch updates are also noisier, so once the policy stumbles onto the “curate every document you see” strategy, there is little pushing the model back to something more reasonable.\n\nOne way to mitigate the reward bias here would be to anneal β over the course of training, as is done in Chroma Context-1.\n\nReward shaping\n\nEvery run so far has used the plain F4 score as the reward. It worked, in the sense that eval scores went up. But reading the rollouts directly and exploring a different set of training metrics shows two other ways we can improve the policy.\n\ntrajectory vs curated recall\n\ntrajectory vs curated recall | step | trajectory recall | curated recall |\n\n|---|---|---|\n\n| 0 | 0.46 | 0.12 |\n\n| 1 | 0.51 | 0.16 |\n\n| 2 | 0.40 | 0.13 |\n\n| 3 | 0.34 | 0.11 |\n\n| 4 | 0.29 | 0.09 |\n\n| 5 | 0.32 | 0.14 |\n\n| 6 | 0.39 | 0.22 |\n\n| 7 | 0.38 | 0.24 |\n\n| 8 | 0.36 | 0.23 |\n\n| 9 | 0.33 | 0.21 |\n\n| 10 | 0.26 | 0.19 |\n\n| 11 | 0.33 | 0.24 |\n\n| 12 | 0.29 | 0.21 |\n\n| 13 | 0.30 | 0.21 |\n\n| 14 | 0.27 | 0.18 |\n\n| 15 | 0.38 | 0.26 |\n\n| 16 | 0.40 | 0.28 |\n\n| 17 | 0.30 | 0.20 |\n\n| 18 | 0.39 | 0.28 |\n\n| 19 | 0.37 | 0.28 |\n\n| 20 | 0.36 | 0.27 |\n\n| 21 | 0.41 | 0.32 |\n\n| 22 | 0.39 | 0.29 |\n\n| 23 | 0.36 | 0.27 |\n\nTrajectory recall measures how many relevant chunks the model encountered while searching.      tool-call format drift\n\ntool-call format drift | step | of calls | of episodes |\n\n|---|---|---|\n\n| 0 | 27.3% | 37.9% |\n\n| 1 | 13.7% | 29.3% |\n\n| 2 | 15% | 26.4% |\n\n| 3 | 9.1% | 16.2% |\n\n| 4 | 11.5% | 18.9% |\n\n| 5 | 12.4% | 18.8% |\n\n| 6 | 14.8% | 17.8% |\n\n| 7 | 25% | 32.2% |\n\n| 8 | 21.2% | 32.2% |\n\n| 9 | 19.7% | 32.4% |\n\n| 10 | 14.6% | 33.6% |\n\n| 11 | 26.8% | 41.6% |\n\n| 12 | 28.3% | 55.9% |\n\n| 13 | 47.2% | 67.2% |\n\n| 14 | 49.6% | 74.2% |\n\n| 15 | 62.6% | 79.9% |\n\n| 16 | 53.8% | 78.3% |\n\n| 17 | 58.5% | 89.3% |\n\n| 18 | 67.3% | 90.2% |\n\n| 19 | 67.9% | 92.6% |\n\n| 20 | 63% | 92.4% |\n\n| 21 | 62.6% | 91% |\n\n| 22 | 67.7% | 96.7% |\n\n| 23 | 75% | 98.8% |\n\nMeasures % of tool calls or episodes with incorrect harmony format.\n\n- Curation recall is bounded by trajectory recall. The model can only curate what it sees as it searches, and in the latter half of training, curation recall consistently sits 0.1 lower than trajectory recall. Trajectory recall also drifts down over training, indicating that the model is getting better at picking relevant docs from searches, but slightly worse at actually searching.\n\n- Tool-call format adherence degrades over training. As training progresses, the share of incorrectly formatted tool calls nearly triples. Luckily, our parser is lenient, so this doesn’t seem to impact training performance too much.\n\nThese observations suggest two ways we can improve the model further: reward the model for encountering more evidence during search, and punish it for bad tool calls.\n\nRewarding discovery\n\nThe trajectory recall chart tells us that the model either isn’t searching well or isn’t searching enough. There are two ways we can influence this behavior.\n\nFirst, we can reward the behavior directly: a bonus per search or grep call, or for more varied queries, and hope that together with the F4 reward, this leads to better coverage. However, a bonus just for searching is easy to hack. The model could just spam irrelevant calls and collect the bonus without learning any performance-improving behavior. We’d need to add a turn penalty to hold it back. This adds a second knob to tune and complicates the reward.\n\nAnother route is to add an explicit trajectory recall term to our reward. This is a better route because we’re rewarding the outcome we want rather than the activity it takes to get there. As our charts show, trajectory recall is directionally aligned with curated recall, and it’s also harder to game — extra searches don’t earn anything unless they are helpful. This is the route we’ll experiment with.\n\nConcretely, we add a term for the fraction of the query’s facts that have a supporting chunk anywhere in the trajectory, curated or not:\n\nr(C)={−0.2F4+0.2⋅Rtrajif ∣C∣=0otherwise\n\nThe explorer below shows the impact of different reward shapes on a group. Seven of the eight rollouts score zero F4 and are indistinguishable under the older reward, but separable under either bonus.\n\nApart from the reward, every setting in this run matches the full run above.\n\neval f1\n\neval f1 | step | w/o trajectory reward | w trajectory reward |\n\n|---|---|---|\n\n| 0 | 0.19 | 0.19 |\n\n| 2 | 0.22 | 0.21 |\n\n| 4 | 0.17 | 0.21 |\n\n| 6 | 0.19 | 0.18 |\n\n| 8 | 0.21 | 0.18 |\n\n| 10 | 0.17 | 0.27 |\n\n| 12 | 0.16 | 0.27 |\n\n| 14 | 0.24 | 0.25 |\n\n| 16 | 0.23 | 0.28 |\n\n| 18 | 0.29 | 0.27 |\n\n| 20 | 0.31 | 0.32 |\n\n| 22 | 0.28 | 0.36 |\n\n| 24 | 0.33 | 0.33 |\n\neval curated recall\n\neval curated recall | step | w/o trajectory reward | w trajectory reward |\n\n|---|---|---|\n\n| 0 | 0.16 | 0.16 |\n\n| 2 | 0.17 | 0.15 |\n\n| 4 | 0.13 | 0.16 |\n\n| 6 | 0.16 | 0.14 |\n\n| 8 | 0.17 | 0.14 |\n\n| 10 | 0.14 | 0.22 |\n\n| 12 | 0.18 | 0.23 |\n\n| 14 | 0.21 | 0.20 |\n\n| 16 | 0.20 | 0.24 |\n\n| 18 | 0.27 | 0.24 |\n\n| 20 | 0.28 | 0.29 |\n\n| 22 | 0.26 | 0.34 |\n\n| 24 | 0.32 | 0.30 |\n\neval trajectory recall\n\neval trajectory recall | step | w/o trajectory reward | w trajectory reward |\n\n|---|---|---|\n\n| 0 | 0.43 | 0.43 |\n\n| 2 | 0.35 | 0.37 |\n\n| 4 | 0.30 | 0.30 |\n\n| 6 | 0.29 | 0.26 |\n\n| 8 | 0.28 | 0.26 |\n\n| 10 | 0.22 | 0.33 |\n\n| 12 | 0.22 | 0.33 |\n\n| 14 | 0.28 | 0.36 |\n\n| 16 | 0.30 | 0.38 |\n\n| 18 | 0.36 | 0.47 |\n\n| 20 | 0.37 | 0.48 |\n\n| 22 | 0.33 | 0.53 |\n\n| 24 | 0.44 | 0.49 |\n\nEval is run 2x, and plotted points are the averages. Reward plots omitted because they are noncomparable.\n\nTrajectory recall is clearly better here: after a dip around step 8, the shaped run recovers sooner and stays above the old, plain-F4 run. Eval F1 and curated recall also trend above the old run and even peak higher. Unfortunately, this doesn’t translate directly to a better final score under the same training budget.\n\nHarmony, the format gpt-oss-20b uses, has a specific header for tool calls. At the start of training, the model gets it wrong on about a quarter of calls.\n\nBecause we use a lenient parse that accepts these slightly off-form calls, the model still gets rewarded whenever the episode they’re in happens to get a good score. This habit ends up compounding. By the final step of training, 99% of episodes contain at least one bad tool call. The model isn’t being told, in any way that reaches the gradient, that those calls are wrong, so it ends up learning that those calls are acceptable.\n\nThe following experiment fixes that directly. Every setting is the same as our full, plain F4 run, except for a slight change to the reward: subtract 0.1 if any tool call in the episode has an off-form header. If two rollouts in a group reach the same F4 but one uses a bad header, the clean one now gets more advantage.\n\nf1\n\nf1 | step | w/o format penalty | w format penalty |\n\n|---|---|---|\n\n| 0 | 0.19 | 0.19 |\n\n| 2 | 0.22 | 0.20 |\n\n| 4 | 0.17 | 0.27 |\n\n| 6 | 0.19 | 0.21 |\n\n| 8 | 0.21 | 0.24 |\n\n| 10 | 0.17 | 0.20 |\n\n| 12 | 0.16 | 0.25 |\n\n| 14 | 0.24 | 0.27 |\n\n| 16 | 0.23 | 0.34 |\n\n| 18 | 0.29 | 0.33 |\n\n| 20 | 0.31 | 0.39 |\n\n| 22 | 0.28 | 0.36 |\n\n| 24 | 0.33 | 0.36 |\n\nrecall\n\nrecall | step | w/o format penalty | w format penalty |\n\n|---|---|---|\n\n| 0 | 0.16 | 0.16 |\n\n| 2 | 0.17 | 0.16 |\n\n| 4 | 0.13 | 0.20 |\n\n| 6 | 0.16 | 0.17 |\n\n| 8 | 0.17 | 0.20 |\n\n| 10 | 0.14 | 0.20 |\n\n| 12 | 0.18 | 0.26 |\n\n| 14 | 0.21 | 0.26 |\n\n| 16 | 0.20 | 0.36 |\n\n| 18 | 0.27 | 0.35 |\n\n| 20 | 0.28 | 0.43 |\n\n| 22 | 0.26 | 0.42 |\n\n| 24 | 0.32 | 0.43 |\n\noff-form tool headers\n\noff-form tool headers | step | w/o format penalty (episodes) | w format penalty (episodes) | w/o format penalty (calls) | w format penalty (calls) |\n\n|---|---|---|---|---|\n\n| 0 | 37.9% | 39.8% | 27.3% | 27.8% |\n\n| 1 | 29.3% | 23.4% | 13.7% | 10.3% |\n\n| 2 | 26.4% | 21.7% | 15% | 6.1% |\n\n| 3 | 16.2% | 9.2% | 9.1% | 2.7% |\n\n| 4 | 18.9% | 6.8% | 11.5% | 2.3% |\n\n| 5 | 18.8% | 3.9% | 12.4% | 0.7% |\n\n| 6 | 17.8% | 3.5% | 14.8% | 1% |\n\n| 7 | 32.2% | 7% | 25% | 2.1% |\n\n| 8 | 32.2% | 4.9% | 21.2% | 0.5% |\n\n| 9 | 32.4% | 9.6% | 19.7% | 2.8% |\n\n| 10 | 33.6% | 12.7% | 14.6% | 3.6% |\n\n| 11 | 41.6% | 8.6% | 26.8% | 1.1% |\n\n| 12 | 55.9% | 8% | 28.3% | 0.8% |\n\n| 13 | 67.2% | 7.2% | 47.2% | 0.6% |\n\n| 14 | 74.2% | 8.2% | 49.6% | 0.5% |\n\n| 15 | 79.9% | 8.6% | 62.6% | 0.4% |\n\n| 16 | 78.3% | 5.9% | 53.8% | 0.4% |\n\n| 17 | 89.3% | 5.1% | 58.5% | 0.2% |\n\n| 18 | 90.2% | 4.9% | 67.3% | 0.2% |\n\n| 19 | 92.6% | 3.1% | 67.9% | 0.1% |\n\n| 20 | 92.4% | 3.1% | 63% | 0.1% |\n\n| 21 | 91% | 4.3% | 62.6% | 0.2% |\n\n| 22 | 96.7% | 4.7% | 67.7% | 0.2% |\n\n| 23 | 98.8% | 8.2% | 75% | 0.4% |\n\nturns per episode\n\nturns per episode | step | w/o format penalty | w format penalty |\n\n|---|---|---|\n\n| 0 | 18 | 18 |\n\n| 1 | 18 | 17 |\n\n| 2 | 18 | 17 |\n\n| 3 | 16 | 13 |\n\n| 4 | 15 | 12 |\n\n| 5 | 13 | 12 |\n\n| 6 | 12 | 12 |\n\n| 7 | 12 | 13 |\n\n| 8 | 12 | 15 |\n\n| 9 | 12 | 14 |\n\n| 10 | 12 | 15 |\n\n| 11 | 12 | 16 |\n\n| 12 | 14 | 20 |\n\n| 13 | 14 | 20 |\n\n| 14 | 16 | 21 |\n\n| 15 | 15 | 23 |\n\n| 16 | 16 | 23 |\n\n| 17 | 16 | 25 |\n\n| 18 | 15 | 26 |\n\n| 19 | 15 | 27 |\n\n| 20 | 14 | 27 |\n\n| 21 | 14 | 28 |\n\n| 22 | 16 | 30 |\n\n| 23 | 17 | 30 |\n\nIn the tool headers chart, solid lines show the share of episodes with off-form calls; dashed lines show the share of individual calls.\n\nThe extra penalty works, fast. Off-form tool calls fall below 1% within a few steps and stay there for the rest of training. Eval metrics and recall also improve against the baseline. Held-out F1 ends at 0.36 vs 0.33, and recall at 0.43 vs 0.32.\n\nThis result surprised me. I expected the penalty to fix tool call formatting, but not necessarily retrieval quality. Yet the model starts taking longer trajectories (does more searches) and achieves higher recall. I don’t have a concrete explanation for this. One possibility is that the penalty changes which search behaviors get reinforced, and well-formed tool calls are happen to be correlated with better search behaviors. Another is that the model just has more coherent long-context understanding when all of its preceding tool calls are well-formed.\n\nEither hypothesis would require more training runs to test, but I’ll stop here for the sake of my training budget. What we should take away from this is that a seemingly narrow reward change can sometimes lead to big shifts in the trained policy. We need to be careful and intentional when working on reward functions.\n\nSome closing thoughts\n\nWe started with a 20B model that struggled to properly execute over our harness, then ran ablations and experiments to drive its eval F1 score from 0.19 up to a peak of 0.39.\n\nThere’s still a lot more we could do here to improve performance. For one, we only trained on a fraction of the total SEC Harness-1 dataset. We could expand the training set or test more learning-rate and batch-size combinations. Tuning one hyperparameter at a time was the right call under our specific budget, but we could have carried over a few configurations to the full run.\n\n## A few other, more interesting experiment directions\n\n- Curriculum learning During reinforcement learning, we typically want to give the model tasks that are on the frontier of the model’s capabilities. One way to do this is by getting an empirical pass rate of each task before training. We can use this pass rate to bucket tasks into different difficulty tiers and then sample with different weights over the course of training.\n\n- SFT Warmup In our scaled up training runs, we end up devoting about a third of our data budget just to teaching the model how to use our harness. We can speed this up by adding an initial warmup stage with SFT over some easier tasks.\n\n- Harness Engineering We generally don’t want to be too prescriptive in the system prompt because we want the model to discover behaviors during training. However, that the model initially does so poorly is a sign that we can improve either our system prompt or harness signals a bit.\n\n- Other reward shapes So far, we’ve kept our reward functions deliberately simple. There’s a lot more we can do here. For example: a turn penalty to encourage shorter rollouts, rewards for tool diversity, etc.","body_html":"<p>This post is a hands-on introduction to reinforcement learning through training a search agent with group-relative policy optimization (GRPO). Search is a fun place to learn RL because it has so many levers, and each of them visibly changes how the model searches. It is also a domain where, in my experience, a well-designed reward function can influence how a model searches more effectively than system prompt changes or harness engineering.</p>\n<p>Our approach is based loosely on the Harness-1 paper from Jiang et al. We use their training dataset and a much-simplified version of their search harness.</p>\n<p>Training is done using the Tinker API from Thinking Machines. Many thanks to the team for the training credits to support this work.</p>\n<p>The dataset</p>\n<p>We use the RL split of the Harness-1 dataset for all experiments. The dataset contains 3,453 queries over 37,113 SEC filings split into 2,115,106 chunks.</p>\n<p>Each query includes a preamble followed by a question. The agent must both retrieve evidence for facts in the preamble and the answer to the question. In the canonical version of this task, retrieving the answer is more important than evidence for facts. However, for simplicity we will treat both the same.</p>\n<p>To keep initial ablations inexpensive, we start with just 256 training queries and 32 evaluation queries over a reduced corpus of 124,395 chunks. For our scaled runs, we expand to 1,024 training queries over 237,533 chunks.</p>\n<p>The harness</p>\n<p>Our agent operates as a search subagent. Given a query, it is meant to retrieve all relevant docs. We keep the harness deliberately simple, only providing lexical search tools. The agent’s output is a curated set it builds up as it goes rather than a single submission at the end. We chose this because it works well for a fixed turn budget: since the set exists from the start, an episode that runs out of turns will will return something gradable.</p>\n<p>| Tool | What it does |</p>\n<p>|---|---|</p>\n<p>| bm25_search | Keyword search over every chunk. Returns an id, title, score and snippet per hit, plus a running count of chunks seen and curated so far. |</p>\n<p>| grep_corpus | A short Python regex scan across the corpus. Returns an id, title, and snippet per hit. |</p>\n<p>| read_document | The title and text of one chunk. Only ids the agent has already seen in a result list can be read. |</p>\n<p>| curate | Add chunk ids to the result set. |</p>\n<p>| drop_curated | Remove ids from the result set. |</p>\n<p>| finish | End the episode and return the set. |</p>\n<p>The system prompt is similarly plain. We name the tools and give a general strategy for search. We don’t try to encourage any specific search characteristics (e.g. high recall or high precision) so that the agent can discover the right behaviors during training.</p>\n<p>Retrieval agent system prompt</p>\n<p>You are a retrieval subagent in a multi-agent system. Your specific role is to identify and retrieve the most relevant documents from a large corpus to help another agent answer questions. You do NOT answer questions yourself - you only find and retrieve relevant documents.</p>\n<p>The user message contains the query you need to find documents for.</p>\n<p><strong>Available Tools:</strong></p>\n<ul><li>bm25_search: Keyword search over the corpus</li><li>grep_corpus: Text pattern matching with a short Python regex</li><li>read_document: Read a specific document that looks promising but incomplete</li><li>curate: Add relevant documents to your curated result set</li><li>drop_curated: Remove curated documents that turned out to be irrelevant</li><li>finish: End the search and return your curated set</li></ul>\n<p><strong>Your Process:</strong></p>\n<ul><li>Break down the query into its key concepts and information needs (list each one explicitly)</li><li>For each key concept, develop a specific search strategy that targets that concept</li><li>Consider what types of documents and evidence would be most helpful for answering this query</li><li>Plan several distinct, non-overlapping search strategies that approach the question from different angles</li><li>Then execute your searches using multiple parallel tool calls.</li></ul>\n<p><strong>Your Thinking:</strong></p>\n<h2 id=\"after-each-round-of-searches-consider-the-following\">After each round of searches, consider the following</h2>\n<ul><li><strong>What do I know?</strong>: List the key topics, themes, or aspects of the question that your curated documents address. What specific information do you have?</li><li><strong>What should I search for next?</strong>: Systematically consider what search approaches, keywords, or document types you haven&#39;t yet tried that might yield valuable information.</li><li><strong>What should I curate or drop?</strong>: Curate documents as soon as you have verified they are relevant; drop curated documents that later look redundant or off-topic.</li><li><strong>Do I have enough information?</strong>: Given the question&#39;s complexity and requirements, do you have sufficient information to help answer it, or are there critical gaps?</li><li>Decide if additional searches are needed (and if so, ensure they use genuinely different approaches and do not duplicate prior searches)</li><li>Avoid getting stuck on a single search strategy - if one approach isn&#39;t yielding results, backtrack and try different approaches</li></ul>\n<p><strong>Tactics to Consider:</strong></p>\n<ul><li>When queries fail, try different approaches or keywords to improve the results</li><li>Avoid duplicate or redundant searches</li><li>Execute multiple tool calls in parallel when possible</li><li>Focus on gathering as much relevant information as possible; it is useful to get multiple perspectives on the same topic to confirm the information you have found is correct</li><li>Follow explicit textual evidence rather than speculation</li></ul>\n<p><strong>Output (IMPORTANT):</strong></p>\n<ul><li>YOU MUST use the curate tool. This is the ONLY way to return documents.</li><li>Every response must contain at least one tool call. Never reply with plain text, and never end a response while still planning a search - emit that search as a tool call instead.</li><li>As soon as you verify a document is relevant, call curate with its ID.</li><li>When your curated set covers the query&#39;s information needs, call finish. Your curated set is scored even if you run out of turns, but finishing cleanly is always better than timing out.</li></ul>\n<p>We use gpt-oss-20b for all training runs. We chose this model because it is cheap to sample and train, allowing us to do many ablations.</p>\n<p>Quick GRPO recap</p>\n<h2 id=\"we-train-with-dr-grpo-a-variant-of-grpo-the-loop-is-as-follows\">We train with Dr. GRPO, a variant of GRPO. The loop is as follows</h2>\n<ul><li>Take a query and run the agent on it N times using the current policy (in this setting, essentially the model’s weights). Each run is a rollout, and the N tries together form a group.</li><li>Score each rollout. This score is its reward.</li><li>Subtract the group’s mean reward from each rollout’s reward. The resulting difference is that rollout’s advantage.  In vanilla GRPO, we would also divide each advantage by the group’s standard deviation. Dr. GRPO drops this term to remove a bias: a query where the eight rollouts nearly agree has a small deviation. Dividing by standard deviation scales up the advantages, so the update can end up dominated by queries the policy is already consistent on.</li><li>Update the policy (model weights) proportionally to the rollout’s advantage, so that the tokens in higher-advantage rollouts become more likely and tokens in lower-advantage rollouts become less likely.</li></ul>\n<p>We would also typically use ratio clipping during the policy update to help stabilize training. However, to simplify training for the sake of pedagogy, we will train in a fully synchronous setting with zero training-inference mismatch (which Tinker provides by default) and only do one update per batch.</p>\n<p>We won’t go deeper into the math behind here, but there are many other great resources out there for understanding GRPO.</p>\n<p>The reward</p>\n<p>Each task here consists of a query, a corpus, a set of facts, and a set of supporting documents for each fact. The agent has to curate a set of documents it determines to be relevant to those facts, and success is evaluated on how well this curated set covers the target facts.</p>\n<p>A reward function is essentially a kind of evaluation scheme for a rollout. A few metrics common to search retrieval work naturally here.</p>\n<p>Recall measures the fraction of relevant documents our agent has curated. Since our search problem is fact-based, this will be the fraction of relevant facts the curated chunks cover. Precision measures the fraction of curated documents that are actually relevant.</p>\n<p>We can’t use just recall or precision alone in our reward function. High recall is desirable in our problem space, but can be trivially maximized by curating every document in the corpus. This is useless for an end user! Why use a search agent at all? High precision can be achieved just as trivially by curating only a single correct chunk per query, which is useless in the other direction.</p>\n<p>We need a metric that punishes both forms of degenerate behavior. F1, the harmonic mean of precision and recall, is the usual choice here:</p>\n<p>F1=P+R2PR</p>\n<p>This metric is only high when both precision and recall are high. This is a totally reasonable metric to choose for our reward function. But thinking about our problem space again… it’s likely that a missing document here costs far more to an end user than an extra, useless one. So, we might want to bias the reward towards recall. With this in mind, we will use F-beta as our reward function. This is a weighted version of F1 with a coefficient for setting how much recall matters more than precision:</p>\n<p>Fβ=β2P+R(1+β2)PR</p>\n<p>We set β = 4 for our reward function, which weighs recall sixteen times more than precision. I chose this number based on my prior experience and notes from the Harness-1 paper rather than a formal ablation. Under this metric, which we call F4, a set with recall 1.0 and precision 0.25 scores 0.85; under F1 the same set would score 0.40.</p>\n<p>One more thing we add to our reward function is a punishment for   not curating anything—an episode that curates nothing at all gets a flat -0.2. This helps with differentiating between rollouts where the model never gets to curate anything and ones where it only curates incorrect documents.</p>\n<p>Putting it together, the reward for a rollout with curated set C is</p>\n<p>r(C)={−0.2F4if ∣C∣=0otherwise</p>\n<p>Initial explorations</p>\n<p>The core idea behind reinforcement learning is to let models explore and reward them for behaviors that lead to good results. This means that RL works specifically by reinforcing behaviors the model already produces. If the model either always reaches for the same actions or always fails, then there is nothing for it to learn from.</p>\n<p>This is especially true in GRPO, where learning depends on variation in rewards within each group. If all rollouts receive the same reward, their advantages are zero and the group provides no gradient.</p>\n<p>This is why it can be helpful to gauge the RL-ability of a task before training. Best-of-N evals are a good proxy for this. Just sample N runs over the same task and report the highest score. A good candidate for RL is a model that has noticeably higher best-of-N scores than best-of-1 — this gap is the headroom RL can close by just sharpening the existing policy.</p>\n<p>Below, we sample gpt-oss-20b four times on each of the eval queries. All rollouts are viewable.</p>\n<p>The gap between best-of-4 and single trial F1 is large: best-of-4 F1 is 0.323 vs 0.195 for a single try. The F1 score is also precision-dominated. The model is likely being over-conservative in its curation.</p>\n<p>Close inspection of the rollout supports this. On seven of the 32 queries, none of the four tries ever calls curate. The model stalls, exhausts all its turns, or runs out of context before curating a single document. Failures like these signal more of a harness problem than a search problem, and that is the kind of behavior reinforcement learning is good at fixing.</p>\n<p>Learning rate sweep</p>\n<p>Learning rate is a critical hyperparameter to get right. Just like in supervised fine-tuning, a bad learning rate can destabilize training. In reinforcement learning, a learning rate too high can cause the policy to collapse; one that is too small means the model never learns anything during training. And even when a run looks fine, I’ve found that further tuning the learning rate can often lead to incrementally better performance.</p>\n<p>That said, if you’re an experienced practitioner and know directionally what the right hyperparameters are, you can skip this step for initial experiments. My opinion is that once training works stably, tuning the learning rate only starts to matter when you’re trying to squeeze out extra performance from your data.</p>\n<p>All runs in this section and the next use the 256-query subset of our larger dataset, with 32 groups of 8 rollouts per step. Here are the results of a sweep over five different learning rates, each run for one epoch:</p>\n<p>train reward</p>\n<p>train reward | step | 5e-6 | 1e-5 | 5e-5 | 1e-4 | 5e-4 |</p>\n<p>|---|---|---|---|---|---|</p>\n<p>| 0 | -0.02 | -0.02 | -0.02 | -0.02 | -0.02 |</p>\n<p>| 1 | -0.03 | -0.03 | -0.02 | -0.02 | 0.01 |</p>\n<p>| 2 | -0.04 | -0.04 | -0.04 | -0.03 | -0.03 |</p>\n<p>| 3 | -0.02 | -0.03 | -0.02 | 0.00 | 0.05 |</p>\n<p>| 4 | 0.01 | -0.01 | 0.02 | 0.05 | 0.12 |</p>\n<p>| 5 | -0.01 | 0.02 | 0.03 | 0.09 | 0.16 |</p>\n<p>| 6 | -0.01 | 0.01 | 0.04 | 0.12 | 0.12 |</p>\n<p>| 7 | -0.04 | -0.02 | 0.02 | 0.15 | 0.01 |</p>\n<p>| 8 | -0.03 | -0.03 | 0.06 | 0.14 | 0.08 |</p>\n<p>valid train rollouts</p>\n<p>valid train rollouts | step | 5e-6 | 1e-5 | 5e-5 | 1e-4 | 5e-4 |</p>\n<p>|---|---|---|---|---|---|</p>\n<p>| 0 | 43% | 43% | 43% | 43% | 43% |</p>\n<p>| 1 | 44% | 39% | 42% | 44% | 46% |</p>\n<p>| 2 | 36% | 34% | 34% | 42% | 53% |</p>\n<p>| 3 | 40% | 39% | 39% | 45% | 73% |</p>\n<p>| 4 | 48% | 45% | 53% | 54% | 91% |</p>\n<p>| 5 | 45% | 47% | 51% | 66% | 100% |</p>\n<p>| 6 | 40% | 43% | 52% | 63% | 100% |</p>\n<p>| 7 | 37% | 41% | 55% | 79% | 75% |</p>\n<p>| 8 | 40% | 38% | 63% | 88% | 100% |</p>\n<p>eval f1</p>\n<p>eval f1 | step | 5e-6 | 1e-5 | 5e-5 | 1e-4 | 5e-4 |</p>\n<p>|---|---|---|---|---|---|</p>\n<p>| 0 | 0.19 | 0.19 | 0.19 | 0.19 | 0.19 |</p>\n<p>| 4 | 0.18 | 0.21 | 0.22 | 0.20 | 0.12 |</p>\n<p>| 8 | 0.16 | 0.14 | 0.16 | 0.19 | 0.03 |</p>\n<p>Training results over one epoch with different learning rates. The valid rollouts chart shows the percentage of training rollouts that curate at least one doc.</p>\n<p>We can see in here the effects of both a too large and a too small learning rate:</p>\n<ul><li>1e-5 and5e-6 barely move train reward. We clearly need a larger gradient signal to meaningfully update the policy.</li><li>5e-4 improves reward rapidly at first and then quickly destabilizes.</li><li>1e-4 and5e-5 train steadily, with1e-4 resulting in a higher final train reward and a higher eval F1.</li></ul>\n<p>Since lr=1e-4 performs the best in the sweep, we’ll use it for the rest of our experiments.</p>\n<p>If you want a better idea of what the training collapse for lr=5e-4 looks like, expand the explorer below to see its eval rollouts at steps 0 and 8.</p>\n<p>Batch size ablation</p>\n<p>Conventional wisdom when training with GRPO is to use large batch sizes, where batch size is groups per step x group size. To study the effects of batch size on training here, we fix group size at 8 and ablate over groups per step.</p>\n<p>The charts below are plotted against prompts seen rather than total updates, since all runs have a different total number of steps.</p>\n<p>train reward</p>\n<p>train reward | step | 8 groups/step | 16 groups/step | 32 groups/step | 64 groups/step |</p>\n<p>|---|---|---|---|---|</p>\n<p>| 0 | -0.03 | -0.03 | -0.03 | -0.03 |</p>\n<p>| 8 | -0.01 |  |  |  |</p>\n<p>| 16 | 0.04 | 0.02 |  |  |</p>\n<p>| 24 | 0.04 |  |  |  |</p>\n<p>| 32 | -0.06 | -0.07 | -0.01 |  |</p>\n<p>| 40 | -0.06 |  |  |  |</p>\n<p>| 48 | -0.00 | -0.06 |  |  |</p>\n<p>| 56 | 0.04 |  |  |  |</p>\n<p>| 64 | 0.25 | 0.04 | -0.04 | -0.05 |</p>\n<p>| 72 | 0.03 |  |  |  |</p>\n<p>| 80 | 0.06 | -0.01 |  |  |</p>\n<p>| 88 | 0.11 |  |  |  |</p>\n<p>| 96 | 0.14 | 0.09 | -0.01 |  |</p>\n<p>| 104 | 0.03 |  |  |  |</p>\n<p>| 112 | 0.07 | 0.04 |  |  |</p>\n<p>| 120 | 0.13 |  |  |  |</p>\n<p>| 128 | 0.21 | 0.20 | 0.01 | -0.01 |</p>\n<p>| 136 | 0.24 |  |  |  |</p>\n<p>| 144 | 0.16 | 0.21 |  |  |</p>\n<p>| 152 | 0.17 |  |  |  |</p>\n<p>| 160 | 0.02 | 0.10 | 0.10 |  |</p>\n<p>| 168 | 0.10 |  |  |  |</p>\n<p>| 176 | 0.13 | 0.14 |  |  |</p>\n<p>| 184 | 0.19 |  |  |  |</p>\n<p>| 192 | 0.09 | 0.14 | 0.08 | 0.07 |</p>\n<p>| 200 | 0.16 |  |  |  |</p>\n<p>| 208 | 0.14 | 0.16 |  |  |</p>\n<p>| 216 | 0.07 |  |  |  |</p>\n<p>| 224 | 0.08 | 0.11 | 0.12 |  |</p>\n<p>| 232 | 0.13 |  |  |  |</p>\n<p>| 240 | 0.06 | 0.07 |  |  |</p>\n<p>| 248 | 0.05 |  |  |  |</p>\n<p>| 256 | 0.17 | 0.12 | 0.16 | 0.05 |</p>\n<p>valid train rollouts</p>\n<p>valid train rollouts | step | 8 groups/step | 16 groups/step | 32 groups/step | 64 groups/step |</p>\n<p>|---|---|---|---|---|</p>\n<p>| 0 | 41% | 41% | 41% | 41% |</p>\n<p>| 8 | 36% |  |  |  |</p>\n<p>| 16 | 55% | 45% |  |  |</p>\n<p>| 24 | 60% |  |  |  |</p>\n<p>| 32 | 39% | 43% | 44% |  |</p>\n<p>| 40 | 66% |  |  |  |</p>\n<p>| 48 | 72% | 48% |  |  |</p>\n<p>| 56 | 85% |  |  |  |</p>\n<p>| 64 | 98% | 61% | 42% | 40% |</p>\n<p>| 72 | 92% |  |  |  |</p>\n<p>| 80 | 97% | 62% |  |  |</p>\n<p>| 88 | 100% |  |  |  |</p>\n<p>| 96 | 98% | 72% | 45% |  |</p>\n<p>| 104 | 100% |  |  |  |</p>\n<p>| 112 | 100% | 75% |  |  |</p>\n<p>| 120 | 100% |  |  |  |</p>\n<p>| 128 | 100% | 90% | 54% | 40% |</p>\n<p>| 136 | 100% |  |  |  |</p>\n<p>| 144 | 100% | 94% |  |  |</p>\n<p>| 152 | 100% |  |  |  |</p>\n<p>| 160 | 100% | 92% | 66% |  |</p>\n<p>| 168 | 100% |  |  |  |</p>\n<p>| 176 | 100% | 95% |  |  |</p>\n<p>| 184 | 100% |  |  |  |</p>\n<p>| 192 | 100% | 94% | 63% | 51% |</p>\n<p>| 200 | 100% |  |  |  |</p>\n<p>| 208 | 100% | 98% |  |  |</p>\n<p>| 216 | 100% |  |  |  |</p>\n<p>| 224 | 98% | 95% | 79% |  |</p>\n<p>| 232 | 100% |  |  |  |</p>\n<p>| 240 | 100% | 97% |  |  |</p>\n<p>| 248 | 100% |  |  |  |</p>\n<p>| 256 | 100% | 95% | 88% | 56% |</p>\n<p>eval f1</p>\n<p>eval f1 | step | 8 groups/step | 16 groups/step | 32 groups/step | 64 groups/step |</p>\n<p>|---|---|---|---|---|</p>\n<p>| 0 | 0.19 | 0.19 | 0.19 | 0.19 |</p>\n<p>| 64 | 0.20 | 0.18 | 0.24 | 0.17 |</p>\n<p>| 128 | 0.13 | 0.20 | 0.20 | 0.21 |</p>\n<p>| 192 | 0.12 | 0.11 | 0.20 | 0.29 |</p>\n<p>| 256 | 0.14 | 0.18 | 0.19 | 0.24 |</p>\n<p>At a glance, the eval results are in line with expectations: at the end of training, 64 groups/step has the highest eval F1, and each smaller batch size monotonically decreases in final eval F1. Larger batches also give smoother reward and eval curves.</p>\n<p>There’s an interesting contradiction in the charts, though. Although 64 groups/step ends with the highest eval score, it has the lowest train reward at the end of training. Also, 8 groups/step has the highest train reward, but the lowest F1 score.</p>\n<p>This has to do with our reward setup and training dynamics — the reward on a rollout that curates nothing is -0.2, which far outweighs the average reward on most valid rollouts. So, the fastest way to raise train reward early on is to just curate any document. Because 8 groups/step takes more updates than any other run, the model starts to produce valid rollouts sooner. The middle chart above shows this: as batch size decreases, the model reaches &gt;75% valid train rollouts faster. The problem from there is that noisy updates (due to small batch size) keep the policy’s curation quality from ever meaningfully improving. What we’re seeing here is actually a form of reward hacking. The next section will explore this more.</p>\n<p>Since 64 groups/step gives the most stable training and has the highest eval score, we will use it for our full training runs.  We could technically increase groups/step further, but for the sake of budget we will stop here.</p>\n<p>Scaling up training</p>\n<p>For our next set of training runs, we will move to the larger dataset of 1,024 queries and 237,533 chunks. This is 4x the data samples and 2x the corpus size of the dataset we used for ablations. We train for 1.5 epochs, which at 64 groups/step is 24 total updates.</p>\n<p>Full training configuration</p>\n<p>| model |</p>\n<p>| model | gpt-oss-20b |</p>\n<p>| adapter | LoRA, rank 32 |</p>\n<p>| optimisation |</p>\n<p>| learning rate | 1e-4 |</p>\n<p>| optimizer | Adam, β₁ 0.9, β₂ 0.95, ε 1e-8 |</p>\n<p>| weight decay | 0 |</p>\n<p>| KL penalty to reference | 0.0 |</p>\n<p>| updates per batch | 1 |</p>\n<p>| sampling temperature | 1.0 |</p>\n<p>| harness |</p>\n<p>| max turns | 40 |</p>\n<p>| max new tokens per turn | 2,048 |</p>\n<p>| context length limit | 30,720 tokens |</p>\n<p>| curated set cap | 30 |</p>\n<p>We first look at training metrics over the full run. In the spirit of openness, every training rollout can be viewed in the explorer below.</p>\n<p>train reward</p>\n<p>train reward | step | train reward |</p>\n<p>|---|---|</p>\n<p>| 0 | -0.06 |</p>\n<p>| 1 | 0.02 |</p>\n<p>| 2 | 0.01 |</p>\n<p>| 3 | 0.03 |</p>\n<p>| 4 | 0.04 |</p>\n<p>| 5 | 0.10 |</p>\n<p>| 6 | 0.17 |</p>\n<p>| 7 | 0.15 |</p>\n<p>| 8 | 0.14 |</p>\n<p>| 9 | 0.15 |</p>\n<p>| 10 | 0.13 |</p>\n<p>| 11 | 0.17 |</p>\n<p>| 12 | 0.15 |</p>\n<p>| 13 | 0.16 |</p>\n<p>| 14 | 0.18 |</p>\n<p>| 15 | 0.21 |</p>\n<p>| 16 | 0.25 |</p>\n<p>| 17 | 0.17 |</p>\n<p>| 18 | 0.22 |</p>\n<p>| 19 | 0.23 |</p>\n<p>| 20 | 0.21 |</p>\n<p>| 21 | 0.27 |</p>\n<p>| 22 | 0.24 |</p>\n<p>| 23 | 0.21 |</p>\n<p>valid rollouts</p>\n<p>valid rollouts | step | valid rollouts |</p>\n<p>|---|---|</p>\n<p>| 0 | 30% |</p>\n<p>| 1 | 42% |</p>\n<p>| 2 | 47% |</p>\n<p>| 3 | 64% |</p>\n<p>| 4 | 79% |</p>\n<p>| 5 | 92% |</p>\n<p>| 6 | 96% |</p>\n<p>| 7 | 98% |</p>\n<p>| 8 | 98% |</p>\n<p>| 9 | 98% |</p>\n<p>| 10 | 98% |</p>\n<p>| 11 | 97% |</p>\n<p>| 12 | 94% |</p>\n<p>| 13 | 95% |</p>\n<p>| 14 | 94% |</p>\n<p>| 15 | 90% |</p>\n<p>| 16 | 92% |</p>\n<p>| 17 | 92% |</p>\n<p>| 18 | 95% |</p>\n<p>| 19 | 96% |</p>\n<p>| 20 | 96% |</p>\n<p>| 21 | 94% |</p>\n<p>| 22 | 94% |</p>\n<p>| 23 | 96% |</p>\n<p>entropy (nats)</p>\n<p>entropy (nats) | step | entropy (nats) |</p>\n<p>|---|---|</p>\n<p>| 0 | 1.1 |</p>\n<p>| 1 | 1.0 |</p>\n<p>| 2 | 1.0 |</p>\n<p>| 3 | 1.0 |</p>\n<p>| 4 | 1.0 |</p>\n<p>| 5 | 0.9 |</p>\n<p>| 6 | 0.8 |</p>\n<p>| 7 | 0.8 |</p>\n<p>| 8 | 0.8 |</p>\n<p>| 9 | 0.8 |</p>\n<p>| 10 | 0.8 |</p>\n<p>| 11 | 0.8 |</p>\n<p>| 12 | 0.8 |</p>\n<p>| 13 | 0.8 |</p>\n<p>| 14 | 0.9 |</p>\n<p>| 15 | 0.9 |</p>\n<p>| 16 | 0.9 |</p>\n<p>| 17 | 0.9 |</p>\n<p>| 18 | 0.9 |</p>\n<p>| 19 | 0.9 |</p>\n<p>| 20 | 0.9 |</p>\n<p>| 21 | 0.9 |</p>\n<p>| 22 | 0.9 |</p>\n<p>| 23 | 0.8 |</p>\n<p>Train reward moves the way we’d expect it to here: up rapidly at first, and then slowing down as training matures. The training shape also mirrors what we see in ablations: the reward first teaches the model to curate. After that, reward moves around as the model starts to learn how to curate well instead of just how to curate. You can think of the -0.2 floor on no curations as jump-starting curation behavior.</p>\n<p>Entropy is another useful metric to watch during reinforcement learning.  Entropy is measured in nats here. One nat is 1 / ln 2, or about 1.44 bits.  It measures how uncertain the policy’s output distribution is over the course of a rollout. High entropy means at several points in the rollout, the model is considering several possible actions (next tokens). As we’ve discussed earlier, this is good for RL because it means the model will explore more and hopefully stumble into behaviors that get rewarded. Low entropy isn’t inherently bad, but a rapid drop during training is a sign the model is becoming overconfident or committing to a single set of behaviors. This is known as entropy collapse and is particularly bad for GRPO — a policy that always takes the same actions produces groups with  no variance, and groups with no variance produce no gradient.</p>\n<p>During healthy training, entropy typically decreases gradually as reward improves. That is exactly what we see here: entropy starts at 1.1 nats, gradually goes to ~0.8 a quarter of the way through training, and then stays there.</p>\n<p>Let’s see how our policy performs on held-out data now.</p>\n<p>eval f1</p>\n<p>eval f1 | step | eval f1 |</p>\n<p>|---|---|</p>\n<p>| 0 | 0.19 |</p>\n<p>| 2 | 0.22 |</p>\n<p>| 4 | 0.17 |</p>\n<p>| 6 | 0.19 |</p>\n<p>| 8 | 0.21 |</p>\n<p>| 10 | 0.17 |</p>\n<p>| 12 | 0.16 |</p>\n<p>| 14 | 0.24 |</p>\n<p>| 16 | 0.23 |</p>\n<p>| 18 | 0.29 |</p>\n<p>| 20 | 0.31 |</p>\n<p>| 22 | 0.28 |</p>\n<p>| 24 | 0.33 |</p>\n<p>eval precision</p>\n<p>eval precision | step | eval precision |</p>\n<p>|---|---|</p>\n<p>| 0 | 0.30 |</p>\n<p>| 2 | 0.38 |</p>\n<p>| 4 | 0.30 |</p>\n<p>| 6 | 0.29 |</p>\n<p>| 8 | 0.32 |</p>\n<p>| 10 | 0.27 |</p>\n<p>| 12 | 0.17 |</p>\n<p>| 14 | 0.35 |</p>\n<p>| 16 | 0.30 |</p>\n<p>| 18 | 0.36 |</p>\n<p>| 20 | 0.40 |</p>\n<p>| 22 | 0.36 |</p>\n<p>| 24 | 0.40 |</p>\n<p>eval recall</p>\n<p>eval recall | step | eval recall |</p>\n<p>|---|---|</p>\n<p>| 0 | 0.16 |</p>\n<p>| 2 | 0.17 |</p>\n<p>| 4 | 0.13 |</p>\n<p>| 6 | 0.16 |</p>\n<p>| 8 | 0.17 |</p>\n<p>| 10 | 0.14 |</p>\n<p>| 12 | 0.18 |</p>\n<p>| 14 | 0.21 |</p>\n<p>| 16 | 0.20 |</p>\n<p>| 18 | 0.27 |</p>\n<p>| 20 | 0.28 |</p>\n<p>| 22 | 0.26 |</p>\n<p>| 24 | 0.32 |</p>\n<p>Progress is uneven early on as the model learns our harness, and eval performance only starts to improve around halfway through training. Mean eval F1 rises from 0.193 at the start to 0.332 at update 24, driven almost entirely by recall, which doubles from 0.16 to 0.32. Precision goes up as well, moving from 0.31 to 0.40. This is a good result because it means we are getting an all around improvement. That is, we aren’t trading off precision for recall in any way.</p>\n<p>The model performance doesn’t come at the cost of test-time compute, either. As the following charts show, the turn count actually falls slightly over training and output tokens go down, which means our final model is finding more relevant docs with less searching.</p>\n<p>turns per episode</p>\n<p>turns per episode | step | turns per episode |</p>\n<p>|---|---|</p>\n<p>| 0 | 18 |</p>\n<p>| 1 | 18 |</p>\n<p>| 2 | 18 |</p>\n<p>| 3 | 16 |</p>\n<p>| 4 | 15 |</p>\n<p>| 5 | 13 |</p>\n<p>| 6 | 12 |</p>\n<p>| 7 | 12 |</p>\n<p>| 8 | 12 |</p>\n<p>| 9 | 12 |</p>\n<p>| 10 | 12 |</p>\n<p>| 11 | 12 |</p>\n<p>| 12 | 14 |</p>\n<p>| 13 | 14 |</p>\n<p>| 14 | 16 |</p>\n<p>| 15 | 15 |</p>\n<p>| 16 | 16 |</p>\n<p>| 17 | 16 |</p>\n<p>| 18 | 15 |</p>\n<p>| 19 | 15 |</p>\n<p>| 20 | 14 |</p>\n<p>| 21 | 14 |</p>\n<p>| 22 | 16 |</p>\n<p>| 23 | 17 |</p>\n<p>output tokens per episode</p>\n<p>output tokens per episode | step | output tokens per episode |</p>\n<p>|---|---|</p>\n<p>| 0 | 3.998k |</p>\n<p>| 1 | 3.47k |</p>\n<p>| 2 | 3.05k |</p>\n<p>| 3 | 2.621k |</p>\n<p>| 4 | 2.202k |</p>\n<p>| 5 | 1.754k |</p>\n<p>| 6 | 1.408k |</p>\n<p>| 7 | 1.438k |</p>\n<p>| 8 | 1.414k |</p>\n<p>| 9 | 1.497k |</p>\n<p>| 10 | 1.698k |</p>\n<p>| 11 | 1.709k |</p>\n<p>| 12 | 1.972k |</p>\n<p>| 13 | 1.954k |</p>\n<p>| 14 | 2.328k |</p>\n<p>| 15 | 2.461k |</p>\n<p>| 16 | 2.527k |</p>\n<p>| 17 | 2.898k |</p>\n<p>| 18 | 2.693k |</p>\n<p>| 19 | 2.654k |</p>\n<p>| 20 | 2.673k |</p>\n<p>| 21 | 2.649k |</p>\n<p>| 22 | 2.825k |</p>\n<p>| 23 | 2.74k |</p>\n<p>So, we have a full training run that works and pushes eval scores noticeably higher. This is a great result! In all likelihood, we haven’t saturated our training setup, either—train reward is still steadily going up around update 24. We’d probably continue to see eval gains by scaling up the dataset or running for another epoch…</p>\n<p>But that’s a boring next step to explore here. There are many other fun ways to squeeze out performance from our dataset.</p>\n<p>Aside: What happened with 8 groups/step?</p>\n<p>In our ablations, we saw that the 8 groups/step run quickly learned to curate, but this failed to translate over meaningfully into eval improvements. We extended that run over the larger dataset to see if more training steps would help. The charts below show the results of training over 8 groups/step vs 64 groups/step over the same first 384 queries of the larger dataset.</p>\n<p>train reward</p>\n<p>train reward | step | 64 groups/step | 8 groups/step |</p>\n<p>|---|---|---|</p>\n<p>| 0 | -0.06 | 0.06 |</p>\n<p>| 8 |  | 0.03 |</p>\n<p>| 16 |  | 0.06 |</p>\n<p>| 24 |  | 0.03 |</p>\n<p>| 32 |  | 0.13 |</p>\n<p>| 40 |  | 0.07 |</p>\n<p>| 48 |  | 0.23 |</p>\n<p>| 56 |  | 0.16 |</p>\n<p>| 64 | 0.02 | 0.15 |</p>\n<p>| 72 |  | 0.21 |</p>\n<p>| 80 |  | 0.30 |</p>\n<p>| 88 |  | 0.18 |</p>\n<p>| 96 |  | 0.27 |</p>\n<p>| 104 |  | 0.12 |</p>\n<p>| 112 |  | 0.09 |</p>\n<p>| 120 |  | 0.20 |</p>\n<p>| 128 | 0.01 | 0.11 |</p>\n<p>| 136 |  | 0.12 |</p>\n<p>| 144 |  | 0.13 |</p>\n<p>| 152 |  | 0.06 |</p>\n<p>| 160 |  | 0.10 |</p>\n<p>| 168 |  | 0.15 |</p>\n<p>| 176 |  | 0.08 |</p>\n<p>| 184 |  | 0.17 |</p>\n<p>| 192 | 0.03 | 0.09 |</p>\n<p>| 200 |  | 0.14 |</p>\n<p>| 208 |  | 0.09 |</p>\n<p>| 216 |  | 0.13 |</p>\n<p>| 224 |  | 0.09 |</p>\n<p>| 232 |  | 0.09 |</p>\n<p>| 240 |  | 0.21 |</p>\n<p>| 248 |  | 0.30 |</p>\n<p>| 256 | 0.04 | 0.10 |</p>\n<p>| 264 |  | 0.18 |</p>\n<p>| 272 |  | 0.23 |</p>\n<p>| 280 |  | 0.03 |</p>\n<p>| 288 |  | 0.10 |</p>\n<p>| 296 |  | 0.14 |</p>\n<p>| 304 |  | 0.25 |</p>\n<p>| 312 |  | 0.19 |</p>\n<p>| 320 | 0.10 | 0.19 |</p>\n<p>| 328 |  | 0.25 |</p>\n<p>| 336 |  | 0.36 |</p>\n<p>| 344 |  | 0.13 |</p>\n<p>| 352 |  | 0.15 |</p>\n<p>| 360 |  | 0.08 |</p>\n<p>| 368 |  | 0.21 |</p>\n<p>| 376 |  | 0.16 |</p>\n<p>| 384 | 0.17 |  |</p>\n<p>curated documents</p>\n<p>curated documents | step | 64 groups/step | 8 groups/step |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 1 | 1 |</p>\n<p>| 8 |  | 1 |</p>\n<p>| 16 |  | 1 |</p>\n<p>| 24 |  | 1 |</p>\n<p>| 32 |  | 1 |</p>\n<p>| 40 |  | 1 |</p>\n<p>| 48 |  | 2 |</p>\n<p>| 56 |  | 3 |</p>\n<p>| 64 | 1 | 2 |</p>\n<p>| 72 |  | 3 |</p>\n<p>| 80 |  | 4 |</p>\n<p>| 88 |  | 4 |</p>\n<p>| 96 |  | 5 |</p>\n<p>| 104 |  | 5 |</p>\n<p>| 112 |  | 6 |</p>\n<p>| 120 |  | 7 |</p>\n<p>| 128 | 1 | 6 |</p>\n<p>| 136 |  | 9 |</p>\n<p>| 144 |  | 8 |</p>\n<p>| 152 |  | 7 |</p>\n<p>| 160 |  | 7 |</p>\n<p>| 168 |  | 7 |</p>\n<p>| 176 |  | 7 |</p>\n<p>| 184 |  | 7 |</p>\n<p>| 192 | 1 | 9 |</p>\n<p>| 200 |  | 12 |</p>\n<p>| 208 |  | 9 |</p>\n<p>| 216 |  | 9 |</p>\n<p>| 224 |  | 11 |</p>\n<p>| 232 |  | 11 |</p>\n<p>| 240 |  | 9 |</p>\n<p>| 248 |  | 13 |</p>\n<p>| 256 | 1 | 18 |</p>\n<p>| 264 |  | 13 |</p>\n<p>| 272 |  | 12 |</p>\n<p>| 280 |  | 17 |</p>\n<p>| 288 |  | 16 |</p>\n<p>| 296 |  | 17 |</p>\n<p>| 304 |  | 21 |</p>\n<p>| 312 |  | 21 |</p>\n<p>| 320 | 2 | 24 |</p>\n<p>| 328 |  | 22 |</p>\n<p>| 336 |  | 23 |</p>\n<p>| 344 |  | 24 |</p>\n<p>| 352 |  | 24 |</p>\n<p>| 360 |  | 24 |</p>\n<p>| 368 |  | 25 |</p>\n<p>| 376 |  | 25 |</p>\n<p>| 384 | 2 |  |</p>\n<p>entropy (nats)</p>\n<p>entropy (nats) | step | 64 groups/step | 8 groups/step |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 1.1 | 1.1 |</p>\n<p>| 8 |  | 1.1 |</p>\n<p>| 16 |  | 1.0 |</p>\n<p>| 24 |  | 0.9 |</p>\n<p>| 32 |  | 0.9 |</p>\n<p>| 40 |  | 0.9 |</p>\n<p>| 48 |  | 0.8 |</p>\n<p>| 56 |  | 0.7 |</p>\n<p>| 64 | 1.0 | 0.8 |</p>\n<p>| 72 |  | 0.7 |</p>\n<p>| 80 |  | 0.7 |</p>\n<p>| 88 |  | 0.7 |</p>\n<p>| 96 |  | 0.7 |</p>\n<p>| 104 |  | 0.7 |</p>\n<p>| 112 |  | 0.7 |</p>\n<p>| 120 |  | 0.7 |</p>\n<p>| 128 | 1.0 | 0.8 |</p>\n<p>| 136 |  | 0.7 |</p>\n<p>| 144 |  | 0.8 |</p>\n<p>| 152 |  | 0.7 |</p>\n<p>| 160 |  | 0.7 |</p>\n<p>| 168 |  | 0.7 |</p>\n<p>| 176 |  | 0.7 |</p>\n<p>| 184 |  | 0.8 |</p>\n<p>| 192 | 1.0 | 0.7 |</p>\n<p>| 200 |  | 0.8 |</p>\n<p>| 208 |  | 0.7 |</p>\n<p>| 216 |  | 0.7 |</p>\n<p>| 224 |  | 0.7 |</p>\n<p>| 232 |  | 0.6 |</p>\n<p>| 240 |  | 0.6 |</p>\n<p>| 248 |  | 0.7 |</p>\n<p>| 256 | 1.0 | 0.5 |</p>\n<p>| 264 |  | 0.6 |</p>\n<p>| 272 |  | 0.6 |</p>\n<p>| 280 |  | 0.5 |</p>\n<p>| 288 |  | 0.5 |</p>\n<p>| 296 |  | 0.5 |</p>\n<p>| 304 |  | 0.4 |</p>\n<p>| 312 |  | 0.5 |</p>\n<p>| 320 | 0.9 | 0.4 |</p>\n<p>| 328 |  | 0.4 |</p>\n<p>| 336 |  | 0.4 |</p>\n<p>| 344 |  | 0.4 |</p>\n<p>| 352 |  | 0.3 |</p>\n<p>| 360 |  | 0.3 |</p>\n<p>| 368 |  | 0.3 |</p>\n<p>| 376 |  | 0.3 |</p>\n<p>| 384 | 0.8 |  |</p>\n<p>eval f1</p>\n<p>eval f1 | step | 64 groups/step | 8 groups/step |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.19 | 0.19 |</p>\n<p>| 128 | 0.25 | 0.12 |</p>\n<p>| 256 | 0.15 | 0.16 |</p>\n<p>| 384 | 0.17 | 0.16 |</p>\n<p>eval precision</p>\n<p>eval precision | step | 64 groups/step | 8 groups/step |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.30 | 0.30 |</p>\n<p>| 128 | 0.38 | 0.11 |</p>\n<p>| 256 | 0.30 | 0.14 |</p>\n<p>| 384 | 0.29 | 0.13 |</p>\n<p>eval recall</p>\n<p>eval recall | step | 64 groups/step | 8 groups/step |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.16 | 0.16 |</p>\n<p>| 128 | 0.17 | 0.14 |</p>\n<p>| 256 | 0.13 | 0.23 |</p>\n<p>| 384 | 0.16 | 0.29 |</p>\n<p>The 8 groups/step run reaches higher training reward, but it gets there by learning to just curate everything — although eval recall rises from 0.16 to 0.29, precision drops from 0.31 to 0.13. Entropy also falls from 1.1 to 0.34 nats over the course of training. This is a clear sign of entropy collapse. It’s likely that learning will outright stop if we were to continue training.</p>\n<p>Part of this is our reward choice. We chose β = 4 to push the model towards recall over precision, and F4 weights recall 16x. Small-batch updates are also noisier, so once the policy stumbles onto the “curate every document you see” strategy, there is little pushing the model back to something more reasonable.</p>\n<p>One way to mitigate the reward bias here would be to anneal β over the course of training, as is done in Chroma Context-1.</p>\n<p>Reward shaping</p>\n<p>Every run so far has used the plain F4 score as the reward. It worked, in the sense that eval scores went up. But reading the rollouts directly and exploring a different set of training metrics shows two other ways we can improve the policy.</p>\n<p>trajectory vs curated recall</p>\n<p>trajectory vs curated recall | step | trajectory recall | curated recall |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.46 | 0.12 |</p>\n<p>| 1 | 0.51 | 0.16 |</p>\n<p>| 2 | 0.40 | 0.13 |</p>\n<p>| 3 | 0.34 | 0.11 |</p>\n<p>| 4 | 0.29 | 0.09 |</p>\n<p>| 5 | 0.32 | 0.14 |</p>\n<p>| 6 | 0.39 | 0.22 |</p>\n<p>| 7 | 0.38 | 0.24 |</p>\n<p>| 8 | 0.36 | 0.23 |</p>\n<p>| 9 | 0.33 | 0.21 |</p>\n<p>| 10 | 0.26 | 0.19 |</p>\n<p>| 11 | 0.33 | 0.24 |</p>\n<p>| 12 | 0.29 | 0.21 |</p>\n<p>| 13 | 0.30 | 0.21 |</p>\n<p>| 14 | 0.27 | 0.18 |</p>\n<p>| 15 | 0.38 | 0.26 |</p>\n<p>| 16 | 0.40 | 0.28 |</p>\n<p>| 17 | 0.30 | 0.20 |</p>\n<p>| 18 | 0.39 | 0.28 |</p>\n<p>| 19 | 0.37 | 0.28 |</p>\n<p>| 20 | 0.36 | 0.27 |</p>\n<p>| 21 | 0.41 | 0.32 |</p>\n<p>| 22 | 0.39 | 0.29 |</p>\n<p>| 23 | 0.36 | 0.27 |</p>\n<p>Trajectory recall measures how many relevant chunks the model encountered while searching.      tool-call format drift</p>\n<p>tool-call format drift | step | of calls | of episodes |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 27.3% | 37.9% |</p>\n<p>| 1 | 13.7% | 29.3% |</p>\n<p>| 2 | 15% | 26.4% |</p>\n<p>| 3 | 9.1% | 16.2% |</p>\n<p>| 4 | 11.5% | 18.9% |</p>\n<p>| 5 | 12.4% | 18.8% |</p>\n<p>| 6 | 14.8% | 17.8% |</p>\n<p>| 7 | 25% | 32.2% |</p>\n<p>| 8 | 21.2% | 32.2% |</p>\n<p>| 9 | 19.7% | 32.4% |</p>\n<p>| 10 | 14.6% | 33.6% |</p>\n<p>| 11 | 26.8% | 41.6% |</p>\n<p>| 12 | 28.3% | 55.9% |</p>\n<p>| 13 | 47.2% | 67.2% |</p>\n<p>| 14 | 49.6% | 74.2% |</p>\n<p>| 15 | 62.6% | 79.9% |</p>\n<p>| 16 | 53.8% | 78.3% |</p>\n<p>| 17 | 58.5% | 89.3% |</p>\n<p>| 18 | 67.3% | 90.2% |</p>\n<p>| 19 | 67.9% | 92.6% |</p>\n<p>| 20 | 63% | 92.4% |</p>\n<p>| 21 | 62.6% | 91% |</p>\n<p>| 22 | 67.7% | 96.7% |</p>\n<p>| 23 | 75% | 98.8% |</p>\n<p>Measures % of tool calls or episodes with incorrect harmony format.</p>\n<ul><li>Curation recall is bounded by trajectory recall. The model can only curate what it sees as it searches, and in the latter half of training, curation recall consistently sits 0.1 lower than trajectory recall. Trajectory recall also drifts down over training, indicating that the model is getting better at picking relevant docs from searches, but slightly worse at actually searching.</li><li>Tool-call format adherence degrades over training. As training progresses, the share of incorrectly formatted tool calls nearly triples. Luckily, our parser is lenient, so this doesn’t seem to impact training performance too much.</li></ul>\n<p>These observations suggest two ways we can improve the model further: reward the model for encountering more evidence during search, and punish it for bad tool calls.</p>\n<p>Rewarding discovery</p>\n<p>The trajectory recall chart tells us that the model either isn’t searching well or isn’t searching enough. There are two ways we can influence this behavior.</p>\n<p>First, we can reward the behavior directly: a bonus per search or grep call, or for more varied queries, and hope that together with the F4 reward, this leads to better coverage. However, a bonus just for searching is easy to hack. The model could just spam irrelevant calls and collect the bonus without learning any performance-improving behavior. We’d need to add a turn penalty to hold it back. This adds a second knob to tune and complicates the reward.</p>\n<p>Another route is to add an explicit trajectory recall term to our reward. This is a better route because we’re rewarding the outcome we want rather than the activity it takes to get there. As our charts show, trajectory recall is directionally aligned with curated recall, and it’s also harder to game — extra searches don’t earn anything unless they are helpful. This is the route we’ll experiment with.</p>\n<p>Concretely, we add a term for the fraction of the query’s facts that have a supporting chunk anywhere in the trajectory, curated or not:</p>\n<p>r(C)={−0.2F4+0.2⋅Rtrajif ∣C∣=0otherwise</p>\n<p>The explorer below shows the impact of different reward shapes on a group. Seven of the eight rollouts score zero F4 and are indistinguishable under the older reward, but separable under either bonus.</p>\n<p>Apart from the reward, every setting in this run matches the full run above.</p>\n<p>eval f1</p>\n<p>eval f1 | step | w/o trajectory reward | w trajectory reward |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.19 | 0.19 |</p>\n<p>| 2 | 0.22 | 0.21 |</p>\n<p>| 4 | 0.17 | 0.21 |</p>\n<p>| 6 | 0.19 | 0.18 |</p>\n<p>| 8 | 0.21 | 0.18 |</p>\n<p>| 10 | 0.17 | 0.27 |</p>\n<p>| 12 | 0.16 | 0.27 |</p>\n<p>| 14 | 0.24 | 0.25 |</p>\n<p>| 16 | 0.23 | 0.28 |</p>\n<p>| 18 | 0.29 | 0.27 |</p>\n<p>| 20 | 0.31 | 0.32 |</p>\n<p>| 22 | 0.28 | 0.36 |</p>\n<p>| 24 | 0.33 | 0.33 |</p>\n<p>eval curated recall</p>\n<p>eval curated recall | step | w/o trajectory reward | w trajectory reward |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.16 | 0.16 |</p>\n<p>| 2 | 0.17 | 0.15 |</p>\n<p>| 4 | 0.13 | 0.16 |</p>\n<p>| 6 | 0.16 | 0.14 |</p>\n<p>| 8 | 0.17 | 0.14 |</p>\n<p>| 10 | 0.14 | 0.22 |</p>\n<p>| 12 | 0.18 | 0.23 |</p>\n<p>| 14 | 0.21 | 0.20 |</p>\n<p>| 16 | 0.20 | 0.24 |</p>\n<p>| 18 | 0.27 | 0.24 |</p>\n<p>| 20 | 0.28 | 0.29 |</p>\n<p>| 22 | 0.26 | 0.34 |</p>\n<p>| 24 | 0.32 | 0.30 |</p>\n<p>eval trajectory recall</p>\n<p>eval trajectory recall | step | w/o trajectory reward | w trajectory reward |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.43 | 0.43 |</p>\n<p>| 2 | 0.35 | 0.37 |</p>\n<p>| 4 | 0.30 | 0.30 |</p>\n<p>| 6 | 0.29 | 0.26 |</p>\n<p>| 8 | 0.28 | 0.26 |</p>\n<p>| 10 | 0.22 | 0.33 |</p>\n<p>| 12 | 0.22 | 0.33 |</p>\n<p>| 14 | 0.28 | 0.36 |</p>\n<p>| 16 | 0.30 | 0.38 |</p>\n<p>| 18 | 0.36 | 0.47 |</p>\n<p>| 20 | 0.37 | 0.48 |</p>\n<p>| 22 | 0.33 | 0.53 |</p>\n<p>| 24 | 0.44 | 0.49 |</p>\n<p>Eval is run 2x, and plotted points are the averages. Reward plots omitted because they are noncomparable.</p>\n<p>Trajectory recall is clearly better here: after a dip around step 8, the shaped run recovers sooner and stays above the old, plain-F4 run. Eval F1 and curated recall also trend above the old run and even peak higher. Unfortunately, this doesn’t translate directly to a better final score under the same training budget.</p>\n<p>Harmony, the format gpt-oss-20b uses, has a specific header for tool calls. At the start of training, the model gets it wrong on about a quarter of calls.</p>\n<p>Because we use a lenient parse that accepts these slightly off-form calls, the model still gets rewarded whenever the episode they’re in happens to get a good score. This habit ends up compounding. By the final step of training, 99% of episodes contain at least one bad tool call. The model isn’t being told, in any way that reaches the gradient, that those calls are wrong, so it ends up learning that those calls are acceptable.</p>\n<p>The following experiment fixes that directly. Every setting is the same as our full, plain F4 run, except for a slight change to the reward: subtract 0.1 if any tool call in the episode has an off-form header. If two rollouts in a group reach the same F4 but one uses a bad header, the clean one now gets more advantage.</p>\n<p>f1</p>\n<p>f1 | step | w/o format penalty | w format penalty |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.19 | 0.19 |</p>\n<p>| 2 | 0.22 | 0.20 |</p>\n<p>| 4 | 0.17 | 0.27 |</p>\n<p>| 6 | 0.19 | 0.21 |</p>\n<p>| 8 | 0.21 | 0.24 |</p>\n<p>| 10 | 0.17 | 0.20 |</p>\n<p>| 12 | 0.16 | 0.25 |</p>\n<p>| 14 | 0.24 | 0.27 |</p>\n<p>| 16 | 0.23 | 0.34 |</p>\n<p>| 18 | 0.29 | 0.33 |</p>\n<p>| 20 | 0.31 | 0.39 |</p>\n<p>| 22 | 0.28 | 0.36 |</p>\n<p>| 24 | 0.33 | 0.36 |</p>\n<p>recall</p>\n<p>recall | step | w/o format penalty | w format penalty |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 0.16 | 0.16 |</p>\n<p>| 2 | 0.17 | 0.16 |</p>\n<p>| 4 | 0.13 | 0.20 |</p>\n<p>| 6 | 0.16 | 0.17 |</p>\n<p>| 8 | 0.17 | 0.20 |</p>\n<p>| 10 | 0.14 | 0.20 |</p>\n<p>| 12 | 0.18 | 0.26 |</p>\n<p>| 14 | 0.21 | 0.26 |</p>\n<p>| 16 | 0.20 | 0.36 |</p>\n<p>| 18 | 0.27 | 0.35 |</p>\n<p>| 20 | 0.28 | 0.43 |</p>\n<p>| 22 | 0.26 | 0.42 |</p>\n<p>| 24 | 0.32 | 0.43 |</p>\n<p>off-form tool headers</p>\n<p>off-form tool headers | step | w/o format penalty (episodes) | w format penalty (episodes) | w/o format penalty (calls) | w format penalty (calls) |</p>\n<p>|---|---|---|---|---|</p>\n<p>| 0 | 37.9% | 39.8% | 27.3% | 27.8% |</p>\n<p>| 1 | 29.3% | 23.4% | 13.7% | 10.3% |</p>\n<p>| 2 | 26.4% | 21.7% | 15% | 6.1% |</p>\n<p>| 3 | 16.2% | 9.2% | 9.1% | 2.7% |</p>\n<p>| 4 | 18.9% | 6.8% | 11.5% | 2.3% |</p>\n<p>| 5 | 18.8% | 3.9% | 12.4% | 0.7% |</p>\n<p>| 6 | 17.8% | 3.5% | 14.8% | 1% |</p>\n<p>| 7 | 32.2% | 7% | 25% | 2.1% |</p>\n<p>| 8 | 32.2% | 4.9% | 21.2% | 0.5% |</p>\n<p>| 9 | 32.4% | 9.6% | 19.7% | 2.8% |</p>\n<p>| 10 | 33.6% | 12.7% | 14.6% | 3.6% |</p>\n<p>| 11 | 41.6% | 8.6% | 26.8% | 1.1% |</p>\n<p>| 12 | 55.9% | 8% | 28.3% | 0.8% |</p>\n<p>| 13 | 67.2% | 7.2% | 47.2% | 0.6% |</p>\n<p>| 14 | 74.2% | 8.2% | 49.6% | 0.5% |</p>\n<p>| 15 | 79.9% | 8.6% | 62.6% | 0.4% |</p>\n<p>| 16 | 78.3% | 5.9% | 53.8% | 0.4% |</p>\n<p>| 17 | 89.3% | 5.1% | 58.5% | 0.2% |</p>\n<p>| 18 | 90.2% | 4.9% | 67.3% | 0.2% |</p>\n<p>| 19 | 92.6% | 3.1% | 67.9% | 0.1% |</p>\n<p>| 20 | 92.4% | 3.1% | 63% | 0.1% |</p>\n<p>| 21 | 91% | 4.3% | 62.6% | 0.2% |</p>\n<p>| 22 | 96.7% | 4.7% | 67.7% | 0.2% |</p>\n<p>| 23 | 98.8% | 8.2% | 75% | 0.4% |</p>\n<p>turns per episode</p>\n<p>turns per episode | step | w/o format penalty | w format penalty |</p>\n<p>|---|---|---|</p>\n<p>| 0 | 18 | 18 |</p>\n<p>| 1 | 18 | 17 |</p>\n<p>| 2 | 18 | 17 |</p>\n<p>| 3 | 16 | 13 |</p>\n<p>| 4 | 15 | 12 |</p>\n<p>| 5 | 13 | 12 |</p>\n<p>| 6 | 12 | 12 |</p>\n<p>| 7 | 12 | 13 |</p>\n<p>| 8 | 12 | 15 |</p>\n<p>| 9 | 12 | 14 |</p>\n<p>| 10 | 12 | 15 |</p>\n<p>| 11 | 12 | 16 |</p>\n<p>| 12 | 14 | 20 |</p>\n<p>| 13 | 14 | 20 |</p>\n<p>| 14 | 16 | 21 |</p>\n<p>| 15 | 15 | 23 |</p>\n<p>| 16 | 16 | 23 |</p>\n<p>| 17 | 16 | 25 |</p>\n<p>| 18 | 15 | 26 |</p>\n<p>| 19 | 15 | 27 |</p>\n<p>| 20 | 14 | 27 |</p>\n<p>| 21 | 14 | 28 |</p>\n<p>| 22 | 16 | 30 |</p>\n<p>| 23 | 17 | 30 |</p>\n<p>In the tool headers chart, solid lines show the share of episodes with off-form calls; dashed lines show the share of individual calls.</p>\n<p>The extra penalty works, fast. Off-form tool calls fall below 1% within a few steps and stay there for the rest of training. Eval metrics and recall also improve against the baseline. Held-out F1 ends at 0.36 vs 0.33, and recall at 0.43 vs 0.32.</p>\n<p>This result surprised me. I expected the penalty to fix tool call formatting, but not necessarily retrieval quality. Yet the model starts taking longer trajectories (does more searches) and achieves higher recall. I don’t have a concrete explanation for this. One possibility is that the penalty changes which search behaviors get reinforced, and well-formed tool calls are happen to be correlated with better search behaviors. Another is that the model just has more coherent long-context understanding when all of its preceding tool calls are well-formed.</p>\n<p>Either hypothesis would require more training runs to test, but I’ll stop here for the sake of my training budget. What we should take away from this is that a seemingly narrow reward change can sometimes lead to big shifts in the trained policy. We need to be careful and intentional when working on reward functions.</p>\n<p>Some closing thoughts</p>\n<p>We started with a 20B model that struggled to properly execute over our harness, then ran ablations and experiments to drive its eval F1 score from 0.19 up to a peak of 0.39.</p>\n<p>There’s still a lot more we could do here to improve performance. For one, we only trained on a fraction of the total SEC Harness-1 dataset. We could expand the training set or test more learning-rate and batch-size combinations. Tuning one hyperparameter at a time was the right call under our specific budget, but we could have carried over a few configurations to the full run.</p>\n<h2 id=\"a-few-other-more-interesting-experiment-directions\">A few other, more interesting experiment directions</h2>\n<ul><li>Curriculum learning During reinforcement learning, we typically want to give the model tasks that are on the frontier of the model’s capabilities. One way to do this is by getting an empirical pass rate of each task before training. We can use this pass rate to bucket tasks into different difficulty tiers and then sample with different weights over the course of training.</li><li>SFT Warmup In our scaled up training runs, we end up devoting about a third of our data budget just to teaching the model how to use our harness. We can speed this up by adding an initial warmup stage with SFT over some easier tasks.</li><li>Harness Engineering We generally don’t want to be too prescriptive in the system prompt because we want the model to discover behaviors during training. However, that the model initially does so poorly is a sign that we can improve either our system prompt or harness signals a bit.</li><li>Other reward shapes So far, we’ve kept our reward functions deliberately simple. There’s a lot more we can do here. For example: a turn penalty to encourage shorter rollouts, rewards for tool diversity, etc.</li></ul>","headings":[{"level":2,"text":"After each round of searches, consider the following","id":"after-each-round-of-searches-consider-the-following"},{"level":2,"text":"We train with Dr. GRPO, a variant of GRPO. The loop is as follows","id":"we-train-with-dr-grpo-a-variant-of-grpo-the-loop-is-as-follows"},{"level":2,"text":"A few other, more interesting experiment directions","id":"a-few-other-more-interesting-experiment-directions"}]}}