This post is a hands-on introduction to reinforcement learning through training a search agent with group-relative policy optimization (GRPO). Search is a fun place to learn RL because it has so many levers, and each of them visibly changes how the model searches. It is also a domain where, in my experience, a well-designed reward function can influence how a model searches more effectively than system prompt changes or harness engineering.

Our approach is based loosely on the Harness-1 paper from Jiang et al. We use their training dataset and a much-simplified version of their search harness.

Training is done using the Tinker API from Thinking Machines. Many thanks to the team for the training credits to support this work.