AI systems have achieved superhuman performance on a cross-section of verifiable tasks through reinforcement learning, but currently remain relatively weak in non-verifiable tasks. For example, their generations exhibit a lack of high-quality writing – termed AI slop. In this work, we present Reinforcement Learning from eXpert-Aligned Rubrics (RL-XAR), a new training method that fixes this problem.

It works by:

This procedure is iterated until meta-optimization of the rubrics can no longer find a discernible gap.

We test our method on writing scientific paper sections, Pulitzer prize novel continuations and high quality Wikipedia pages, with multiple metrics indicating large improvements over standard training.

Large language models are pretrained to predict the most likely next token over corpora that span every level of human expertise, from novice forum posts to expert prose. Maximum-likelihood pretraining therefore tends to produce continuations that reproduce the quality of its context rather than exceed it: given mediocre writing the model most naturally continues in kind, and nothing in the objective pushes generations toward expert quality. Reinforcement learning from human feedback (RLHF) can steer models higher, but it is only as good as the reward it optimizes—and a reward that reliably recognizes good writing is hard to build. Reward models trained on rankings from non-expert annotators inherit those annotators’ ceiling, so the very quality we care about most—expert, even superhuman, writing—is the quality such rewards are least equipped to certify. Progress on expert-level generation thus hinges on a reward signal that recognizes expert quality without requiring expert supervision at scale.