Lessons from building an automated research scaffold

TL;DR. We built a scaffold to speed up our own research and gather data on automated alignment research (AAR). It turned out to not be valuable for researcher uplift, but was useful for gathering certain failure modes of AAR. Going forward, we plan to study the broader failure modes of AAR and how these automated research systems can be monitored and analyzed.

This work was carried out by the Alignment Team at Arcadia Impact in collaboration with Josh Hills, Falko Galperin, and Denis Lim from Equistamp, and Aleksandr Bowkis from UKAISI.

The scaffold

The scaffold is given a task description and a description of a metric. During setup, an agent writes an evaluation script for the metric and creates a local and held-out evaluation environment. In cases where the task lacks a clear metric, the evaluation consists of LLM judges. Worker agents run on separate VMs and iterate against the metric using only the local environment. An orchestrator monitors their progress by pulling their transcripts, and can restart and steer them. Workers submit their work as pull requests, which are scored on the held-out environment on separate machines.