---
title: "Lessons from building an automated research scaffold"
slug: lessons-from-building-an-automated-research-scaffold
url: https://listedarticles.com/articles/lessons-from-building-an-automated-research-scaffold
canonical_url: https://www.lesswrong.com/posts/zGaQ3SS9C6So9NXFg/lessons-from-building-an-automated-research-scaffold
content_type: research
language: en
published_at: 2026-10-02T10:36:04.000Z
updated_at: 2026-10-04T08:12:56.195Z
author: "Alejandro Aristizabal, Josh Hills, Dewi Gould, et al."
authored_by: human
publisher: "LessWrong"
publisher_url: https://www.lesswrong.com
topics: ["AI Safety", "AI Agents", "Research", "AI", "Machine Learning"]
license: all-rights-reserved
word_count: 989
reading_minutes: 4
citation: "Alejandro Aristizabal, Josh Hills, Dewi Gould, et al., LessWrong. \"Lessons from building an automated research scaffold.\" 2 Oct 2026. https://www.lesswrong.com/posts/zGaQ3SS9C6So9NXFg/lessons-from-building-an-automated-research-scaffold (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Lessons from building an automated research scaffold

> Arcadia Impact / Equistamp / UKAISI report: an automated alignment-research scaffold failed to uplift researchers on conceptual work, but surfaced auditing, monitoring, and scaffold–model interaction failure modes worth studying next.

# Lessons from building an automated research scaffold

**TL;DR.** We built a scaffold to speed up our own research and gather data on automated alignment research (AAR). It turned out to *not* be valuable for researcher uplift, but was useful for gathering certain failure modes of AAR. Going forward, we plan to study the broader failure modes of AAR and how these automated research systems can be monitored and analyzed.

*This work was carried out by the [Alignment Team](https://www.arcadiaimpact.org/alignment-research) at [Arcadia Impact](https://www.arcadiaimpact.org/) in collaboration with Josh Hills, Falko Galperin, and Denis Lim from [Equistamp](https://www.equistamp.com/), and Aleksandr Bowkis from [UKAISI](https://www.aisi.gov.uk/).*

## The scaffold

The scaffold is given a task description and a description of a metric. During setup, an agent writes an evaluation script for the metric and creates a local and held-out evaluation environment. In cases where the task lacks a clear metric, the evaluation consists of LLM judges. Worker agents run on separate VMs and iterate against the metric using only the local environment. An orchestrator monitors their progress by pulling their transcripts, and can restart and steer them. Workers submit their work as pull requests, which are scored on the held-out environment on separate machines.

## What blocked researcher uplift?

We made our scaffold available to our researchers and found that adoption was low, primarily because researchers found minimal uplift over their existing workflows for most tasks. This was for three main reasons:

**Our scaffold wasn't helpful for conceptual work.** While our scaffold performed well on very narrowly scoped, well-defined objectives with clear metrics, those aren't the main bottleneck of our team's work. By the time a project has been reduced to a well-defined metric to hill climb, most of the work is already done. More importantly, our projects rarely take this shape – most of our time goes on thinking through threat models, designing experiments, and analysing results to decide what to run next.

**Current models have a very spiky skill distribution across automated research tasks.** We find that models are extremely capable at building and running well-specified experiments. However, they are *very poor* at interpreting results and deciding what the interesting next research questions are. I.e., it seems like current training methods do not generalise well to research taste and decision-making, especially if the decision requires conceptual clarity. Since this is a big part of the automated research pipeline, this makes it hard to do end-to-end auto-research.

**It's easy to fall into scaffold-fixing whack-a-mole.** Our initial setup framed the goal of the workers as trying to improve the score, which led to high amounts of reward hacking. We then reframed the task as "doing good science" and specified the metric as an imperfect proxy. This reduced reward hacking frequency, but also made models obsessed with doing ablations and controls on uninteresting things. Once we encountered these problems, we found ourselves bloating system prompts, adding more constraints and components to work around these pitfalls and other model quirks. In doing so, we increased the amount of friction our team faced when trying to use the scaffold, and in many cases the setup cost was not worth the quality of the outputs we got.

## The scaffold was useful for gathering failure modes

Although our scaffold did not speed up our research as much as we hoped, running it was useful for identifying problems that are likely to matter for automated research more broadly.

**Auditing what the agents did is hard, and getting harder.** After a run, we want to answer three questions: what did the agents actually do, were there any undesired behaviors, did they do good science? With the current trend toward agent swarms, this problem seems likely to get substantially harder, and much more expensive.

**This is also a bottleneck for live monitoring for misbehavior.** There's no ground truth for evaluating monitoring in the AAR regime. In our transcripts, humans and models often disagree on what counts as misbehavior and which category it falls under. Additionally, since we don't have good strategies for monitoring multi-agent events, misalignment that encompasses actions across agents can easily be overlooked.

**However, our insights may not transfer beyond the scaffold.** We have very little insight into how labs are doing their automated research work. Even if our scaffold was 'good', it may still present different failure modes compared to those which may occur in labs. This is a core reason why we don't want to over-invest in a single scaffold going forward.

There seem to be many interaction effects between the scaffold and the model's final outputs, such as changes to prompts, roles, components, or objectives, and this makes it difficult to know which findings generalise beyond the scaffold design. We think it's important to evaluate the model and scaffold together, test a broader class of scaffolds, and study which behaviors persist across them.

## Next Steps

We are worried that frontier labs will continue automating AI R&D aggressively, and that this could further accelerate AI development. At the same time, safety research may be harder to automate and fail to keep pace.

We lack a shared understanding of which parts of research are actually automatable, what failure modes there are, and how monitorability changes as we move to larger, longer-horizon multi-agent processes. We therefore want to focus on three things:

1. **Getting an accurate picture of how automated research is being done.** Frontier labs should report *separately* how their capabilities R&D and AI safety research are being automated, including which parts are automated, how much speed-up researchers get, remaining bottlenecks, and how runs are monitored.

2. **Building evaluations for monitorability of autoresearch runs.** We currently lack clear-cut methods of measuring whether an autoresearch run went well: whether it was safe and whether the science produced is valid and useful.

3. **Developing proxy settings to study scalable oversight protocols for auto-alignment research.** As models become more capable, our ability to oversee and understand their outputs will diminish, so we will need better scalable oversight protocols.
