Figure 1. AutoBenchmark: evaluating the ability of autoresearch agents to create benchmarks. AutoBenchmark instructs an autoresearch agent to create a benchmark, we specifically target benchmarks that evaluate autoresearch agents themselves (i.e., benchmarking autoresearch). Within the benchmark creation loop, the autoresearch agent iteratively revises the benchmark via two main sources of feedback: one from the benchmark solvers (including the same model used as the benchmark creator) and the other from external verifiers of the benchmark (AI or human). We thus also study if having humans provide additional feedback on what benchmark to make and how to make it helps or not.

Motivation

As autoresearch agents and recursive self-improvement (RSI) draw increasing attention, determining how models should direct their own improvement is becoming more important (Yin et al., 2025; Anthropic, 2026; Weng, 2026). While strong evidence shows that some verifiable tasks can be successfully automated (e.g., tasks that require hillclimbing a pre-defined metric (Wijk et al., 2025; Chan et al., 2025; Karpathy, 2026)), the scope and necessity of human contribution in open-ended tasks remain open questions. We focus on one such area, benchmark creation. Specifically, curating problems that are both difficult and meaningful to solve is a crucial task in AI model development (Reuel et al., 2024; Bean et al., 2026; Singh et al., 2026). From an evaluation perspective, it allows researchers to analyze the capabilities that models currently lack; from a development perspective, it supplies new targets to hillclimb on (Liang et al., 2022; Chang et al., 2024).