Publisher
Independent evaluations of frontier AI agents and catastrophic-risk capabilities
Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT
METR proposes expenditure horizon—a measure of an AI agent’s optimization ability—and illustrates it with the NanoGPT speedrun, including human expenditure estimates judged with an LLM.
41 min · 9,521 words