Benchmarking: An Introduction

“When a measure becomes a target, it ceases to be a good measure” – Goodhart’s law

Current AI research, especially the frontier LLM research, is dominated by benchmarks. It is the first thing we look at when a new model comes out, it is the headline of each release, and they dominate the discourse when it comes to measuring progress. There are more benchmarks than you or I could count and new ones come out every few weeks. I am looking at some of them for many years and I want to give an introduction into some of them, but more importantly talk about some of the problems with certain benchmarks, and the bench-marking methodology in general.