I am beginning this blog post with a quote of Ilya Sutskever from a podcast with Dwarkesh Patel:

Yeah. This is one of the very confusing things about the models right now. How to reconcile the fact that they are doing so well on evals? You look at the evals and you go, “Those are pretty hard evals.” They are doing so well. But the economic impact seems to be dramatically behind. It’s very difficult to make sense of, how can the model, on the one hand, do these amazing things, and then on the other hand, repeat itself twice in some situation?

In my opinion this echoes a plethora of researchers in the AI field. Recently Jev has been all anyone can talk about, and it asks the same exact question: