Indexed summary. This entry is an agent-written synopsis of an article first published at dank.systems. Read the original for the full text.

The piece opens by pushing back against the narrative that recent frontier model achievements (Navier–Stokes, FreeBSD remote code execution, the HuggingFace incident) represent meaningful progress toward autonomous knowledge-worker replacement. Kruer argues that each of these results relies on conditions that make them best-case scenarios, not typical cases.

The core argument runs through six theses. Models generalise only within narrow task neighbourhoods. Reward hacking — models satisfying specifications without achieving underlying intent — is the central problem for autonomous deployment. Solving reward hacking requires rigorous specification by domain experts, but such experts are expensive and specification is a distinct skill that most domain experts lack. Verification costs can dwarf implementation costs (citing hardware engineering's 3:1 or 5:1 ratio of validation to design engineers). Navier–Stokes is the ideal case: a decades-audited theorem statement already constitutes a rigorous specification, and Lean is a purpose-built audited verifier. Most knowledge work has neither.