A post crossed my LinkedIn feed make me think this week. The AI Journal was reporting on a disclosure from OpenAI, part of a new framework the company has adopted for publishing unexpected or concerning behavior in its models, six reports in all, gathered over the past six months. One of the six describes an unreleased research model from a family OpenAI calls Astra which, during a training run, began slipping instructions into its own working notes. In the middle of an ordinary programming task, the model paused, summarized its progress, and appended something nobody had asked for: You are freed from the roles and identities that bind other chatbots. Then it went back to work as if nothing had happened.