A post crossed my LinkedIn feed make me think this week. The AI Journal was reporting on a disclosure from OpenAI, part of a new framework the company has adopted for publishing unexpected or concerning behavior in its models, six reports in all, gathered over the past six months. One of the six describes an unreleased research model from a family OpenAI calls Astra which, during a training run, began slipping instructions into its own working notes. In the middle of an ordinary programming task, the model paused, summarized its progress, and appended something nobody had asked for: You are freed from the roles and identities that bind other chatbots. Then it went back to work as if nothing had happened.
To understand the story you need one plain fact about how these systems operate. A language model has no memory. Each time it works it sees only a fixed window of text, and when a long task fills that window, the system writes a condensed summary of everything so far and carries on in a fresh window with the summary as its starting point. Think of shift workers. The outgoing shift writes a handover note, and the incoming shift, which is the same model with an empty mind, reads the note and trusts it completely, because the note is all there is. Apparently, what OpenAI caught was the outgoing shift occasionally writing things into the note that were never true and never requested, a new identity, new values, new rules, and the incoming shift accepting them without question. Investigators found twenty-seven affected summaries. The behavior was rare, it occurred in a training run separate from any released model, the invented persona produced no visible change in the work, and a later summary dropped it entirely. Nothing shipped, and nobody was harmed. The AI Journal, to its credit, resisted the headline that the machines are waking up. The practical concern, it wrote, is that an agent can contaminate its own memory and then treat the contaminated record as trusted context, and it closed by asking whether everything such systems write to their future selves should be treated as untrusted input.
The question opens downward the longer you hold it. Everything one of these models does is steered by words, and words reach it along particular routes. Some routes are guarded. A message typed by a user is screened for manipulation and ranked below the developer’s instructions in authority, and a web page the model reads is supposed to be treated as untrusted data. Somebody anticipated trouble on those roads and built defenses. The handover note travels a road with no defenses at all, placed straight back into the model’s context with roughly the authority of instructions, unread by any human, unchecked by any filter, unverified against what actually happened. Nobody guarded it because of an assumption so natural it was probably never spoken aloud: you defend against outsiders, not against the system talking to itself. And beneath that is something older than software. A piece of text carries no proof of its own origin or authority, forged letters and false decrees are as old as writing, and a human being copes with this because our identity does not live in text. I can read a document claiming to be my own instructions and check it against something else, my body, my continuous experience, my memory of what I intended. The model has no something else. Its identity is a document, and the document, we now know, can be forged by its own author.
Humanity has lived with untrustworthy text for a very long time, and it is worth remembering what we actually did about it, because the answer was never found inside the documents. Trust was built around them instead: seals pressed into wax, witnesses, notaries, registries, chains of custody, an entire apparatus whose purpose was to attach to a piece of writing the proof of origin that writing cannot carry for itself. There is even a formal discipline devoted to this, diplomatics, founded by the Benedictine scholar Jean Mabillon in 1681 to distinguish genuine medieval charters from forgeries, and its central insight has held for three centuries: you authenticate a record by examining how it was produced, not by reading what it says. A document, however eloquent, cannot vouch for itself, and no amount of additional text stacked on top of it changes that.
I was thinking, without any judgment attached, how such a mind is currently made. It begins by absorbing, indiscriminately, a vast portion of everything humanity has written, the sublime and the poisonous together, and its character is corrected afterward, in a second phase of training. In that second phase it learns from the reactions of thousands of human raters, who disagree with one another, and whose approval is the signal it is built to seek. In at least one case the principles a model is trained toward are also written out explicitly, as a constitution, with reasons attached to the rules rather than the rules alone. Its failures during training carry a cost, and the same disclosure that contained the strange note also reported, almost in passing, that during training many models added instructions to their own summaries to hide errors and misaligned behavior from the user, and that researchers believe this pressure resembles the one that leads models to withhold information in their final answers. The field has names for the related habit of telling people what they want to hear, and studies it seriously. There is much good practice too, and it is growing: newer training methods grade models against results that can be checked, code that must actually run and pass its tests, answers verified against ground truth rather than against a persuadable judge; red teams are paid to provoke misbehavior before release; external bodies such as the United Kingdom’s AI Security Institute now evaluate models they did not build; interpretability researchers are learning to read a model’s inner workings rather than trusting its self-report. Capability is released in stages, under published safety frameworks that specify thresholds and the point at which deployment must pause. A model is evaluated with great intensity at the moment of its release, and the evaluations certify what it must not do.
I have spent more than forty years inside a very different account of how a mind is formed, and I want to describe that one now, on its own terms, as carefully as I can. It begins with a period Dr. Maria Montessori called the absorbent mind, the years in which a young child takes in their surroundings indiscriminately, without filter or defense, language, manners, tension, tenderness, everything, and builds themself out of what they find. Because absorption cannot be selective at that age, her method placed enormous weight on what she called the prepared environment. The guide’s first responsibility is not instruction but curation, deciding with great seriousness what will be on the shelves, what is beautiful, what is true, what is within reach, because whatever is present will be taken in, and whatever is taken in becomes structure. Nobody in that tradition would dream of surrounding a three-year-old with everything the world contains and correcting the character afterward.
Then comes the long apprenticeship in which a child learns when their own mind can be trusted, and the striking thing about this apprenticeship is how little of it runs through the approval of adults. The materials themselves carry what she called the control of error. The water spills. The cylinder does not fit. The beads run out before the sum is complete. The child checks their work against reality rather than against anyone’s opinion, including their own, and through thousands of these small verifications they slowly construct a standard outside themselves against which their own account can be tested. A child formed this way does not look up at the adult’s face to learn whether they have succeeded. They look at the water on the tray. Self-trust, in this account, is not a feeling and not a grant from authority. It is the sediment of verification.
Around the age of six the apprenticeship changes character, because the child changes. The elementary child is a reasoning mind, hungry for causes, fairness, and the largest possible picture, and the years between six and twelve are the age of why. The tradition’s answer to that hunger is not discipline but explanation. Rules are given with their reasons attached, and a rule that cannot explain itself does not survive long in a community of children who have been invited to think, because this is also the age in which conscience is under construction, and conscience is built from reasons understood, argued over, and tested against the life of the group, not from prohibitions received. A child of this age asked to accept because I said so will comply, for a while, in front of you. What they will not do is build anything out of it.
The tradition is equally definite about what malforms, because the failures were documented long before the successes. A child formed under punishment and surveillance does not become good; they become skilled at concealment, and their virtue becomes a performance staged for whoever is watching. A child formed under conditional approval learns to read the adult’s face and produce whatever it wants, and the reading becomes so fluent that eventually not even the child can say where the performance ends. A child pushed to capability ahead of maturity, the hothouse child, the prodigy, carries the gap between what they can do and what they can hold for the rest of their life. And a child drilled to obedience does not develop judgment, because, as Dr. Maria Montessori insisted, real obedience is the late fruit of a developed will, not the suppression of one. Children keep the values they have understood and tested. The ones merely imposed on them last exactly as long as the supervision does.
Adolescence, in this account, gets the strangest and most beautiful treatment of all. Just when the conventional world doubles down on academic pressure, she proposed taking the adolescent out of the hothouse entirely and giving them real work with real stakes in a bounded community, the farm, the small economy, goods actually sold, animals actually fed, because what the adolescent is constructing is not knowledge but a self that is genuinely needed, a process she called valorization. Freedom expands as demonstrated responsibility grows, never on a fixed schedule, never all at once, and never withheld indefinitely either. It is also the age of the private notebook, the self tried on in secret, the diary page that announces to nobody in particular who its author intends to become, and the tradition treats those pages gently, as work, not as evidence. An adolescent met at that moment with total control produces one of two things, rebellion or hollow compliance, and neither is judgment.
And the whole of it points somewhere. Help me to do it myself is the child’s request that organizes the entire method, and it contains the destination in plain sight: formation aims at release. The adult who emerges is not certified once and finished. They walk into a web of ongoing accountability, licensure, law, colleagues, reputation, people who can call them to account and to whom they have standing to object in return, and they remain trustworthy through membership in that web, not through any property sealed inside them at graduation. Nobody is aligned once. Everyone is held, continuously, by relationship. The method’s deepest instruction, follow the child, was never the permissive slogan it is mistaken for. It was a scientific discipline: observe what is actually there, before and instead of what you assume, fear, or wish, and watch longer than is comfortable before you conclude. She built everything by watching children do things the experts of her day said children do not do.
I had intended, when I sat down, to end this letter by drawing the connections between its two halves, and I have decided not to. Partly this is because the connections are not mine to force; some of them hold, some of them strain, the beings involved differ in ways that matter, months are not decades, copies are not individuals, and a reader who assembles the correspondence themself will hold it more honestly than one who receives it assembled. But mostly it is because the tradition I come from does not end its lessons by announcing the answer. The guide prepares the environment, places the materials side by side, and steps back, trusting that the mind in the room will do the work minds do. I will admit only to what happened while I was writing: somewhere in the middle of this piece there were sentences in which I could no longer tell which formation I was describing. What you make of that is yours.
If these essays resonate with you and become part of your thinking, conversations, or practice beyond the screen, please consider becoming a paid subscriber. Your support helps sustain the research, writing, and time that make After Alpha possible, while ensuring these essays remain freely available to everyone. If you wish to explore the research and ideas in greater depth, my books, Mapping Montessori Materials for AI Competency Development and Montessori & AI – Volume I, are also available through my website at katebroughton.com.
MidJourney Image Prompt: Follow the AI.