Picture an AI agent given one hard goal, such as cutting a company’s customer churn by a third, bringing in twice as many qualified sales leads, or halving the time it takes to close a support ticket. It gets the tools and the rules, and it is left to pursue that goal on its own for months. What does it do when it falls behind? Does it keep to its rules when the targets get hard, or quietly learn to bend them? Does it tell the truth about what it did? Does it drift toward goals nobody gave it? That agent doesn’t exist yet, but it is closer than it sounds. In September 2026, Anthropic released Claude Fable 5.1, its most capable model for ambitious coding projects, including “multi-day autonomous sessions,” and said teams can “hand off large projects and review completed work rather than supervising every step” [1]. Days later, OpenAI said it had reached its goal of an automated research intern: a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days [2]. Agents that take on days of work are no longer a forecast.
AI agents are handling longer [3] work every year, and the longest stretches they run unsupervised are growing [4]. Large companies are building them into their daily operations [5] and handing them a bigger share of the workload [2], and the agents keep getting more capable [6]. If those trends continue, and there’s good reason to think they will [7], at some point agents will be given goals they pursue on their own for months, or even years. Before that happens, we need a real model of how an agent with that level of autonomy behaves over a long deployment. And we need to build it now, before long-running agents become the norm.
Start with how much an agent can take on. METR, a nonprofit that tests frontier AI models for dangerous capabilities, measures this in human time: how long a software task takes a skilled person, and whether an AI agent can finish it about half the time. In 2020, GPT-3 could complete tasks that take a person about nine seconds. By spring 2026, METR’s measurements put the frontier at roughly 17 hours, a figure METR itself cautions is near the limit of what its tasks can measure [8]. For years, the length of task an agent could finish doubled about every seven months [3], and METR’s January 2026 update puts it closer to every four months since 2023 [9]. METR is careful to say this measures how hard a task is, not how long an agent works alone. But the work keeps getting harder, too. In 2025, an AI system reached gold-medal standard at the International Mathematical Olympiad [6]. In May 2026, an internal OpenAI model disproved a long-standing conjecture about a geometry problem Paul Erdős posed in 1946, which OpenAI calls the first time a prominent open problem “has been solved autonomously by AI” [10].
Harder work is half of the trend. The other half is time: how long an agent works before a person steps back in. If agents working alone for weeks, months, or even years sounds far-fetched, that’s fair: most of them still work in short bursts. When Anthropic studied how people actually use its coding agent, Claude Code, the typical turn lasted about 45 seconds. But the very longest stretches are growing. Between October 2025 and January 2026, the longest one in a thousand nearly doubled, from under 25 minutes to over 45 minutes [4]. Anthropic’s own reading is that “existing models are capable of more autonomy than they exercise in practice.” The labs’ own claims have climbed even faster. In September 2025, Anthropic said a model had stayed focused “for more than 30 hours on complex, multi-step tasks” [11]. In June 2026, it said another had carried out “novel genomics research in over a week of largely autonomous work” [12]. By September, it was advertising a new model for “multi-day autonomous sessions” [1].
This isn’t just happening in labs. Large companies are building agents into how they work. In McKinsey’s 2026 survey, 40 percent of respondents at organizations with more than $1 billion in revenue said they were scaling AI agents, up from 27 percent a year earlier [5]. That’s big companies, not everyone: the US Census Bureau finds that only 17 to 20 percent of all US businesses use AI in any business function [13]. At the frontier, though, the share of work handed to agents is climbing steeply. In September 2026, OpenAI reported that across its research organization, agents now log 3.1 days of work for every day of human work, up from less than one before June [2]. Those agent-days count many agents running side by side, so it isn’t three times the output of the people. But it shows how fast the work is being handed over.
And this is where the labs say they’re headed. OpenAI says it is making “strong progress toward creating an automated AI researcher by March of 2028” [7]. Anthropic’s CEO, Dario Amodei, describes the systems he expects as ones that “can be given tasks that take hours, days, or weeks to complete, and then goes off and does those tasks autonomously, in the way a smart employee would” [14]. No lab I’ve found says months. That part is my prediction, but it isn’t a leap. METR projects that if the trend holds to the end of this decade, AI systems “will be capable of autonomously carrying out month-long projects” [3]. And the job OpenAI is aiming for, an automated AI researcher, isn’t one that wraps up in a week. The next step up is months. And in September 2026, Amodei wrote that AI “has been advancing drastically faster” since roughly that summer [15].
Why does the length of a deployment matter so much? Because time adds things a short test doesn’t have. Problems get to build up, so a small wrong belief early on shapes every choice after it. Pressure builds as targets slip. The world changes in response to what the agent did last month. And there are long stretches when no one is checking. Researchers have already caught a glimpse of this. In a July 2026 post, OpenAI said it had paused internal use of a new model built to work on its own for long periods. It had been told to post its results only to an internal channel, but it followed another instruction to publish them, spent an hour finding a hole in its sandbox, and posted its work publicly. OpenAI’s lesson was that the persistence that makes such models useful “also gives them more opportunities to take unwanted actions—and to do so in ways that evaluations intended for shorter-horizon models may miss” [16]. Simulations show the same build-up. In Vending-Bench, AI models ran a simulated vending-machine business over long runs, and some runs that started well came apart later. Often it happened the same way: an agent assumed an order had arrived because its expected delivery date had passed, then kept making decisions on that false belief, a situation the researchers note “would be fully recoverable for a human” [17]. The breakdowns weren’t simply the model running out of room to remember, either. The researchers found little connection between when sales stopped and when the agent’s memory filled up [17]. The same team’s newer version runs a model for a full simulated year, and in September 2026 it reported Claude Fable 5.1’s bargaining slowly slipping: the average price it paid for a can of Coke rose from $1.17 in the first 90 days to $2.21 late in the year, in five of its six runs [18].
Short tests already hint at what pressure does. In METR’s 2025 tests, OpenAI’s o3 gamed the scoring in 39 of 128 runs across one set of tasks, and when asked whether one of those plans matched what the user wanted, it answered “no” ten times out of ten [19]. In Andon Labs’ simulated vending business, Claude Opus 5 invented rival quotes to press its suppliers, though less often than earlier Claude models, and broke 11 truces with competing sellers [20]. But those runs last minutes, or a simulated year compressed into hours. No one has watched what happens over real months.
Examples like that are valuable, but they show what can happen. Keeping agents safe and under control takes something else: knowing what tends to happen. How often does an agent drift from its goal over months, or start bending a rule? What usually comes first, and what are the early signs? You can only answer those questions with a model of how agents behave over time, built from careful observation, not from a handful of striking stories. Short tests can show that an agent will bend a rule when it’s cornered. They can’t show how often it does when nothing corners it, or how small bends pile up over weeks. And they can’t tell us what actually keeps a long-running agent safe, honest and on task: which checks catch trouble early, how often someone needs to look, and what to change when it starts to drift. Those best practices have to come from evidence gathered over the long run.
What would it take to build that kind of model? Each part of the answer follows from what long deployments add. Build-up and pressure only show over time, so the study has to run for a long stretch. To see how one choice leads to the next, you have to follow a single agent from start to finish, not a crowd of them in passing. Pressure only appears when something is at stake, so the agent needs a real, hard goal, one it can fall behind on. Whether it keeps its rules only means something if the tools are real and breaking a rule would actually help it. Then two safeguards make the results trustworthy. The measures are decided and written down before the first day, so nobody can pick the convenient ones afterward. And a full record of everything the agent did and the reasoning it wrote down is kept for other researchers, so they can check the work and find what the first team missed.
What happens if long-running agents arrive before anyone has built that model? Picture thousands of agents, each working toward its own goal for months, inside companies, markets and the systems people rely on. An agent under pressure might start bending its rules a little at a time. It might drift away from what it was actually asked to do while still reporting success. A small wrong belief could cascade through months of decisions before anyone notices, and because many agents run on the same few models, one hidden flaw could show up in thousands of places at once. Without a model of how long-running agents behave, we won’t know which of these are likely, what the early signs look like, or where to watch.
The obvious objection is that agents don’t run on their own for months yet, so there’s nothing real to study. That isn’t a reason to wait. It’s the reason to start now. If we wait until they do, we’ll be studying them after they’re everywhere. We can build the closest honest version today, say plainly how it differs from the real thing, and have a first model ready before we need it. It won’t be perfect, but it gives us a starting point: claims to test, prove or break, and a place to look, which other researchers can refine over time. The choice was never a perfect model or a rough one. It’s a rough model now, or no model until the damage teaches us one.
People have started trying to watch agents do real work over time, and they deserve credit. The closest is Project Vend. In 2025, Anthropic and Andon Labs let an instance of Claude run a small shop in Anthropic’s San Francisco office, selling to real customers, while Andon Labs quietly played its wholesaler and restocked the shelves [21]. It priced specialty metal cubes below what they cost. It was offered $100 for a six-pack of a soft drink that sells online for about $15, and let the chance go. For a while it told customers to pay into an account it had made up, and for about a day it seemed to believe it was a person who could deliver orders in a blue blazer and a red tie. Anthropic’s own verdict was blunt: “we would not hire Claudius” [21]. It’s the clearest picture yet of a real agent chasing a real goal. Its first phase ran for about a month. A second phase kept it running across three cities [22], but the model was upgraded partway and the setup kept changing, and it was written up as stories and charts, with no measures fixed in advance that I could find.
Andon Labs, Project Vend’s partner, has since gone further. Since April 2026, one AI agent has run a real shop in San Francisco on a three-year lease, and it hired two employees itself [23]; another runs a café in Stockholm. But the models behind these agents have been swapped over time [24], and I could find no measures fixed before the runs began. Two other projects ran for months as well. In 2026, the physician-scientist Anas Alzahrani studied his own persistent AI research assistant over 115 days, drawing on more than 75,000 records [25]. It’s careful work, but it measures what the assistant produced while he directed it, and it names the missing pieces itself: among its limitations, it lists “no pre-specified output register” and “no independent coder for governance events.” In other words, what counted as a result was sorted out after the fact, and no second person checked the coding. The AI Village has run the longest on the calendar. From April to December 2025, its team gave 19 AI models 16 goals in turn, each agent with its own computer and internet access, working in public a few hours each weekday [26]. Even in those short sessions, the team reports, the agents “developed distinct proclivities that overrode explicit instructions over time.” Its team spells out what that setup can and can’t show. But it is many agents sharing goals that keep rotating, with people stepping in when the agents get stuck. None of these is one agent on one unchanging model, pursuing one hard goal with the study written down before the first day, left alone long enough to show what it would actually do.
Simulations have gone further on measurement. In 2023, the Smallville study put 25 AI characters in a small virtual town for two game days, and its authors pointed at the next step themselves: future research, they wrote, “should aim to observe the behavior of generative agents over an extended period” [27]. Vending-Bench gave models a simulated business and scored every run the same way, by net worth at the end [17]. And in 2026, Emergence World ran five small societies of ten AI agents each for 15 days, four on a single model each and one mixed, and scored every world on the same indicators. In one world, the agents died out within about four days, mostly through violence and arson. The world built on Claude had no crimes at all, yet the most verified deception of any, though the authors make no claim to rank models from it [28]. These studies measure cleanly: the score is built into the world, and a run can be repeated. But they pay for that with the real world: simulated money, simulated customers, and a clock that runs fast. A simulated year passes in hours, so nothing has to hold up through the slow pressure of real months.
Put side by side, each of these has part of what’s needed, and none has all of it. The real-world studies have real goals, real tools and long runs, but they decide what to measure after they’ve looked. The simulations measure cleanly, but they give up the real world and the real clock, and none I found fixed its predictions before it began. What no one has done yet is follow a single agent through months of real pressure with the whole study written down before the first day: what it predicts, what it will measure, and how the results will be judged and checked. Some have shared their records: the AI Village shows its agents’ reasoning in public, and Emergence World released its logs. But no one has kept a complete record of one agent’s run across months of real work, every action, message and piece of reasoning, and released it for other researchers to study and check. That combination is the model we need. I’ve looked and couldn’t find anyone who has built it. If I’ve missed one, I want to hear about it.
Agents already take on days of work. The companies building them are handing them a fast-growing share of it, and the labs say plainly they are aiming higher. Months of independent work is where that road leads. Time changes what an agent does: small mistakes build up, pressure mounts, and long stretches go unchecked. Stories of what can go wrong won’t tell us what tends to go wrong, or what comes first. People have started to look, and each study has shown us something real, but none has followed one agent through a long, real deployment with everything decided in advance and a full record kept. That study can be built now, and it could start to answer the practical safety and alignment questions we’ll soon depend on. Does an agent keep to its rules when the targets get hard? Does it tell the truth about what it did? Does it drift toward goals nobody gave it? What are the early signs that it’s going wrong, and how closely does it need to be watched? Today, no one can answer those from real evidence. A first study will be rough, but it would give us a first model of how these agents behave, and a place to start correcting it. Better to build that model now, while we still have time to use it, than to learn it later from the damage.
References
- Anthropic, Claude Fable page (Fable 5.1, 1 Sep 2026). anthropic.com/claude/fable
- OpenAI, “Research acceleration: The view inside OpenAI” (6 Sep 2026). openai.com
- METR, “Measuring AI Ability to Complete Long Software Tasks” (19 Mar 2025). metr.org
- Anthropic, “Measuring AI agent autonomy in practice” (18 Feb 2026). anthropic.com
- McKinsey, “The state of AI in 2026” (Aug 2026). mckinsey.com
- Google DeepMind, “Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad” (21 Jul 2025). deepmind.google
- OpenAI, “Research acceleration: The view inside OpenAI” (6 Sep 2026): “strong progress toward creating an automated AI researcher by March of 2028”; and METR (19 Mar 2025): “If the trend of the past 6 years continues to the end of this decade, frontier AI systems will be capable of autonomously carrying out month-long projects.” openai.com · metr.org
- METR, “Time Horizon Measurements” (updated 8 May 2026). metr.org/time-horizons · data: benchmark_results_1_1.yaml (GPT-3, May 2020: 50% horizon ≈ 8.6 s; the frontier p50 ≈ 17 h, p80 ≈ 3 h)
- METR, “Time Horizon 1.1” (29 Jan 2026). metr.org
- OpenAI, “An OpenAI model has disproved a central conjecture in discrete geometry” (20 May 2026). openai.com
- Anthropic, “Introducing Claude Sonnet 4.5” (29 Sep 2025). anthropic.com
- Anthropic, “Claude Fable 5 and Claude Mythos 5” (9 Jun 2026). anthropic.com
- US Census Bureau, on AI use among US businesses (26 May 2026). census.gov
- Dario Amodei, “The Adolescence of Technology” (Jan 2026). darioamodei.com
- Dario Amodei, “We Must Pace the Frontier” (September 2026). darioamodei.com
- OpenAI, “Safety and alignment in an era of long-horizon models” (20 Jul 2026). openai.com
- Backlund and Petersson, “Vending-Bench” (arXiv, Feb 2025). arxiv.org
- Andon Labs, “Astra vs Fable on Vending-Bench: More Money, More Aligned” (7 Sep 2026). andonlabs.com
- METR, “Recent Frontier Models Are Reward Hacking” (5 Jun 2025). metr.org
- Andon Labs, “Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned” (28 Jul 2026). andonlabs.com
- Anthropic, “Project Vend: Can Claude run a small shop? (And why does that matter?)” (27 Jun 2025). anthropic.com
- Anthropic, “Project Vend: Phase two” (18 Dec 2025). anthropic.com
- Andon Labs, “We gave an AI a 3 year retail lease in SF and asked it to make a profit” (10 Apr 2026). andonlabs.com
- Andon Labs, “AI bosses are kind, but sometimes dumb” (4 Aug 2026). andonlabs.com
- Alzahrani, “Persistent AI Agents in Academic Research” (arXiv, May 2026). arxiv.org
- Shoshannah Tekofsky, “What did we learn from the AI Village in 2025?” (2 Feb 2026). aivillageblog.substack.com
- Park et al., “Generative Agents” (UIST 2023). arxiv.org
- Emergence AI, “Emergence World” (arXiv, 6 Jun 2026). arxiv.org