---
title: "The Hamster Paradox"
slug: the-hamster-paradox
url: https://listedarticles.com/articles/the-hamster-paradox
canonical_url: https://pub.towardsai.net/the-hamster-paradox-36dd0d950568
content_type: essay
language: en
published_at: 2026-10-04T00:01:02.151Z
updated_at: 2026-10-04T05:15:06.172Z
author: "Younss"
authored_by: human
publisher: "Towards AI"
publisher_url: https://pub.towardsai.net
topics: ["ai", "productivity", "engineering-leadership", "opinion"]
license: all-rights-reserved
word_count: 2310
reading_minutes: 10
citation: "Younss, Towards AI. \"The Hamster Paradox.\" 4 Oct 2026. https://pub.towardsai.net/the-hamster-paradox-36dd0d950568 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# The Hamster Paradox

> Younss argues that AI can raise individual task throughput without raising organizational throughput—like a hamster spinning faster without traveling—because queues, handoffs, and decision bottlenecks still dominate.

The Hamster Paradox
AI speeds up work. There is no guarantee it speeds up the organization
The views expressed here are my own and do not represent those of my employer.
Imagine a hamster in its wheel. It runs faster and faster, 30% faster, then 50%. If we measure its performance in revolutions per minute, every indicator is green. Yet it has not travelled a single metre.
This image often comes to mind when I hear about AI productivity gains. I do not doubt those gains, which are increasingly well documented. What gives me pause is what we mean by productivity, and especially where we choose to measure it. Generative AI genuinely increases productivity across many activities. But when one activity becomes much faster without the rest of the system keeping pace, the gain does not disappear. It simply runs into the next constraint.
Operations management, one of my favourite fields, knows this problem very well. Goldratt formalized it in the theory of constraints, which holds that improving one part of a system does not necessarily improve the system as a whole. Widening one stretch of highway rarely ends a traffic jam. It usually just moves the jam to the next interchange. Generative AI gives this longstanding idea a new dimension, however, because it brings the cost and time required to produce knowledge work close to zero.
We therefore need to look at what AI allows us to produce faster, but also at what happens around that production and after it.
The gains are real
Brynjolfsson, Li and Raymond studied 5,172 customer service agents. Access to generative AI increased the number of issues resolved per hour by 15% on average, with even larger gains among less experienced agents.
The experiment that Dell’Acqua and colleagues ran with 758 BCG consultants reached a similar conclusion, with an important qualification. On tasks inside GPT-4’s capability frontier, AI-assisted consultants completed 12.2% more tasks, took 25.1% less time and produced higher-quality work.
These gains are real, and in many settings AI makes us more productive. The open question is at what level.
It depends on where we draw the boundary
Early in my career, I was a developer at several large European organizations. I left them long ago, but some of the code I wrote may still be running. Had someone wanted to measure my productivity at the time, they would probably have looked at what I delivered over a given period. That measure would have made sense, but it would have been incomplete.
That code then had to live inside an organization. Other people would have to understand it, test it, integrate it, operate it, modify it and maintain it. Some of my decisions would make their work easier years later, and others would cost them time (I gave an example in my previous essay, When Anyone Can Build Software, Who Decides What Not to Build?).
That experience makes me cautious about some of today’s measures of AI productivity. A developer who produces more code with AI Code Assistant may genuinely be more productive. But once that code leaves the developer’s workstation, it enters a system where it must be reviewed, tested, secured, integrated, deployed, operated and maintained. And before any of that, it should solve a problem worth solving.
If I measure only the developer, I am measuring an activity. If I widen the boundary to the development process, I begin to measure throughput. If I extend it to what the software changes for the customer or the organization, I am measuring an outcome. All three perspectives are useful. The problem comes when we use the first as evidence of the third.
When production is no longer the constraint
For a long time, producing knowledge work was expensive. An analysis took time, and exploring several scenarios took even more. Writing code, documenting an architecture, preparing a recommendation or building a presentation occupied people for hours, sometimes days. That scarcity forced us to choose.
With today’s AI models, part of that constraint disappears. I can request another analysis, three more scenarios, a new architecture or another implementation at a very low marginal cost (the economics of tokens).
This raises a question we ask less often. What happens when our capacity to produce grows much faster than our capacity to judge and validate what we produce?
Consider a deliberately simplified chain:
Problem > Analysis > Creation > Validation > Decision > Integration > Value
AI can now accelerate several of these stages considerably, but nothing guarantees that every capacity in the system will improve at the same pace. A committee does not make five times as many decisions because it receives five times as many analyses. An architecture team cannot evaluate five times as many solutions. An organization does not absorb five times as much change because its developers can generate it.
At some point the constraint shifts, and that is probably where we need to look.
A question for you, the reader: have you ever wondered whether it would have been simpler to build a presentation yourself than to ask AI for one and iterate on it many times?
Judgement throughput
I use the term judgement throughput for the rate at which an organization turns what it produces into something it understands well enough to validate, decide how to use and act on.
If generation grows faster than this throughput, the surplus does not automatically become value. It accumulates as work in progress. Code waits for review, analyses pile up, several recommendations coexist for the same problem, decisions are deferred, and information sits somewhere without changing anything the organization does. Sometimes rework follows.
From this I draw a hypothesis, which I do not present as an established finding:
When generation capacity persistently grows faster than judgement throughput, local productivity gains should translate less and less proportionally into value for the system.
It can lead to synaptic laziness. Also, it would be disproved if organizations that measure their end-to-end throughput saw local gains consistently turn into outcomes without changing how they validate, decide or integrate. Several recent findings justify, at a minimum, taking it seriously.
Consider DORA’s research. In 2024, its study associated a 25% increase in AI adoption with an estimated 1.5% decrease in delivery throughput and a 7.2% decrease in stability. In 2025, the relationship with throughput turned positive, while stability continued to deteriorate.
I do not see this as proof that AI slows down or speeds up development. I read it instead as a sign that the effect depends on the system into which the technology is introduced. DORA now describes AI as an amplifier. On that view, AI amplifies the strengths of a well-run organization as much as the weaknesses of one with fragile processes.
METR offers another signal. In its 2025 randomized trial, 16 experienced open-source developers working on 246 real tasks took 19% longer on average when they had access to the AI tools being studied. Afterwards, they still estimated that those tools had made them about 20% faster.
METR has since said that newer tools probably speed developers up more, while noting that its 2026 data cannot measure the size of that effect reliably. What I take from the study is therefore less the 19% figure than the possible gap between what we feel we are accelerating and what the system actually produces.
What about quality?
Our productivity dashboards also do a poor job of capturing quality.
The BCG experiment contains a finding that strikes me as even more important than the speed gains. On problems well suited to GPT-4’s capabilities, AI improved the consultants’ performance. On a task deliberately placed outside that frontier, those using AI were 19 percentage points less likely to reach the correct solution.
Get Younss’s stories in your inbox
Join Medium for free to get updates from this writer.
The technology and the population were the same. The outcome was not.
It therefore becomes hard to speak of a single “AI productivity gain” that applies uniformly across an organization. The outcome depends on the task, the context, the user’s expertise, the model’s capabilities and the system in which the task sits. Applying an average rate to thousands of employees tells me much less than understanding exactly where those gains appear and what happens to them afterwards.
The hamster paradox
Let us return to the hamster. We can certainly measure its revolutions per minute, and if AI lets it complete 50% more, we have indeed improved that indicator. But if our goal was to cover a distance, something is missing from our measurement.
This is what I call the hamster paradox:
an organization can substantially improve the measured productivity of certain activities without a proportional improvement in system throughput or outcomes.
The hamster is therefore not a criticism of AI. The problem becomes visible precisely because AI works. It raises the capacity of an activity enough to reveal what was limiting it downstream.
For a leader, this changes the question. I would distinguish three levels: Activity > Throughput > Outcome
Activity tells us what we produced. Throughput tells us what moved through the system, and outcome what changed as a result.
Knowing that 10,000 employees save two hours a week through AI is useful, but I would then want to know what became of those 20,000 hours. Did we shorten the time it takes to make a decision? Deliver faster? Improve quality? Reduce a risk? Create more value for our customers? Or did we simply use the freed-up time to produce more or to do team bu?
After adoption
The first phase of enterprise AI naturally focused on access: licences, copilots, use cases, training and adoption. The next will force us to examine the organization as a system in which certain capacities have just changed radically.
Before asking where to add AI, I would therefore start with another question. What constraint currently limits our ability to create value? I would then ask what happens to that constraint if we increase the capacity of the preceding activity fivefold.
The answer will vary from one system to another. Sometimes creation really is the constraint, and AI can transform throughput. Elsewhere, the constraint will be validation, decision-making, integration, architecture or governance. There are probably also activities that create too little value to be worth accelerating, and the right decision would then be to eliminate them.
That is why I increasingly see AI transformation as a question of system design rather than a matter of technology adoption.
Measuring distance
At the activity level, the classic definition of productivity remains useful: output relative to the resources consumed. To discuss the organization as a whole, the boundary has to widen. I would represent it this way:
System productivity = (Value delivered) / (Total resources consumed by the system)
It is a measurement discipline, one that stops us from ending the count the moment AI finishes its response. What has been generated may still need to be verified, corrected, coordinated, integrated, governed, operated and maintained, and those resources are part of the system too.
The question is less appealing than a number of hours saved, but far more useful. How much additional value have we moved through the system?
What deserves to be produced
The more I observe these transformations, the less their main challenge seems technological.
For a long time, production accounted for a large part of the constraint on knowledge work. Analyzing, writing, coding or exploring several possibilities took time. AI reduces that constraint, so we will produce more.
That does not mean we will choose better.
The scarce resource may be shifting towards our ability to understand the problem, tell a plausible answer from reliable knowledge, weigh competing solutions and decide what deserves to enter the system. When production becomes abundant, judgement becomes more valuable.
The hamster simply runs faster, as it is asked to. It is up to us not to mistake the revolution counter for the distance travelled.
This Article’s Paradox
As with all my recent articles, I use AI to structure my ideas and write my blog posts. Here’s how I do it. I speak to the LLM on my phone during my commute and explain everything about the article, in French: the structure, where I want to go, the main idea, the challenges, the storytelling. Then I pass it to my systems-thinking agents to validate my hypothesis and point of view. It’s a way of challenging myself again. Then I iterate. Once I have the fil conducteur, the common thread, I pass it to another model to write the article, then to yet another model to challenge it again, restructure it, translate it from French to English and write some sections. These models are Claude, Gemini and ChatGPT. What took time wasn’t writing the article. It was deciding: the narrative that sounds like me, and which data and references are true and legitimate.
Bref, generation is faster. My judgement and thinking didn’t move.
I’m the hamster in this story too ;)
Thank you for reading until the end.
References
- Becker, J., Rush, N., Barnes, E. and Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR (https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/).
- Brynjolfsson, E., Li, D. and Raymond, L. (2025). “Generative AI at Work.” The Quarterly Journal of Economics, 140(2), 889–942. DOI: 10.1093/qje/qjae044 (https://academic.oup.com/qje/article/140/2/889/7990658).
- Dell’Acqua, F. et al. (2026). “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.” Organization Science, 37(2).
- DORA (2024). Accelerate State of DevOps Report. Google Cloud (https://dora.dev/research/2024/dora-report/).
- DORA (2025). State of AI-assisted Software Development. Google Cloud (https://dora.dev/research/2025/dora-report/).
- Goldratt, E. M. and Cox, J. (1984). The Goal. North River Press.
- METR (2026). We Are Changing Our Developer Productivity Experiment Design. METR (https://metr.org/blog/2026-02-24-uplift-update/).
