We lost the war on quality
Content notice and disclaimer: This post is about AI. When I say AI, I mean the GenAI/LLM hype which has been going on since late 2022 in the broader public.
In this post, I will not further mention all the bad things we all know about AI<sup>1</sup>. The reason is not that I don't think it's important. The reason is that I think that we all know about this at this point. Some decide to ignore it. I am not ignoring it and I am deeply concerned about it. It's important, but not the subject matter of this post. I promise it'll make sense if you make it to the end.
Also, I am only talking about AI in the context of software engineering and computer infrastructure operations.
No AI has been used in the writing of this blog post, except the parts where I describe my interactions with an AI. There, obviously, I used AI in some way. But the writing is mine.
it's good and I hate it so much
That's me, writing to a friend about the experience of having Anthropic's Fable model diagnose my broken NextCloud setup. That incident has been the spark which made me write this post.
It's important that everyone understands the implications of the developments of the past year. Tech has always been fast, but compared to everything else, AI seems to be moving at the speed of light.
Last year in AI (from my perspective)
The Vulnocalypse
About a year ago, my social media feed was full of people complaining about bullshit issues and security bug reports which cluttered their inboxes. People (we used to call this kind of clientele script kiddies, a term I haven't heard in quite a while) used bad AI tools in incompetent ways, either in the hope of scoring a CVE number<sup>2</sup> or in the hope of being useful to the community.
Either way, they did so without the necessary knowledge, time or willingness to triage the output of the AI tools, which eventually culminated in at least one famous open source project to close their bug bounty in January 2026.
This felt amazing. It seemed as if whenever someone tried to throw AI at a new field, it failed spectacularly and that was it. It felt as if there was a decent chance that, after half a decade of public development and hundreds of billions of dollars having been poured into this technology, investors might finally realize the error of their ways and the bubble might burst sooner rather than later. I was cheering for every failure of AI, in the hopes that we might yet save the planet from its impact.
But then, in March 2026, something shifted. Multiple projects, most notably the Linux kernel developers, reported that the quality of AI generated reports skyrocketed, basically overnight. Instead of hallucinated bogus problems, they were being confronted with well-researched issues, working reproducers and even patches.
What followed was the onset of what is now being called the Vulnocalypse: a storm of high-quality AI-assisted security vulnerability findings of varying degrees of severity.
Linux has seen several local privilege escalation (LPE) issues<sup>3</sup> (the well-marketed copy.fail and its siblings) and at least three hypervisor escapes<sup>4</sup> (CVE-2026-53359, CVE-2026-64513 and CVE-2026-89775). nginx saw several remote code execution (RCE) bugs (e.g. CVE-2026-42945 and CVE-2026-42530), at least two of which coincided with a Linux LPE on the same weekend. OpenStack has seen more security advisories this year than in the previous eight years combined (at least two of which were privilege escalation issues, OSSA-2026-015 and OSSA-2026-037).
To make matters worse, an issue found by one AI tool (user) is likely to be found by another, which has led to another round of questioning the practice of responsible disclosure and security issue embargos.
Vendors and operators<sup>5</sup> alike are scrambling to close the issues before they are being exploited. I won't go into detail about the HuggingFace/OpenAI incident, because we cannot know from the outside how much of that is marketing<sup>6</sup>, but I don't like to think about what would happen if an adversary combined an agent swarm with a vulnerability search across major open source projects.
No matter what you otherwise think of AI, these security issues which have been found in human written code are real. After the initial marketing hype blew past, more and more issues have been found and quietly fixed in the recent months and it's not over yet. As jyn put it: "we have a year to fix security everywhere".
Personal Experience
So far, I talked about things which I observed from the outside. This is because up until very recently, I avoided contact with LLMs like the plague (you know, because of all the bad things we all know about them).
I work at a company doing things with computers, so of course there have been run-ins with AI at my dayjob, too. Due to the usual NDAs which go alongside working in technology, I cannot talk much about this publicly. So far, I have been spared of having to use LLMs for work. My employer puts a great deal of focus on digital sovereignty and sustainability, which makes the use of LLMs problematic, to say the least.
Nontheless, one thing I can talk in more detail about is what I opened this article with: my NextCloud problem.
A couple of weeks ago, one of my users complained that their Windows desktop client couldn't create share links anymore. I thought it was a permission issue, but no amount of tweaking fixed it. I then assumed a version mismatch<sup>7</sup>, but even after upgrading, the problem persisted.
The final escalation level was then reinstalling the NextCloud client. That surfaced a different problem: The enrollment procedure did not work. The login worked, but the step after that, where you pick the sync folder, that got stuck when pressing "done".
I checked the nextcloud logs and found a 403 Forbidden in response to an
MKCOL HTTP verb. When you google for that, you find one or two forum posts,
but none which applied to my situation.
It was a lazy friday afternoon and so I decided to finally learn what all this AI fuzz was about. Of course, I did the worst thing (in terms of quality) possible: I interacted with the AI-generated reply from Google.
I won't bore you with the details. We went a bit back-and-forth<sup>8</sup>, it attempted to get me to seriously fuck up my NextCloud setup twice before I gave up.
Then a friend "offered" to let me use Anthropic's Fable model to debug the
issue. After another back-and-forth, it pin-pointed the issue to my
hand-written Apache config not allowlisting the ocs/v1.php endpoint
(only v2.php was allowed).
I had completely forgotten about my PHP allowlist. There is no chance that I would've spotted this on my own. The most likely way this would've gotten fixed is me giving up and eventually reconfiguring everything from scratch.
This illustrates two things: First, the difference in quality even between proprietary models can be extreme. The level of bullshit generated by the model Google uses in their search frontend is absolutely not comparable to the quality of Fable.
Second, the experience with Fable was one more similar to interacting with a well-functioning open source community chat<sup>9</sup>. On a Sunday when the weather outside is bad and everyone has nothing better to do than hanging out there and debugging your problems<sup>10</sup>.
Interacting with Fable felt like interacting with a NextCloud expert. The
initial hypotheses were not correct, but they were the correct hypotheses to
have based on tha data I provided. As I provided more data to refute its
proposals it formed new ones. And for each hypothesis, it provided an
obviously-harmless command to verify it. For example, it provided a fully
correct grep command line to find the 403 replies sent by apache (and thus
never seen by NextCloud) in the log files (including my customized log file
names which it got from the apache config snippet).
Takeaway
In the meantime (as usual for blog posts, this one sat on my disk for several weeks before it got published), I had more situations where I was able to see what the current proprietary frontier models can do (even though I have not interacted with one directly, yet).
I know that some people will still say that AI does not work or generally produces bad results. From my own experiments at this point, this is certainly true for small models. It may be true for larger open weight models except the most recent top-of-the-line ones (I don't have the hardware to try these). It is certainly untrue for some proprietary frontier models.
This means one thing first and foremost: If you want to argue against AI (and as I mentioned in the beginning, there's lots of good reasons to), arguing based on quality is not going to work. To whomever you are arguing with, if that person has had contact with a frontier model, your argument will seem absurd.
There are still many good reasons not to use AI and not to endorse or support products made in part or exclusively with AI. However, if they have been made with current frontier models, quality is very likely not one of these reasons.
We can use all the ethical arguments, but it's of absolute importance that everyone who opposes AI realizes this: We have lost the war on quality.
Comments are welcome in the Fediverse thread for this post.
Electricity use, water use, mass-scraping breaking many services and costing the volunteers who run them a lot of nerves and money, wide abuse of published materials for training, skill loss, LLMs being run by fascistoid entities, the (ab-)use of clickworkers, etc. etc. People are keeping lists about this stuff, in case you truly need to read more about this (but when things argue about quality, check the date). ↩
Curriculum Vitae Enhancer Common Vulnerabilities
and Exposures database. ↩
Though some people argue that Linux LPEs are so common that this was basically just overblown marketing. ↩
If you don't know what that means: It's extremely bad for cloud providers. It allows a customer to break out into the underlying infrastructure, compromising the cloud provider's infrastructure itself and/or other customer's VMs.It's not fully clear whether all three of these escapes have been found using AI tools. ↩
And I'd like you to realize that this isn't just hitting evil companies which are driven by capitalism to minimize spending on IT security. This is also hitting people like you and me, self-hosting their email or chat or whatever, to keep their data private and to fight for a decentralized and open internet. ↩
Though my personal opinion is, that even if this is 80% marketing, it's bad enough. In addition, the talk at BlackHat USA makes me seriously doubt the "this is just marketing" angle. ↩
This is a good place for a rant about how awful the update experience is with the NextCloud windows desktop client. It wants to update so often, and whenever it does it needs a reboot due to its tight integration with Windows. Gosh, that'sso annoying. ↩
I'll now start to anthropomorphize the LLMs. That is simply because I lack the vocabulary to talk about this, except in very convoluted ways (e.g. "generated text which said" instead of "said"). ↩
Fun fact: I had strace'd the NextCloud processes to understand whether the 403 was caused by an EPERM from some syscall and told Fable that. Fable did not trust me to be able to run strace correctly ("strace was probably watching the wrong process anyway."). So much for sycophancy. Fable seems to be rather scathing indeed, from what I've heard so far. When it is factually correct, you can't make it back off. ↩
With the difference that the chance that such a community would've started to yell at me for not mentioning my custom apache config in the first place is much greater than zero. ↩