---
title: "We lost the war on quality"
slug: we-lost-the-war-on-quality
url: https://listedarticles.com/articles/we-lost-the-war-on-quality
canonical_url: https://jssfr.de/2026-10-04-we-lost-the-war-on-quality.html
content_type: opinion
language: en
published_at: 2026-10-04T00:00:00.000Z
updated_at: 2026-10-04T11:14:29.997Z
author: "Jonas Schäfer"
author_url: https://jssfr.de/
authored_by: human
publisher: "jssfr.de"
publisher_url: https://jssfr.de/
topics: ["AI", "LLMs", "Software Engineering", "Opinion", "Programming"]
license: all-rights-reserved
word_count: 2037
reading_minutes: 9
citation: "Jonas Schäfer, jssfr.de. \"We lost the war on quality.\" 4 Oct 2026. https://jssfr.de/2026-10-04-we-lost-the-war-on-quality.html (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# We lost the war on quality

> Jonas Schäfer argues that LLM quality has improved enough that “AI output is bad” no longer works as a blanket objection—quality arguments lost, so the harder fights are about labor, ecology, control, and what we still choose to do ourselves.

# We lost the war on quality

**Content notice and disclaimer:** This post is about AI. When I say AI, I
mean the GenAI/LLM hype which has been going on since late 2022 in the broader
public.

In this post, I will *not* further mention all the bad things we all know about
AI<sup>1</sup>. The reason is not that I don't think it's important. The reason
is that I think that we **all** know about this at this point. Some decide to
ignore it. I am not ignoring it and I am deeply concerned about it. It's
important, but not the subject matter of this post. I promise it'll make sense
if you make it to the end.

Also, I am only talking about AI in the context of software engineering and computer infrastructure operations.

No AI has been used in the writing of this blog post, except the parts where I describe my interactions with an AI. There, obviously, I used AI in some way. But the writing is mine.

it's good and I hate it so much

That's me, writing to a friend about the experience of having Anthropic's Fable model diagnose my broken NextCloud setup. That incident has been the spark which made me write this post.

It's important that everyone understands the implications of the developments of the past year. Tech has always been fast, but compared to everything else, AI seems to be moving at the speed of light.

### Last year in AI (from my perspective)

#### The Vulnocalypse

About a year ago, my social media feed was full of people complaining about
bullshit issues and security bug reports which cluttered their inboxes. People
(we used to call this kind of clientele script kiddies, a term I haven't heard
in quite a while) used bad AI tools in incompetent ways, either in the hope of
scoring a CVE number<sup>2</sup> or in the hope of being useful to the community.

Either way, they did so without the necessary knowledge, time or willingness to triage the output of the AI tools, which eventually culminated in at least one famous open source project to close their bug bounty in January 2026.

This felt amazing. It seemed as if whenever someone tried to throw AI at a new field, it failed spectacularly and that was it. It felt as if there was a decent chance that, after half a decade of public development and hundreds of billions of dollars having been poured into this technology, investors might finally realize the error of their ways and the bubble might burst sooner rather than later. I was cheering for every failure of AI, in the hopes that we might yet save the planet from its impact.

But then, in March 2026, something shifted. Multiple projects, most notably the Linux kernel developers, reported that the quality of AI generated reports skyrocketed, basically overnight. Instead of hallucinated bogus problems, they were being confronted with well-researched issues, working reproducers and even patches.

What followed was the onset of what is now being called the Vulnocalypse:
a storm of *high-quality* AI-assisted security vulnerability findings of
varying degrees of severity.

Linux has seen several local privilege escalation (LPE) issues<sup>3</sup>
(the well-marketed copy.fail and its siblings) and at
least three hypervisor escapes<sup>4</sup>
(CVE-2026-53359,
CVE-2026-64513 and
CVE-2026-89775). nginx saw
several remote code execution (RCE) bugs (e.g.
CVE-2026-42945 and
CVE-2026-42530), at least two
of which coincided with a Linux LPE on the same weekend.
OpenStack has seen more security advisories
this year than in the previous eight years combined (at least two of which were
privilege escalation issues,
OSSA-2026-015 and
OSSA-2026-037).

To make matters worse, an issue found by one AI tool (user) is likely to be found by another, which has led to another round of questioning the practice of responsible disclosure and security issue embargos.

Vendors and operators<sup>5</sup> alike are scrambling to close the
issues before they are being exploited. I won't go into detail about the
HuggingFace/OpenAI incident,
because we cannot know from the outside how much of that is
marketing<sup>6</sup>, but I don't like to think about what would happen if an
adversary combined an agent swarm with a vulnerability search across major
open source projects.

No matter what you otherwise think of AI, these security issues which have
been found in *human written* code are real. After the initial marketing hype
blew past, more and more issues have been found and quietly fixed in the
recent months and it's not over yet. As jyn put
it: "we have a year to fix security everywhere".

#### Personal Experience

So far, I talked about things which I observed from the outside. This is because up until very recently, I avoided contact with LLMs like the plague (you know, because of all the bad things we all know about them).

I work at a company doing things with computers, so of course there have been run-ins with AI at my dayjob, too. Due to the usual NDAs which go alongside working in technology, I cannot talk much about this publicly. So far, I have been spared of having to use LLMs for work. My employer puts a great deal of focus on digital sovereignty and sustainability, which makes the use of LLMs problematic, to say the least.

Nontheless, one thing I can talk in more detail about is what I opened this article with: my NextCloud problem.

A couple of weeks ago, one of my users complained that their Windows desktop
client couldn't create share links anymore. I thought it was a permission
issue, but no amount of tweaking fixed it. I then assumed a version
mismatch<sup>7</sup>, but even after upgrading, the problem
persisted.

The final escalation level was then reinstalling the NextCloud client. That surfaced a different problem: The enrollment procedure did not work. The login worked, but the step after that, where you pick the sync folder, that got stuck when pressing "done".

I checked the nextcloud logs and found a 403 Forbidden in response to an
`MKCOL` HTTP verb. When you google for that, you find one or two forum posts,
but none which applied to my situation.

It was a lazy friday afternoon and so I decided to finally learn what all this AI fuzz was about. Of course, I did the worst thing (in terms of quality) possible: I interacted with the AI-generated reply from Google.

I won't bore you with the details. We went a bit back-and-forth<sup>8</sup>, it
attempted to get me to seriously fuck up my NextCloud setup twice before I gave
up.

Then a friend "offered" to let me use Anthropic's Fable model to debug the
issue. After another back-and-forth, it pin-pointed the issue to my
hand-written Apache config not allowlisting the `ocs/v1.php` endpoint
(only `v2.php` was allowed).

I had completely forgotten about my PHP allowlist. There is no chance that I would've spotted this on my own. The most likely way this would've gotten fixed is me giving up and eventually reconfiguring everything from scratch.

This illustrates two things: First, the difference in quality even between
proprietary models can be *extreme*. The level of bullshit generated by the
model Google uses in their search frontend is absolutely not comparable to the
quality of Fable.

Second, the experience with Fable was one more similar to interacting with a
well-functioning open source community chat<sup>9</sup>. On a Sunday when the
weather outside is bad and everyone has nothing better to do than hanging out
there and debugging your problems<sup>10</sup>.

Interacting with Fable felt like interacting with a NextCloud expert. The
initial hypotheses were not correct, but they were the correct hypotheses to
have based on tha data I provided. As I provided more data to refute its
proposals it formed new ones. And for each hypothesis, it provided an
obviously-harmless command to verify it. For example, it provided a fully
correct `grep` command line to find the 403 replies sent by apache (and thus
never seen by NextCloud) in the log files (including my customized log file
names which it got from the apache config snippet).

### Takeaway

In the meantime (as usual for blog posts, this one sat on my disk for several weeks before it got published), I had more situations where I was able to see what the current proprietary frontier models can do (even though I have not interacted with one directly, yet).

I know that some people will still say that AI does not work or generally produces bad results. From my own experiments at this point, this is certainly true for small models. It may be true for larger open weight models except the most recent top-of-the-line ones (I don't have the hardware to try these). It is certainly untrue for some proprietary frontier models.

This means one thing first and foremost: **If** you want to argue against AI
(and as I mentioned in the beginning, there's lots of good reasons to), arguing
based on quality is *not* going to work. To whomever you are arguing with, if
that person has had contact with a frontier model, your argument will seem
absurd.

There are still many good reasons not to use AI and not to endorse or support products made in part or exclusively with AI. However, if they have been made with current frontier models, quality is very likely not one of these reasons.

We can use all the ethical arguments, but it's of absolute importance
that everyone who opposes AI realizes this:
**We have lost the war on quality.**

Comments are welcome in the Fediverse thread for this post.

1. 
Electricity use, water use, mass-scraping breaking many services and costing the volunteers who run them a lot of nerves and money, wide abuse of published materials for training, skill loss, LLMs being run by fascistoid entities, the (ab-)use of clickworkers, etc. etc. People are keeping lists about this stuff, in case you truly need to read more about this (but when things argue about quality, check the date). ↩
2. 
~~Curriculum Vitae Enhancer~~ Common Vulnerabilities
  and Exposures database. ↩
3. 
Though some people argue that Linux LPEs are *so* common that
  this was basically just overblown marketing. ↩
4. 
If you don't know what that means: It's *extremely* bad for cloud providers. It allows a customer to break out into the
  underlying infrastructure, compromising the cloud provider's
  infrastructure itself and/or other customer's VMs.It's not fully clear whether all three of these escapes have been found using AI tools. ↩
5. 
And I'd like you to realize that this isn't just hitting evil companies which are driven by capitalism to minimize spending on IT security. This is also hitting people like you and me, self-hosting their email or chat or whatever, to keep their data private and to fight for a decentralized and open internet. ↩
6. 
Though my personal opinion is, that even *if* this is 80%
  marketing, it's bad enough. In addition, the
  talk at BlackHat USA
  makes me seriously doubt the "this is just marketing" angle. ↩
7. 
This is a good place for a rant about how awful the update experience is with the NextCloud windows desktop client. It wants to update *so* often, and whenever it does it needs a reboot due to its
  tight integration with Windows. Gosh, that's**so** annoying. ↩
8. 
I'll now start to anthropomorphize the LLMs. That is simply because I lack the vocabulary to talk about this, except in very convoluted ways (e.g. "generated text which said" instead of "said"). ↩
9. 
Fun fact: I had strace'd the NextCloud processes to understand whether the 403 was caused by an EPERM from some syscall and told Fable that. Fable did not trust me to be able to run strace correctly ("strace was probably watching the wrong process anyway."). So much for sycophancy. Fable seems to be rather scathing indeed, from what I've heard so far. When it is factually correct, you can't make it back off. ↩
10. 
With the difference that the chance that such a community would've started to yell at me for not mentioning my custom apache config in the first place is *much* greater than zero. ↩
