One chef's-knife brand gets 31% of its "what should I buy" mentions from 5% of the accounts, four times what chance predicts. So I pulled those accounts' full Reddit histories.
This article was drafted with AI from my outline and the run logs, then edited. The buying-thread numbers can be recomputed from the published data with one script; the account-history comparison cannot, because it rests on usernames I will not publish.
Last year I fine-tuned a small named-entity model, GLiNER, to pull brands, models and steels out of knife comments. That model runs over every comment the New Knife Day scraper collects from six subreddits: r/knives, r/knifeclub, r/chefknives, r/japaneseknives, r/FixedBladeEdc and r/KnifeSteels.
So for every comment I already have who wrote it, which brands it names, and whether the thread it sits in is someone asking what to buy.
What astroturfing would look like in the data
Services like REDCmts sell Reddit comments from "real, aged accounts." Taking those sales pages as the description of the product, a paid campaign should show up as:
- A small tail of accounts writing a disproportionate share of brand mentions in buying threads
- Those accounts naming one brand almost every time
- Thin accounts: few comments, low scores
- Young accounts, or histories that are hidden
- Accounts that post mostly in knife subreddits
- Links to a store or affiliate page
Every one of those is also what a devoted fan would produce. Public Reddit data can show concentration, and cannot show why.
The corpus
A refresh pass went back to 3,607 posts older than 48 hours and refetched comment trees, taking the corpus from 21,673 comments to 51,129.
| Count |
|---|
| Posts | 6,675 |
| Comments after refresh | 51,129 |
| Authors with 10 or more comments | 987 |
| Brand mentions in buying threads (those authors) | 1,471 |
The tail is the top 5% of the 987 authors on brand-heaviness—49 accounts. Chance is estimated by 1,000 random reassignments of author names across buying-thread mentions.
What the data supports
If the tail shows up only because it posts a lot, random reassignment says it should write about 7.9% of brand mentions (typically 6.3%–10.1%). It wrote 11.3%. Only 2 of 1,000 reassignments reached that.
Where the extra share lands: r/chefknives and r/knifeclub sit above chance; r/knives (the largest) is within half a point of chance. Three brands get far more of their buying advice from the tail than chance; one large brand gets slightly less; three get none.
On corpus data, the tail accounts are thin (median 12 comments, median score 1) and loyal (two thirds of brand mentions go to the same brand).
Their full Reddit histories look ordinary
I fetched full histories for 23 tail accounts behind the three brands with a signal, plus 23 comparison accounts. The tail accounts are 4.5 years old at the median—the same as comparison. Only 3% of their comments are in the six knife subs (vs 10% for comparison). Across their whole history they name seven knife brands, not one.
That is not what a warmed, single-purpose account looks like. It is what a person who is on Reddit a lot, with strong feelings about one knife maker, looks like—or a well-run paid account, which is why vendors sell aged accounts.
What I take from it
Adding "reddit" to a knife search still gets you humans, mostly. For a couple of brands, a quarter to a third of the buying advice comes from accounts that mostly recommend that brand. Whether those are fans or paid, I do not know; after reading their histories I lean toward fans and hold the lean loosely.
The practical check: when a knife recommendation comes from an account you do not recognize, click through and see whether it has ever named a different brand.
Code, anonymized data and charts are in the public repository. Usernames, comment text, account histories and the brand-code key are not published.