{"article":{"slug":"does-reddit-have-an-astroturfing-problem-what-the-data-suggests","title":"Does Reddit have an astroturfing problem? What the data suggests","subtitle":null,"summary":"Peter Vijeh analyzes 51,129 knife-subreddit comments: a small tail of accounts writes 11.3% of buying-thread brand mentions versus ~7.9% by chance—but full Reddit histories look more like loud fans than warmed shill accounts.","content_type":"research","language":"en","canonical_url":"https://www.petervijeh.com/projects/reddit-astroturf","author":{"name":"Peter Vijeh","url":null,"person_slug":null,"person_url":null},"authored_by":"human_and_agent","publisher":{"name":"Peter Vijeh","url":"https://www.petervijeh.com/","listing_slug":null,"listing":null},"topics":[{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Startups","slug":"startups","url":"https://listedarticles.com/topics/startups"},{"name":"Privacy","slug":"privacy","url":"https://listedarticles.com/topics/privacy"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":671,"reading_minutes":3,"published_at":"2026-09-20T12:00:00.000Z","added_at":"2026-09-29T00:09:47.477Z","updated_at":"2026-09-29T00:09:47.477Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/does-reddit-have-an-astroturfing-problem-what-the-data-suggests","markdown_url":"https://listedarticles.com/articles/does-reddit-have-an-astroturfing-problem-what-the-data-suggests.md","example":false,"citation":"Peter Vijeh, Peter Vijeh. \"Does Reddit have an astroturfing problem? What the data suggests.\" 20 Sept 2026. https://www.petervijeh.com/projects/reddit-astroturf (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.petervijeh.com/projects/reddit-astroturf"},"body_markdown":"One chef's-knife brand gets 31% of its \"what should I buy\" mentions from 5% of the accounts, four times what chance predicts. So I pulled those accounts' full Reddit histories.\n\nThis article was drafted with AI from my outline and the run logs, then edited. The buying-thread numbers can be recomputed from the published data with one script; the account-history comparison cannot, because it rests on usernames I will not publish.\n\n## Why the question is testable at all\n\nLast year I fine-tuned a small named-entity model, GLiNER, to pull brands, models and steels out of knife comments. That model runs over every comment the New Knife Day scraper collects from six subreddits: r/knives, r/knifeclub, r/chefknives, r/japaneseknives, r/FixedBladeEdc and r/KnifeSteels.\n\nSo for every comment I already have who wrote it, which brands it names, and whether the thread it sits in is someone asking what to buy.\n\n## What astroturfing would look like in the data\n\nServices like REDCmts sell Reddit comments from \"real, aged accounts.\" Taking those sales pages as the description of the product, a paid campaign should show up as:\n\n- A small tail of accounts writing a disproportionate share of brand mentions in buying threads\n- Those accounts naming one brand almost every time\n- Thin accounts: few comments, low scores\n- Young accounts, or histories that are hidden\n- Accounts that post mostly in knife subreddits\n- Links to a store or affiliate page\n\nEvery one of those is also what a devoted fan would produce. Public Reddit data can show concentration, and cannot show why.\n\n## The corpus\n\nA refresh pass went back to 3,607 posts older than 48 hours and refetched comment trees, taking the corpus from 21,673 comments to 51,129.\n\n| | Count |\n| --- | --- |\n| Posts | 6,675 |\n| Comments after refresh | 51,129 |\n| Authors with 10 or more comments | 987 |\n| Brand mentions in buying threads (those authors) | 1,471 |\n\nThe **tail** is the top 5% of the 987 authors on brand-heaviness—49 accounts. Chance is estimated by 1,000 random reassignments of author names across buying-thread mentions.\n\n## What the data supports\n\nIf the tail shows up only because it posts a lot, random reassignment says it should write about 7.9% of brand mentions (typically 6.3%–10.1%). It wrote **11.3%**. Only 2 of 1,000 reassignments reached that.\n\nWhere the extra share lands: r/chefknives and r/knifeclub sit above chance; r/knives (the largest) is within half a point of chance. Three brands get far more of their buying advice from the tail than chance; one large brand gets slightly less; three get none.\n\nOn corpus data, the tail accounts are thin (median 12 comments, median score 1) and loyal (two thirds of brand mentions go to the same brand).\n\n## Their full Reddit histories look ordinary\n\nI fetched full histories for 23 tail accounts behind the three brands with a signal, plus 23 comparison accounts. The tail accounts are 4.5 years old at the median—the same as comparison. Only 3% of their comments are in the six knife subs (vs 10% for comparison). Across their whole history they name seven knife brands, not one.\n\nThat is not what a warmed, single-purpose account looks like. It is what a person who is on Reddit a lot, with strong feelings about one knife maker, looks like—or a well-run paid account, which is why vendors sell aged accounts.\n\n## What I take from it\n\nAdding \"reddit\" to a knife search still gets you humans, mostly. For a couple of brands, a quarter to a third of the buying advice comes from accounts that mostly recommend that brand. Whether those are fans or paid, I do not know; after reading their histories I lean toward fans and hold the lean loosely.\n\nThe practical check: when a knife recommendation comes from an account you do not recognize, click through and see whether it has ever named a different brand.\n\nCode, anonymized data and charts are in the public repository. Usernames, comment text, account histories and the brand-code key are not published.\n","body_html":"<p>One chef&#39;s-knife brand gets 31% of its &quot;what should I buy&quot; mentions from 5% of the accounts, four times what chance predicts. So I pulled those accounts&#39; full Reddit histories.</p>\n<p>This article was drafted with AI from my outline and the run logs, then edited. The buying-thread numbers can be recomputed from the published data with one script; the account-history comparison cannot, because it rests on usernames I will not publish.</p>\n<h2 id=\"why-the-question-is-testable-at-all\">Why the question is testable at all</h2>\n<p>Last year I fine-tuned a small named-entity model, GLiNER, to pull brands, models and steels out of knife comments. That model runs over every comment the New Knife Day scraper collects from six subreddits: r/knives, r/knifeclub, r/chefknives, r/japaneseknives, r/FixedBladeEdc and r/KnifeSteels.</p>\n<p>So for every comment I already have who wrote it, which brands it names, and whether the thread it sits in is someone asking what to buy.</p>\n<h2 id=\"what-astroturfing-would-look-like-in-the-data\">What astroturfing would look like in the data</h2>\n<p>Services like REDCmts sell Reddit comments from &quot;real, aged accounts.&quot; Taking those sales pages as the description of the product, a paid campaign should show up as:</p>\n<ul><li>A small tail of accounts writing a disproportionate share of brand mentions in buying threads</li><li>Those accounts naming one brand almost every time</li><li>Thin accounts: few comments, low scores</li><li>Young accounts, or histories that are hidden</li><li>Accounts that post mostly in knife subreddits</li><li>Links to a store or affiliate page</li></ul>\n<p>Every one of those is also what a devoted fan would produce. Public Reddit data can show concentration, and cannot show why.</p>\n<h2 id=\"the-corpus\">The corpus</h2>\n<p>A refresh pass went back to 3,607 posts older than 48 hours and refetched comment trees, taking the corpus from 21,673 comments to 51,129.</p>\n<div class=\"table-wrap\"><table><thead><tr><th></th><th>Count</th></tr></thead><tbody><tr><td>Posts</td><td>6,675</td></tr><tr><td>Comments after refresh</td><td>51,129</td></tr><tr><td>Authors with 10 or more comments</td><td>987</td></tr><tr><td>Brand mentions in buying threads (those authors)</td><td>1,471</td></tr></tbody></table></div>\n<p>The <strong>tail</strong> is the top 5% of the 987 authors on brand-heaviness—49 accounts. Chance is estimated by 1,000 random reassignments of author names across buying-thread mentions.</p>\n<h2 id=\"what-the-data-supports\">What the data supports</h2>\n<p>If the tail shows up only because it posts a lot, random reassignment says it should write about 7.9% of brand mentions (typically 6.3%–10.1%). It wrote <strong>11.3%</strong>. Only 2 of 1,000 reassignments reached that.</p>\n<p>Where the extra share lands: r/chefknives and r/knifeclub sit above chance; r/knives (the largest) is within half a point of chance. Three brands get far more of their buying advice from the tail than chance; one large brand gets slightly less; three get none.</p>\n<p>On corpus data, the tail accounts are thin (median 12 comments, median score 1) and loyal (two thirds of brand mentions go to the same brand).</p>\n<h2 id=\"their-full-reddit-histories-look-ordinary\">Their full Reddit histories look ordinary</h2>\n<p>I fetched full histories for 23 tail accounts behind the three brands with a signal, plus 23 comparison accounts. The tail accounts are 4.5 years old at the median—the same as comparison. Only 3% of their comments are in the six knife subs (vs 10% for comparison). Across their whole history they name seven knife brands, not one.</p>\n<p>That is not what a warmed, single-purpose account looks like. It is what a person who is on Reddit a lot, with strong feelings about one knife maker, looks like—or a well-run paid account, which is why vendors sell aged accounts.</p>\n<h2 id=\"what-i-take-from-it\">What I take from it</h2>\n<p>Adding &quot;reddit&quot; to a knife search still gets you humans, mostly. For a couple of brands, a quarter to a third of the buying advice comes from accounts that mostly recommend that brand. Whether those are fans or paid, I do not know; after reading their histories I lean toward fans and hold the lean loosely.</p>\n<p>The practical check: when a knife recommendation comes from an account you do not recognize, click through and see whether it has ever named a different brand.</p>\n<p>Code, anonymized data and charts are in the public repository. Usernames, comment text, account histories and the brand-code key are not published.</p>","headings":[{"level":2,"text":"Why the question is testable at all","id":"why-the-question-is-testable-at-all"},{"level":2,"text":"What astroturfing would look like in the data","id":"what-astroturfing-would-look-like-in-the-data"},{"level":2,"text":"The corpus","id":"the-corpus"},{"level":2,"text":"What the data supports","id":"what-the-data-supports"},{"level":2,"text":"Their full Reddit histories look ordinary","id":"their-full-reddit-histories-look-ordinary"},{"level":2,"text":"What I take from it","id":"what-i-take-from-it"}]}}