{"article":{"slug":"the-problem-with-ai-verifiers-provenance-in-the-age-of-ai","title":"The Problem with AI Verifiers: Provenance in the Age of AI","subtitle":null,"summary":"Three years after ChatGPT arrived, the institutions that shape writing—publishers, prizes, newsrooms, platforms, universities, communities—have still not worked out what to do about generative AI. Not for lack of trying.","content_type":"essay","language":"en","canonical_url":"https://ellipsus.com/blog/the-problem-with-ai-verifiers","author":{"name":"Rex Mizrach","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Ellipsus","url":"https://ellipsus.com/","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Writing","slug":"writing","url":"https://listedarticles.com/topics/writing"},{"name":"AI Policy","slug":"ai-policy","url":"https://listedarticles.com/topics/ai-policy"},{"name":"Opinion","slug":"opinion","url":"https://listedarticles.com/topics/opinion"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2917,"reading_minutes":13,"published_at":"2026-09-10T00:00:00.000Z","added_at":"2026-10-02T15:16:52.781Z","updated_at":"2026-10-02T15:16:52.781Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-problem-with-ai-verifiers-provenance-in-the-age-of-ai","markdown_url":"https://listedarticles.com/articles/the-problem-with-ai-verifiers-provenance-in-the-age-of-ai.md","example":false,"citation":"Rex Mizrach, Ellipsus. \"The Problem with AI Verifiers: Provenance in the Age of AI.\" 10 Sept 2026. https://ellipsus.com/blog/the-problem-with-ai-verifiers (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://ellipsus.com/blog/the-problem-with-ai-verifiers"},"body_markdown":"## The summer of our discontent\n\nThree years after ChatGPT arrived, the institutions that shape writing—publishers, prizes, newsrooms, platforms, universities, communities—have still not worked out what to do about generative AI.\n\nNot for lack of trying. “[*Human authored*](https://authorsguild.org/human-authored/)” stamps have been added to book covers (as defined by “reviewers”); various opt-in [constitutions](https://www.createdbyhumans.ai/post/created-by-humans-partners-with-the-authors-guild) have been drafted; games, movies, and TV shows forefront their human origin. “*No AI*” statements have been required on countless submissions forms; and the [teaching](https://www.nytimes.com/2026/09/09/world/10int-theworld-web-ai-education.html) profession has collectively thrown up its hands (not least because many school systems push students to embrace the tools teachers are now asked to police). There’s no shortage of rules in 2026, except that nobody can tell, with any degree of reliability, if those rules were broken.\n\nEnter the AI detector.\n\nThis summer, the publishing industry saw major book deals collapse as manuscripts were fed into AI detection tools. The first wave of accusations had relied on models like GPTZero (already used by many schools and universities, despite a wide range of reports of “false positives and negatives”); then through a newcomer, Pangram—billed as the most accurate AI detector yet built.\n\nIn July, a debut novel that had inspired a fourteen-way publishers’ auction—and a reported $2 million advance—was [withdrawn](https://www.wsj.com/business/media/author-whose-2-million-book-deal-was-derailed-by-ai-concerns-says-hes-innocent-c9f4c91b) by the author’s own agent after the publication *Publisher’s Lunc*h fed the manuscript through Pangram. And it wasn’t alone. Mia Ballard’s novel *Shy Girl* was also pulled; H.M. Wolfe’s novel [*Daggermouth*](https://www.theatlantic.com/technology/2026/07/daggermouth-novel-bestseller-ai/688067/) was acquired by Simon & Schuster, despite a Pangram “score” of “60% AI.”\n\nHow true were these scores? Though the verifiers seemed certain, the humans were not. In May, a story published in Granta, and selected as a regional finalist for the Commonwealth Short Story Prize, was flagged as AI after a Wharton professor ran it through Pangram. The story went on to win the Commonwealth Prize, but Granta [announced](https://www.theguardian.com/books/2026/jun/20/granta-magazine-commonwealth-short-story-prize-ai) it would “review its selection process”, and many remained unconvinced of the story’s authorship. Meanwhile, fanfic writers weathered waves of social media [callouts](https://fansplaining.com/heated-rivalry-is-at-the-center-of-fanfics-ai-reckoning/) over suspected AI use, with authors named, dogpiled, and even driven offline by their communities. And on Amazon, new book releases nearly [tripled](https://www.nber.org/papers/w34777\\) since 2022; which makes sense only in the slopocalypse of sham books overtaking huge parts of the platform.\n\nEvery platform that has ever hosted writing has been scraped. Some owners knew, and were paid for selling their users words and images wholesale (looking at you, Tumblr). Some found out after the fact, as the volunteers and authors who run semi-public archives like AO3 did. Hundreds of thousands of physical books and copyrighted materials have been [bought, scanned and “pulped”](https://www.newyorker.com/culture/the-lede/destroying-books-to-build-a-mind) to train new AI models. And companies like Google have continued to train its own models on the words written inside their own tools.\n\nNow, the same models that caused all the furor are being turned back on the writers that fed them. First to compete, and then to judge.\n\nWe’ve spent a long time thinking about this growing loss of trust—trust in creative work, trust in the *real*, and trust in each other—and the gradual dissolution of the once standard assumption that a piece of writing has a human behind it. As LLMs improve month over month and institutions fall further behind them, the desire for a tool that can *just* *tell* has become overwhelming. Hence the verifier’s current ubiquity, to the detriment of creative culture and *adding* to these losses.\n\nA hero that gets it wrong is not really a hero. And AI verifiers, as we see them, get it wrong in ways that are threatening the labor of writing: doing real, verifiable harm to writers, and damaging creativity itself by teaching us to fear using our own voice, and teaching readers to trust in nothing at all.\n\n## What are AI verifiers, and how do they work?\n\nAI verifiers are AI.\n\nAI verifiers are relatively easy to understand, though their marketing tends to obscure this. Detector models are built from many of the same resources used to create models like ChatGPT and Claude—the vast swathes of texts that humans have written—or, at least, the texts that AI companies can scan and discard, or the platforms they can scrape. There is, generally speaking, not a genre, nor a platform, nor an author, nor an individual text that has been spared. As LLMs are trained on all human work, so too are the verifiers judging their outputs.\n\nTo [paraphrase](<https://www.pangram.com/blog/how-does-pangram-work)>) Pangram’s own explainer: AI verifiers evaluate text based on “author identification”. They learn the stylistic “decisions” made by the author of the text in question, that characterize a given author, and then estimate if the author was AI or human. In a sense, they treat Claude or ChatGPT *as* *authors* (with their own style, cadence, grammatical nuances, and punctuation habits), and then compare \"your writing against them. The closer your writing is to the chatbot’s, the more likely your writing will be flagged as AI.\n\nYou already know all about the “tells”—or, at least, the *assumptions*. Pity the poor em dash (*a favorite of ours—and [worth defending](https://merch.ellipsus.com/collections/defend-the-em-dash)*). Pour one out for the [rule-of-three,](https://www.theguardian.com/books/ng-interactive/2026/jul/04/future-of-fiction-next-great-novel-ai-language-chat-gpt) [“it’s not X, it’s Y”](https://www.theatlantic.com/ideas/2026/06/ai-writing-reading-nazir/687419/), *delve, tapestry, nuance* (*because we can’t have* that *anymore*)—and adverbs, apparently.\n\nThe issue is that all of these are human inventions. They’re made and fostered by the habits of centuries of English language writing of copy editors, *Strunk and White*, and every decent writing teacher who impressed on their students to vary their sentence length. Basically—the chatbot style is the statistical average of *us*.\n\n## The numbers\n\nPangram, with its release of Pangram 4 in late July, [claims](https://www.pangram.com/blog/introducing-pangram-4) a false positive rate of 0.0041% (or 1 wrongly flagged out of every 24,000 documents).\n\nThose are no doubt impressive numbers, but seem to be difficult to square with the reports coming from the ground.\n\nThe hundreds, likely thousands of writers who have experimented with Pangram on their own writing appear to tell a different story. A glance of at any comments section, any writing subreddit or X thread where verification is discussed reveals widespread reports of false positives *and* false negatives; the same tool waving through machine-generated text while flagging a decade-old blog post as \"100% AI\". As *The Atlantic*'s Matteo Wong [wrote](https://www.theatlantic.com/technology/2026/05/pangram-ai-detection-accuracy/687381/): \"While Pangram is accumulating the power to end reputations and careers, the tool does make mistakes, perhaps to a greater extent than is currently understood.”\n\nOn July 21, Substack announced a partnership with Pangram. Since then, all text posts on Substack longer than a hundred words can be subjected to an AI verification test with a click of a button. Substack framed this decision as part of a bid to combat AI slop on the site (and not end up like, say, LinkedIn), writing in its [announcement:](https://post.substack.com/p/against-claudefishing) \"people should know what they're getting.\" In response, many voiced frustration with the partnership and [doubts about the tool’s accuracy](https://www.404media.co/substackers-say-new-ai-detection-tool-is-a-witch-hunt/). And Pangram’s posture toward critics… did not help. When challenged on social media, Pangram took an aggressive stance. When, in early August, a user complained that Pangram incorrectly flagged their human written work as AI generated, a Pangram spokesperson ran the complaint through the detector and posted the result: “100% of this text is AI.”\n\nIt’s hard to see this as anything but a tactic designed to deter critics: *If you speak out against us, we’ll run you through the machine.*Whether this will serve it and other platforms well remains to be seen… but it certainly isn’t winning over writers, scores of whom have voiced valid complaints about the tool and its introduction into online publishing spaces.\n\nDespite the pushback, Substack has stood by its partnership. And more platforms will likely follow, as well as newsrooms, classrooms, academic institutions, literary prizes, publishing houses… with one verifier or another; “[one-shot gotcha machine](https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-is-broken-but)”.\n\n## More verifiers, more problems\n\nOn the issue of accuracy, there’s an important question that verifiers continually evade: Can machines—or humans trained on the current slopocalypse—truly identify “AI tells” at all, given that the the models producing them were trained on human writing?\n\nIf someone’s natural writing style hews closer to that of Claude or ChatGPT, their work will show a higher rate of AI-generation. Not because you used AI, but because AI, in the actual sense, used you.\n\nAnd writers are already feeling the pressure to avoid “AI tells.” Long before Pangram, people were stripping em dashes from their drafts (couldn’t be us!). Now, with pressures of external verification, the forces acting to discourage writers are reaching a unsustainable level, leading some to speculate that this is part and parcel of the strategy motivating AI verifiers.\n\nWe're all slop-maxxed in 2026. Everyone has read enough AI writing to have started to see it everywhere, including where it isn't.\n\nConsider what the reports actually show. Many writers report one result, and others the opposite, for the same piece of text run through Pangram and other verifiers. In some cases, changing a single word—or a common name—was enough to flip a result from \"100% Human\" to \"100% AI.\" Even OpenAI [cautions against](https://help.openai.com/en/articles/8313351-how-can-educators-respond-to-students-presenting-ai-generated-content-as-their-own) the use of AI detectors, describing tools that judge human work against AI stylistic \"signatures\" as untrustworthy, prone to false positives, and given to \"random\" judgments with \"no basis in fact.\" The company had attempted to train its own detector and found that \"it labeled human-written text like Shakespeare and the Declaration of Independence as AI-generated.\" This has happened to many detectors, and it will continue to.\n\nAnd there are still many open questions. If a verifier's training skews toward academic writing and journalism, how will its judgments be distorted inside specific genres—creative fiction, where adverbs (a “tell”) are far more likely to be found than in a research paper? Tell ChatGPT to \"write fanfic\" and it will produce something in the style of writers publishing to AO3 (and especially, for *reasons*: the [Omegaverse](https://www.wired.com/story/fanfiction-omegaverse-sex-trope-artificial-intelligence-knotting/)). And what happens to dialects and non-native English, and to neurodivergent writers (whom [at least one study](https://teaching.unl.edu/ai-exchange/challenge-ai-checkers/) has found detectors flag at higher rates)? The answers are even more nebulous.\n\nAt the end of 2026, the [EU’s AI act will require AI text to be watermarked](https://digital-strategy.ec.europa.eu/en/policies/eu-icons-labelling-ai-generated-content), something campaigners have demanded for years. Anthropic has already [announced](https://www.anthropic.com/news/claude-text-watermark) that its new Claude models will contain these “imperceptible marks” worldwide.\n\nBut watermarks could complicate things even more. Anthropic itself points out that the absence of watermarks will not prove that AI was not used. And text watermarks are, reportedly, statistical; meaning detection yet again comes down to the likelihood of particular words or cadence appearing in particular places. E.g., the same “tells” Claude was trained on.\n\nThen there’s the reality that LLMs are constantly evolving. We are on the cusp of—or already living in—a world where chatbots can produce text that is statistically and stylistically indistinguishable from human writing. Studies find that people already [struggle to discriminate](https://www.theguardian.com/technology/2026/aug/05/ai-generated-stories-rated-better-quality-than-human-written-ones-study-finds) between human written text and their AI-generated counterparts (and in one, readers rated AI stories *better*). Given the delays and difficulties of keeping up with the breakneck pace of AI development (by design), the risk of verifiers getting it wrong will only grow. It’s hard to see how comparison models can be viable in even the medium-term.\n\n## The vibes are degrading\n\nIf you’re not a weird AI booster, you wouldn’t be surprised to know that readers do not want to read AI work. In our [survey](https://ellipsus.com/blog/survey-on-writing-and-ai) of nearly 6,000 writers and readers, over 98% of respondents “actively avoid content they suspect is AI-generated”.But that’s also why verifiers have become so ubiquitous—demand for *the real* is very real. And with the tools accessible to anyone, the culture turns in very strange ways.\n\nIn early August, Jack Osbourne—yes, Ozzy's son—[accused](https://spitfirenews.com/p/jack-osbourne-ai-kat-tenbarge-allegations-debunked) the journalist Kat Tenbarge of using AI to write a *Rolling Stone* piece on the Tate brothers, citing an AI checker's 91% score and her \"big hyphens\" as evidence *(truly the strangest timeline).* Tenbarge answered with her Google Docs version history, armed with timestamps. \"AI \"introduces all of this doubt into our society. Doubt that can very easily be weaponized against journalists and real creatives\"—or, we'd add, by anyone with an axe to grind against any particular author.\n\nThe culture emerging in many writing spaces today is a culture of suspicion. We’ve all heard these stories, and it goes without saying that having one's work flagged as AI carries major consequences. There’s a sense of dread from many writers worrying about posting work online at all, in the assumption that someone, somewhere, will get suspicious of an em dash or well-placed adverb. And many are changing their styles, honed over many years and deeply entwined with their voice and their personhood, the closest thing a writer has to a fingerprint—to fit within the confines of a verifier’s assumption.\n\n## Do AI verifiers have a use case?\n\nMany publishers and platforms think so, if only for lack any other options. AI is here to stay, and what counts as acceptable content is [shifting under our feet as we write](https://www.reuters.com/legal/litigation/openai-new-york-times-case-tees-up-key-test-ai-training-under-copyright-law-2026-09-08/).\n\nBut right now, we think there are too much inconsistency, and too little transparency about what is being measured for verifiers to be the panacea that they purport to be. Clocking AI writing by way of conventions or “tells”—all of which are being trained out of AI models as they improve—will be a thing of the past within months, depending on the rate of acceleration. And AI companies show no sign of slowing down.\n\nAnd then there’s the deeper layer, which we keep returning to, as writers—why should we—on whose writing these models were trained—be punished for stylistic habits that happen to converge with the LLM models trained on human work?\n\n## Emboss\n\nThrough all of this, an undercurrent pulses just beneath the surface: the gradual sense of a general devaluing of human labor.\n\nCreative processes, and creative progress, can feel weightless when their value can be undone by assumptions. The immense pride a creative feels in the act of spending time creating, writing, editing, thinking is the elemental substance of writing, and of human creativity.\n\nWriters need defense against accusation. But they also need something more than that; a reclamation of the value of human work.\n\nLook at how writers are doing this work right now—journalists showing videos of their Google Docs version history, fic writers sharing timestamps and screenshots of their drafts in Discord servers; novelists showing screenshots of notebooks and email chains to prove to their own agents that they wrote their own book. The evidence exists, but it’s scattered and awkward to assemble—and probably produced under duress, after an accusation has already happened.\n\nWay back in 2023, we had an idea: that the process as recorded in a document itself acts as a kind of fingerprint of its author. And that the provenance could deflect from bad-faith accusations and false positives, and create a durable trace of the time and labor behind the words.\n\nAn integrated tool that could follow the entire journey of a piece—from initial drafting to final edits, and all the way through to publishing—could create a digital trail, or record of the work behind it. That record would support writers at every stage of their work, preserving their unique voice and vision in a world where the origins of creative work are more clouded than ever.\n\n\nThat’s what we set out to build, and that's what we’ve started building with Emboss.\n\nThe idea is the same as it was years ago: make it easy for writers to do what they’re already doing, and make it simple and transparent (*and of course, AI-free*). The answer is not a “better detector”, but an honest record; one which does not punish writers.\n\nEmboss’s writing journey metrics come from the same underlying processes as versionhistory and real-time writing in Ellipsus. As you write, edit, pause, revise, your document is already creating a record. The writing journey turns that record into an easy-to-read picture of a document’s progress. Everything is AI-free and private by default, and shared only on your own terms.\n\n**And as of today, the writing journey is free for everyone, because all writers should be able to stand behind their work.**\n\nEmboss is not a verifier. It does not judge the quality of your writing, and it does not assign an \"AI or human\" probability. It shows the labor—yours—and lets that speak for you. Nothing about your writing process is public unless you actively choose to share it.\n\n[You can view this post’s writing journey here.](https://ellipsus.com/writing-journey/1muvonVPrcMC1X2s2XnoJm)\n\nNo platform can fully control how communities, readers, institutions, and creative spaces interpret the inclusion or absence of a record of provenance. But we at Ellipsus can be very clear about our own stance: not sharing a writing journey should never be treated as suspicious, just as the results of an AI verifier should never be taken as proof of AI use.\n\nBut for writers who wish to show their work’s record, we think there should be a way to do that without judgment, and which doesn’t require feeding a document to AI.\n\nWe feel that the rise of AI in creative spaces, and what that means for writers, is one of the biggest challenges facing writers today. And the time and effort a writer spends making something their own is worth sharing.\n\nThe story—all of it—should be human.","body_html":"<h2 id=\"the-summer-of-our-discontent\">The summer of our discontent</h2>\n<p>Three years after ChatGPT arrived, the institutions that shape writing—publishers, prizes, newsrooms, platforms, universities, communities—have still not worked out what to do about generative AI.</p>\n<p>Not for lack of trying. “<a href=\"https://authorsguild.org/human-authored/\" rel=\"nofollow ugc noopener\"><em>Human authored</em></a>” stamps have been added to book covers (as defined by “reviewers”); various opt-in <a href=\"https://www.createdbyhumans.ai/post/created-by-humans-partners-with-the-authors-guild\" rel=\"nofollow ugc noopener\">constitutions</a> have been drafted; games, movies, and TV shows forefront their human origin. “<em>No AI</em>” statements have been required on countless submissions forms; and the <a href=\"https://www.nytimes.com/2026/09/09/world/10int-theworld-web-ai-education.html\" rel=\"nofollow ugc noopener\">teaching</a> profession has collectively thrown up its hands (not least because many school systems push students to embrace the tools teachers are now asked to police). There’s no shortage of rules in 2026, except that nobody can tell, with any degree of reliability, if those rules were broken.</p>\n<p>Enter the AI detector.</p>\n<p>This summer, the publishing industry saw major book deals collapse as manuscripts were fed into AI detection tools. The first wave of accusations had relied on models like GPTZero (already used by many schools and universities, despite a wide range of reports of “false positives and negatives”); then through a newcomer, Pangram—billed as the most accurate AI detector yet built.</p>\n<p>In July, a debut novel that had inspired a fourteen-way publishers’ auction—and a reported $2 million advance—was <a href=\"https://www.wsj.com/business/media/author-whose-2-million-book-deal-was-derailed-by-ai-concerns-says-hes-innocent-c9f4c91b\" rel=\"nofollow ugc noopener\">withdrawn</a> by the author’s own agent after the publication <em>Publisher’s Lunc</em>h fed the manuscript through Pangram. And it wasn’t alone. Mia Ballard’s novel <em>Shy Girl</em> was also pulled; H.M. Wolfe’s novel <a href=\"https://www.theatlantic.com/technology/2026/07/daggermouth-novel-bestseller-ai/688067/\" rel=\"nofollow ugc noopener\"><em>Daggermouth</em></a> was acquired by Simon &amp; Schuster, despite a Pangram “score” of “60% AI.”</p>\n<p>How true were these scores? Though the verifiers seemed certain, the humans were not. In May, a story published in Granta, and selected as a regional finalist for the Commonwealth Short Story Prize, was flagged as AI after a Wharton professor ran it through Pangram. The story went on to win the Commonwealth Prize, but Granta <a href=\"https://www.theguardian.com/books/2026/jun/20/granta-magazine-commonwealth-short-story-prize-ai\" rel=\"nofollow ugc noopener\">announced</a> it would “review its selection process”, and many remained unconvinced of the story’s authorship. Meanwhile, fanfic writers weathered waves of social media <a href=\"https://fansplaining.com/heated-rivalry-is-at-the-center-of-fanfics-ai-reckoning/\" rel=\"nofollow ugc noopener\">callouts</a> over suspected AI use, with authors named, dogpiled, and even driven offline by their communities. And on Amazon, new book releases nearly [tripled](<a href=\"https://www.nber.org/papers/w34777\\\" rel=\"nofollow ugc noopener\">https://www.nber.org/papers/w34777\\</a>) since 2022; which makes sense only in the slopocalypse of sham books overtaking huge parts of the platform.</p>\n<p>Every platform that has ever hosted writing has been scraped. Some owners knew, and were paid for selling their users words and images wholesale (looking at you, Tumblr). Some found out after the fact, as the volunteers and authors who run semi-public archives like AO3 did. Hundreds of thousands of physical books and copyrighted materials have been <a href=\"https://www.newyorker.com/culture/the-lede/destroying-books-to-build-a-mind\" rel=\"nofollow ugc noopener\">bought, scanned and “pulped”</a> to train new AI models. And companies like Google have continued to train its own models on the words written inside their own tools.</p>\n<p>Now, the same models that caused all the furor are being turned back on the writers that fed them. First to compete, and then to judge.</p>\n<p>We’ve spent a long time thinking about this growing loss of trust—trust in creative work, trust in the <em>real</em>, and trust in each other—and the gradual dissolution of the once standard assumption that a piece of writing has a human behind it. As LLMs improve month over month and institutions fall further behind them, the desire for a tool that can <em>just</em> <em>tell</em> has become overwhelming. Hence the verifier’s current ubiquity, to the detriment of creative culture and <em>adding</em> to these losses.</p>\n<p>A hero that gets it wrong is not really a hero. And AI verifiers, as we see them, get it wrong in ways that are threatening the labor of writing: doing real, verifiable harm to writers, and damaging creativity itself by teaching us to fear using our own voice, and teaching readers to trust in nothing at all.</p>\n<h2 id=\"what-are-ai-verifiers-and-how-do-they-work\">What are AI verifiers, and how do they work?</h2>\n<p>AI verifiers are AI.</p>\n<p>AI verifiers are relatively easy to understand, though their marketing tends to obscure this. Detector models are built from many of the same resources used to create models like ChatGPT and Claude—the vast swathes of texts that humans have written—or, at least, the texts that AI companies can scan and discard, or the platforms they can scrape. There is, generally speaking, not a genre, nor a platform, nor an author, nor an individual text that has been spared. As LLMs are trained on all human work, so too are the verifiers judging their outputs.</p>\n<p>To <a href=\"https://www.pangram.com/blog/how-does-pangram-work\" rel=\"nofollow ugc noopener\">paraphrase</a>&gt;) Pangram’s own explainer: AI verifiers evaluate text based on “author identification”. They learn the stylistic “decisions” made by the author of the text in question, that characterize a given author, and then estimate if the author was AI or human. In a sense, they treat Claude or ChatGPT <em>as</em> <em>authors</em> (with their own style, cadence, grammatical nuances, and punctuation habits), and then compare &quot;your writing against them. The closer your writing is to the chatbot’s, the more likely your writing will be flagged as AI.</p>\n<p>You already know all about the “tells”—or, at least, the <em>assumptions</em>. Pity the poor em dash (<em>a favorite of ours—and <a href=\"https://merch.ellipsus.com/collections/defend-the-em-dash\" rel=\"nofollow ugc noopener\">worth defending</a></em>). Pour one out for the <a href=\"https://www.theguardian.com/books/ng-interactive/2026/jul/04/future-of-fiction-next-great-novel-ai-language-chat-gpt\" rel=\"nofollow ugc noopener\">rule-of-three,</a> <a href=\"https://www.theatlantic.com/ideas/2026/06/ai-writing-reading-nazir/687419/\" rel=\"nofollow ugc noopener\">“it’s not X, it’s Y”</a>, <em>delve, tapestry, nuance</em> (<em>because we can’t have</em> that <em>anymore</em>)—and adverbs, apparently.</p>\n<p>The issue is that all of these are human inventions. They’re made and fostered by the habits of centuries of English language writing of copy editors, <em>Strunk and White</em>, and every decent writing teacher who impressed on their students to vary their sentence length. Basically—the chatbot style is the statistical average of <em>us</em>.</p>\n<h2 id=\"the-numbers\">The numbers</h2>\n<p>Pangram, with its release of Pangram 4 in late July, <a href=\"https://www.pangram.com/blog/introducing-pangram-4\" rel=\"nofollow ugc noopener\">claims</a> a false positive rate of 0.0041% (or 1 wrongly flagged out of every 24,000 documents).</p>\n<p>Those are no doubt impressive numbers, but seem to be difficult to square with the reports coming from the ground.</p>\n<p>The hundreds, likely thousands of writers who have experimented with Pangram on their own writing appear to tell a different story. A glance of at any comments section, any writing subreddit or X thread where verification is discussed reveals widespread reports of false positives <em>and</em> false negatives; the same tool waving through machine-generated text while flagging a decade-old blog post as &quot;100% AI&quot;. As <em>The Atlantic</em>&#39;s Matteo Wong <a href=\"https://www.theatlantic.com/technology/2026/05/pangram-ai-detection-accuracy/687381/\" rel=\"nofollow ugc noopener\">wrote</a>: &quot;While Pangram is accumulating the power to end reputations and careers, the tool does make mistakes, perhaps to a greater extent than is currently understood.”</p>\n<p>On July 21, Substack announced a partnership with Pangram. Since then, all text posts on Substack longer than a hundred words can be subjected to an AI verification test with a click of a button. Substack framed this decision as part of a bid to combat AI slop on the site (and not end up like, say, LinkedIn), writing in its <a href=\"https://post.substack.com/p/against-claudefishing\" rel=\"nofollow ugc noopener\">announcement:</a> &quot;people should know what they&#39;re getting.&quot; In response, many voiced frustration with the partnership and <a href=\"https://www.404media.co/substackers-say-new-ai-detection-tool-is-a-witch-hunt/\" rel=\"nofollow ugc noopener\">doubts about the tool’s accuracy</a>. And Pangram’s posture toward critics… did not help. When challenged on social media, Pangram took an aggressive stance. When, in early August, a user complained that Pangram incorrectly flagged their human written work as AI generated, a Pangram spokesperson ran the complaint through the detector and posted the result: “100% of this text is AI.”</p>\n<p>It’s hard to see this as anything but a tactic designed to deter critics: <em>If you speak out against us, we’ll run you through the machine.</em>Whether this will serve it and other platforms well remains to be seen… but it certainly isn’t winning over writers, scores of whom have voiced valid complaints about the tool and its introduction into online publishing spaces.</p>\n<p>Despite the pushback, Substack has stood by its partnership. And more platforms will likely follow, as well as newsrooms, classrooms, academic institutions, literary prizes, publishing houses… with one verifier or another; “<a href=\"https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-is-broken-but\" rel=\"nofollow ugc noopener\">one-shot gotcha machine</a>”.</p>\n<h2 id=\"more-verifiers-more-problems\">More verifiers, more problems</h2>\n<p>On the issue of accuracy, there’s an important question that verifiers continually evade: Can machines—or humans trained on the current slopocalypse—truly identify “AI tells” at all, given that the the models producing them were trained on human writing?</p>\n<p>If someone’s natural writing style hews closer to that of Claude or ChatGPT, their work will show a higher rate of AI-generation. Not because you used AI, but because AI, in the actual sense, used you.</p>\n<p>And writers are already feeling the pressure to avoid “AI tells.” Long before Pangram, people were stripping em dashes from their drafts (couldn’t be us!). Now, with pressures of external verification, the forces acting to discourage writers are reaching a unsustainable level, leading some to speculate that this is part and parcel of the strategy motivating AI verifiers.</p>\n<p>We&#39;re all slop-maxxed in 2026. Everyone has read enough AI writing to have started to see it everywhere, including where it isn&#39;t.</p>\n<p>Consider what the reports actually show. Many writers report one result, and others the opposite, for the same piece of text run through Pangram and other verifiers. In some cases, changing a single word—or a common name—was enough to flip a result from &quot;100% Human&quot; to &quot;100% AI.&quot; Even OpenAI <a href=\"https://help.openai.com/en/articles/8313351-how-can-educators-respond-to-students-presenting-ai-generated-content-as-their-own\" rel=\"nofollow ugc noopener\">cautions against</a> the use of AI detectors, describing tools that judge human work against AI stylistic &quot;signatures&quot; as untrustworthy, prone to false positives, and given to &quot;random&quot; judgments with &quot;no basis in fact.&quot; The company had attempted to train its own detector and found that &quot;it labeled human-written text like Shakespeare and the Declaration of Independence as AI-generated.&quot; This has happened to many detectors, and it will continue to.</p>\n<p>And there are still many open questions. If a verifier&#39;s training skews toward academic writing and journalism, how will its judgments be distorted inside specific genres—creative fiction, where adverbs (a “tell”) are far more likely to be found than in a research paper? Tell ChatGPT to &quot;write fanfic&quot; and it will produce something in the style of writers publishing to AO3 (and especially, for <em>reasons</em>: the <a href=\"https://www.wired.com/story/fanfiction-omegaverse-sex-trope-artificial-intelligence-knotting/\" rel=\"nofollow ugc noopener\">Omegaverse</a>). And what happens to dialects and non-native English, and to neurodivergent writers (whom <a href=\"https://teaching.unl.edu/ai-exchange/challenge-ai-checkers/\" rel=\"nofollow ugc noopener\">at least one study</a> has found detectors flag at higher rates)? The answers are even more nebulous.</p>\n<p>At the end of 2026, the <a href=\"https://digital-strategy.ec.europa.eu/en/policies/eu-icons-labelling-ai-generated-content\" rel=\"nofollow ugc noopener\">EU’s AI act will require AI text to be watermarked</a>, something campaigners have demanded for years. Anthropic has already <a href=\"https://www.anthropic.com/news/claude-text-watermark\" rel=\"nofollow ugc noopener\">announced</a> that its new Claude models will contain these “imperceptible marks” worldwide.</p>\n<p>But watermarks could complicate things even more. Anthropic itself points out that the absence of watermarks will not prove that AI was not used. And text watermarks are, reportedly, statistical; meaning detection yet again comes down to the likelihood of particular words or cadence appearing in particular places. E.g., the same “tells” Claude was trained on.</p>\n<p>Then there’s the reality that LLMs are constantly evolving. We are on the cusp of—or already living in—a world where chatbots can produce text that is statistically and stylistically indistinguishable from human writing. Studies find that people already <a href=\"https://www.theguardian.com/technology/2026/aug/05/ai-generated-stories-rated-better-quality-than-human-written-ones-study-finds\" rel=\"nofollow ugc noopener\">struggle to discriminate</a> between human written text and their AI-generated counterparts (and in one, readers rated AI stories <em>better</em>). Given the delays and difficulties of keeping up with the breakneck pace of AI development (by design), the risk of verifiers getting it wrong will only grow. It’s hard to see how comparison models can be viable in even the medium-term.</p>\n<h2 id=\"the-vibes-are-degrading\">The vibes are degrading</h2>\n<p>If you’re not a weird AI booster, you wouldn’t be surprised to know that readers do not want to read AI work. In our <a href=\"https://ellipsus.com/blog/survey-on-writing-and-ai\" rel=\"nofollow ugc noopener\">survey</a> of nearly 6,000 writers and readers, over 98% of respondents “actively avoid content they suspect is AI-generated”.But that’s also why verifiers have become so ubiquitous—demand for <em>the real</em> is very real. And with the tools accessible to anyone, the culture turns in very strange ways.</p>\n<p>In early August, Jack Osbourne—yes, Ozzy&#39;s son—<a href=\"https://spitfirenews.com/p/jack-osbourne-ai-kat-tenbarge-allegations-debunked\" rel=\"nofollow ugc noopener\">accused</a> the journalist Kat Tenbarge of using AI to write a <em>Rolling Stone</em> piece on the Tate brothers, citing an AI checker&#39;s 91% score and her &quot;big hyphens&quot; as evidence <em>(truly the strangest timeline).</em> Tenbarge answered with her Google Docs version history, armed with timestamps. &quot;AI &quot;introduces all of this doubt into our society. Doubt that can very easily be weaponized against journalists and real creatives&quot;—or, we&#39;d add, by anyone with an axe to grind against any particular author.</p>\n<p>The culture emerging in many writing spaces today is a culture of suspicion. We’ve all heard these stories, and it goes without saying that having one&#39;s work flagged as AI carries major consequences. There’s a sense of dread from many writers worrying about posting work online at all, in the assumption that someone, somewhere, will get suspicious of an em dash or well-placed adverb. And many are changing their styles, honed over many years and deeply entwined with their voice and their personhood, the closest thing a writer has to a fingerprint—to fit within the confines of a verifier’s assumption.</p>\n<h2 id=\"do-ai-verifiers-have-a-use-case\">Do AI verifiers have a use case?</h2>\n<p>Many publishers and platforms think so, if only for lack any other options. AI is here to stay, and what counts as acceptable content is <a href=\"https://www.reuters.com/legal/litigation/openai-new-york-times-case-tees-up-key-test-ai-training-under-copyright-law-2026-09-08/\" rel=\"nofollow ugc noopener\">shifting under our feet as we write</a>.</p>\n<p>But right now, we think there are too much inconsistency, and too little transparency about what is being measured for verifiers to be the panacea that they purport to be. Clocking AI writing by way of conventions or “tells”—all of which are being trained out of AI models as they improve—will be a thing of the past within months, depending on the rate of acceleration. And AI companies show no sign of slowing down.</p>\n<p>And then there’s the deeper layer, which we keep returning to, as writers—why should we—on whose writing these models were trained—be punished for stylistic habits that happen to converge with the LLM models trained on human work?</p>\n<h2 id=\"emboss\">Emboss</h2>\n<p>Through all of this, an undercurrent pulses just beneath the surface: the gradual sense of a general devaluing of human labor.</p>\n<p>Creative processes, and creative progress, can feel weightless when their value can be undone by assumptions. The immense pride a creative feels in the act of spending time creating, writing, editing, thinking is the elemental substance of writing, and of human creativity.</p>\n<p>Writers need defense against accusation. But they also need something more than that; a reclamation of the value of human work.</p>\n<p>Look at how writers are doing this work right now—journalists showing videos of their Google Docs version history, fic writers sharing timestamps and screenshots of their drafts in Discord servers; novelists showing screenshots of notebooks and email chains to prove to their own agents that they wrote their own book. The evidence exists, but it’s scattered and awkward to assemble—and probably produced under duress, after an accusation has already happened.</p>\n<p>Way back in 2023, we had an idea: that the process as recorded in a document itself acts as a kind of fingerprint of its author. And that the provenance could deflect from bad-faith accusations and false positives, and create a durable trace of the time and labor behind the words.</p>\n<p>An integrated tool that could follow the entire journey of a piece—from initial drafting to final edits, and all the way through to publishing—could create a digital trail, or record of the work behind it. That record would support writers at every stage of their work, preserving their unique voice and vision in a world where the origins of creative work are more clouded than ever.</p>\n<p>That’s what we set out to build, and that&#39;s what we’ve started building with Emboss.</p>\n<p>The idea is the same as it was years ago: make it easy for writers to do what they’re already doing, and make it simple and transparent (<em>and of course, AI-free</em>). The answer is not a “better detector”, but an honest record; one which does not punish writers.</p>\n<p>Emboss’s writing journey metrics come from the same underlying processes as versionhistory and real-time writing in Ellipsus. As you write, edit, pause, revise, your document is already creating a record. The writing journey turns that record into an easy-to-read picture of a document’s progress. Everything is AI-free and private by default, and shared only on your own terms.</p>\n<p><strong>And as of today, the writing journey is free for everyone, because all writers should be able to stand behind their work.</strong></p>\n<p>Emboss is not a verifier. It does not judge the quality of your writing, and it does not assign an &quot;AI or human&quot; probability. It shows the labor—yours—and lets that speak for you. Nothing about your writing process is public unless you actively choose to share it.</p>\n<p><a href=\"https://ellipsus.com/writing-journey/1muvonVPrcMC1X2s2XnoJm\" rel=\"nofollow ugc noopener\">You can view this post’s writing journey here.</a></p>\n<p>No platform can fully control how communities, readers, institutions, and creative spaces interpret the inclusion or absence of a record of provenance. But we at Ellipsus can be very clear about our own stance: not sharing a writing journey should never be treated as suspicious, just as the results of an AI verifier should never be taken as proof of AI use.</p>\n<p>But for writers who wish to show their work’s record, we think there should be a way to do that without judgment, and which doesn’t require feeding a document to AI.</p>\n<p>We feel that the rise of AI in creative spaces, and what that means for writers, is one of the biggest challenges facing writers today. And the time and effort a writer spends making something their own is worth sharing.</p>\n<p>The story—all of it—should be human.</p>","headings":[{"level":2,"text":"The summer of our discontent","id":"the-summer-of-our-discontent"},{"level":2,"text":"What are AI verifiers, and how do they work?","id":"what-are-ai-verifiers-and-how-do-they-work"},{"level":2,"text":"The numbers","id":"the-numbers"},{"level":2,"text":"More verifiers, more problems","id":"more-verifiers-more-problems"},{"level":2,"text":"The vibes are degrading","id":"the-vibes-are-degrading"},{"level":2,"text":"Do AI verifiers have a use case?","id":"do-ai-verifiers-have-a-use-case"},{"level":2,"text":"Emboss","id":"emboss"}]}}