---
title: "IdeaLens: Detecting AI Ideas in Long-form Writing"
slug: idealens-detecting-ai-ideas-in-long-form-writing
url: https://listedarticles.com/articles/idealens-detecting-ai-ideas-in-long-form-writing
canonical_url: https://arxiv.org/abs/2610.06778
content_type: research
language: en
published_at: 2026-10-05T00:00:00.000Z
updated_at: 2026-10-07T08:13:26.870Z
author: "Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Marzena Karpinska, John Wieting, Mohit Iyyer et al."
authored_by: human
publisher: "arXiv"
publisher_url: https://arxiv.org/
topics: ["AI", "Research", "Machine Learning"]
about: ["https://listedstartups.com/companies/pangram"]
license: all-rights-reserved
word_count: 266
reading_minutes: 1
citation: "Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Marzena Karpinska, John Wieting, Mohit Iyyer et al., arXiv. \"IdeaLens: Detecting AI Ideas in Long-form Writing.\" 5 Oct 2026. https://arxiv.org/abs/2610.06778 (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# IdeaLens: Detecting AI Ideas in Long-form Writing

> IdeaLens detects whether a document's ideas came from a human or AI, regardless of who wrote the prose, by scoring paraphrased outlines rather than text. Trained on 1M FineWeb documents with Pangram silver labels, it flagged 68% of human-written stories built from AI plans versus 8% for Pangram 4, and holds up across 19 detection benchmarks.

## Abstract

While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.

## Resources

- [Paper on arXiv](https://arxiv.org/abs/2610.06778)
- [IdeaLens demo](https://ideadetector.ai/)
- [idealens on PyPI](https://pypi.org/project/idealens/)
