For our article ‘Erin Reed Has Been Hiding AI-Generated Text in Her Reporting For Years’, we applied analysis techniques to every article Erin Reed published under her solo byline at Erin in the Morning, from the outlet’s launch in July 2022 through August 2026: 754 articles totaling 713,163 words.
Our conclusions come from looking at her entire body of work, not just a few egregious examples.
First, I programatically scraped Erin Reed’s substact via the Archive feed Substack puts out.
Then, I filtered this collection of everything EITM ever published to only everything Reed wrote on her solo byline.
Then, I looked for tells of AI writing, first using the Wikipedia editors’ tells of AI and a standard quantitative battery in quanteda. I learned how to use quanteda in my Master’s program. That Master’s thesis relied on the same core packages in the programming language R, but applied in different ways. Once I felt there was some smoke, such as the changes to the average reading level (Flesch Kincaid scores), I dug deeper.
In addition to running the various quantitative tests on my workstation (both as preparation for this investigation and in the initial phases), I reviewed literature from various academics and scientists on how to detect AI mixed writing, or if it can even be done. A few different papers became central to my methodology. They are cited in the original article, and at the end of this one.
The tests I ran included recreating multiple papers’ methodologies to the best of my ability. I also looked for simple tells, like punctuation (such as the rate of hyphens versus em dashes) and spent 50+ hours fact checking every one of Erin’s 100+ articles published from January to August 2026. The reason this time period was given particular attention was due to the changes in her writing style which became observable in February 2026. Particularly, the pasting once again became obvious in Feb 2026 after the release of Claude Opus 4.6.
In this case, pasting went from not obvious to obvious because of the synthetic corpus we ran tests against confirming how often different chatbots use different punctuation marks.
Finally, most of these p-values are so large because we have a massive sample size available to work with. When you have something like 750,000 words and hundreds of articles, the p values for something that is actually a trend within such a large dataset gets astronomical. P-values are the probability that random chance or something unobserved could explain the hypothesis equally well as the tested explanation.
Every p value below indicates the odds of the same author writing in the same style and making such different groups of work. For example, the difference between pasted sentence or paragraph lengths at this sample size leads to such a p value.
For more details on how to use this methodology to investigate other public figures for robotic plagiarism, please contact our team at editors<at>theneedlenews<dot>com (or click below), as we are willing to replicate our process on a media consulting basis for mission-aligned groups.
click to show our email addressResults and P values with this methodology are below:
Our Source Data
The Corpus
Population: Every solo-byline article at Erin in the Morning. 754 articles, 713,163 words.
Window: July 2022 through August 2026, the last month with complete data.
Her Work Pre-ChatGPT
ChatGPT reached the public on Nov. 30, 2022. Everything Reed wrote before that is our 100% human baseline.
Window: July through November 2022 (5, 3, 2, 1, then 11 articles by month).
Role: Treated as 100% human, and used as the control for every later era.
We Compared Against Her Employee
S. Baum is the other full-time writer at Erin in the Morning. Baum gives us a second known author to measure Reed against.
Window: Restricted to the period Baum and Reed worked together.
Use: Second-author reference for the stylometric tests.
The Five Eras
The drift tests split Reed’s career into five stretches:
- Before ChatGPT
- December 2022 through 2023
- 2024
- 2025 to Feb. 5, 2026
- Feb. 9, 2026 to the end of the collection period
Stuff She Pasted
How We Knew It Was Pasted
A straight quotation mark (U+0022) and a straight apostrophe (U+0027) are different characters from the curly, aka “smart”, versions that most word processors insert on their own. Substack, where Reed publishes, curls them by default. A straight mark that survives to the page came from somewhere else.
The Presence Rule
Pasted: A paragraph holding any straight mark.
Typed: Every other paragraph.
We also ran a stricter test, counting a paragraph as pasted only when the majority of its marks were straight. It returned largely the same results. For the side-by-side length and reading-level comparisons, we used only paragraphs that were fully straight-marked or fully curly, never a mix of the two.
Pasted = Different
Population: Pasted vs. typed paragraphs and sentences, across the full corpus and inside the window after Feb. 9, 2026.
- -Pasted paragraphs run 134 words to typed paragraphs’ 53, a gap of about 81 words.
- -In 2026, sentences in pasted paragraphs are 12.51 words longer and read 5.56 grade levels harder.
- -Since Feb. 9, 2026, a pasted sentence sits at reading level 16.83 against a typed sentence’s 12.55 (up 4.28 grades), and runs 27.58 words long against 15.88 (up 73.7%).
- -In that window, 65.7% of her published words fell in pasted paragraphs. Of the 642 paragraphs from Jan. 1 to the Claude Opus 4.6 release, one was clearly pasted; ten articles after it ran 99% or more of their words in pasted paragraphs.
Some Odds Are Smaller Than The Universe Is Large
Paragraph length: p = 7.0e-126, ~449 heads in a row. The likelihood that the two sources of writing use their different paragraph lengths by random chance.
Reading level: p= 4.51e-199, ~659 heads. The likelihood that the two sources of writing write at their varying reading levels by random chance.
Sentence length: p= 4.63e-211 ~700 heads. The likelihood that the two sources of writing use their different sentence lengths by random chance.
Reed v. The Machines
Which Machines Agree?
Tier 1: live/web interfaces (Sept. 2026): we gave one identical prompt to Gemini, Grok, ChatGPT, Claude, and Perplexity. Current Gemini and Claude both produced straight marks every time.
Tier 2: dated corpus: each model wrote 1,000 documents, each seeded with a headline drawn from a fixed set of 50 real Reed headlines matched to the model’s year [1].
Models: ChatGPT (2023-26), Claude (2024-26), Gemini (2024-26). No 2023 Claude or Gemini could still be run.
Result:
- -Gemini used straight marks less often than Reed does.
- -ChatGPT flipped from straight to curly between its 2024 and 2025 models.
- -Reed’s own straight marks disappear in 2025.
- -Claude output straight marks 100% of the time across each year measured.
- -Her marks return inside a four-day window from Feb. 5-9, 2026, the days surrounding the launch of Claude Opus 4.6.
Style: Personable or Robotic?
Small grammatical words shift when human and machine prose mix [2]. We tracked them across Reed’s career.
Features: “This” and “these” fall, while “the” rises; “would” falls while “or” and “but” rise; depersonalizing language (“its” in place of “their”) climbs. Each shift is statistically significant.
Punctuation only: This could distinguish the eras with 93.4% certainty, p= 2.68×10⁻¹¹⁰, ~364 heads.
Words + punctuation: Separated 2022-23 from 2025-26 at 94% across 274 articles a side, p= 7.64×10⁻¹¹⁴, ~376 heads.
Self-drift: Across all features, p= 4.95×10⁻⁸⁷, ~287 heads
Distance to Claude
We built this one to give Reed every benefit of the doubt, then measured how close her writing sits to fully-machine text [2].
Human baseline: Writing at The Needle: Artemis’s byline, Jane’s byline, and our joint byline.
Machine baseline: Claude Opus 5 rewrites of each 2025-26 Reed article, under two prompts: one told not to sound like Reed, and one given only her headline and word count.
Result:
- -Her writing edges closer to Claude with each era.
- -Two of her 2025-26 articles read as more Claude-like than anything Claude wrote for us; 21 beat the average Claude output, all after Feb. 5, 2026.
- -Her 2023-25 work matched the era’s ChatGPT more than Claude.
Workflow fit: A “ChatGPT 2023-24, no AI in 2025, Claude from Feb. 2026” pattern is 1.4% as likely as the best fit, “latest ChatGPT through 2025, then Claude in 2026.”
Author Drift and Fingerprinting
This is the same method of stylometrics used by courts and journalists across the world.
Burrows’s Delta
A classical authorship-attribution distance, the same family of method that unmasked Robert Galbraith as J.K. Rowling and helped identify the Unabomber [3].
Reed vs. Reed: Her 2022-23 writing separates from her 2025-26 writing 93.5% of the time.
Reed vs. Baum: Over the window they shared, the two separate 97.7% of the time.
Distance: Reed in 2025-26 sits farther from her own 2022 self than from Baum today, and each era lands farther from the pre-ChatGPT baseline than the last.
Machine-Infected Writing
She’s Not a Professor, Right?
We scored every month on the Flesch-Kincaid grade level [4].
Before: Monthly average 11.28, a high-school junior.
After April 2023: Monthly average 14.23, a college junior/senior.
Peak: 15.71, 4.64 grades above her pre-LLM baseline. Six separate months landed at postgraduate difficulty.
So. Many. Words.
Sentence length: From about 17 words in her first six months, to about 23 in her latest six, up 33.6%.
Pasted vs. typed: Pasted paragraphs run 64.4% more words, 36.6% higher grade level, and 33.6% more words per sentence.
Dash Switcharoo
Reed used to use the spaced hyphen exclusively. She switched to the em dash, the version preferred by LLMs. The shift began about two months after ChatGPT launched.
Start: Zero em dashes in her first six months. Her first came in Jan. 2023; they climbed from Feb. 2023.
Peak: 12.05 em dashes per 1,000 words in Oct. 2025, about one every 83 words. The rate never falls back below the summer-2024 or pre-ChatGPT floor.
Odds: Draw one 2022-23 article and one 2025-26 article at random, and the latter carries more em dashes 95.7% of the time.
Reed’s errors
Facts? In This Economy?
Population: Reed’s 106 solo-byline articles, January through August 2026.
Result: Errors, misstatements, or failed citations in many of them, 50+ errors in all. We found no correction or retraction for any of them.
Methods and sources
- Synthetic model corpus, generated from year-matched Reed headlines. Method writeup. Method · performed and replicable by The Needle · primary
- Function-word and topic-model separation of human and machine prose. PLOS ONE (0335369). Peer-reviewed method · ~99.8% reported separation of fully-human from fully-AI text
- Burrows’s Delta, a classical measure of stylistic difference and likely authorship. Burrows (2002). Method · standard in authorship attribution
- Flesch-Kincaid grade level. Reference. Readability standard · word and sentence length
Corpus, class definitions, and the per-test outputs are held by The Needle.
Editor's note: This was edited after release to further explain our process rather than simply use the results values.
