Ad
Skip to content

AI-generated books are flooding Amazon and tanking sales for human authors

Image description
Nano Banana Pro prompted by THE DECODER

An analysis of more than 14,000 self-published Amazon books shows that AI-generated titles are displacing human authors through sheer volume, not quality. Revenue per book is falling even for titles where no AI text was detected.

The researchers analyzed 14,419 randomly selected self-published e-books released between January 2023 and March 2026. For each title, they pulled daily sales figures from an internal dataset maintained by one of the five major US publishers. That dataset tracks about 500,000 Amazon titles and covers roughly 95 percent of all e-books sold daily on the platform, according to the researchers.

Unlike earlier studies that tried to detect AI text from short book previews alone, the researchers classified each book based on its full text. They used the Pangram v3.3 detector, whose developers report a false-positive rate of 0.04 percent. Pangram 4 has since been released. Books were sorted into three bands based on the share of text flagged as AI-generated: none, light (up to 25 percent), and substantial (over 25 percent).

Big catalog presence, below-average sales

Books with substantial AI content make up 20 percent of the catalog studied but account for only 12.1 percent of sales and 11.3 percent of revenue. Books with no detected AI text represent 62.9 percent of the catalog and generate 72.5 percent of revenue. At first glance, that seems to confirm the view that AI books remain low-quality "slop" stuck at the bottom of the market.

Three-part chart on AI text in self-publishing: line graphs show the rising share of new titles with over 25 percent AI text from 2023 to 2026 across eight genres, bar charts break down sales-rank tiers by AI-text band, and a slope graph compares title, sales, and revenue shares across the three AI bands.
Books with substantial AI content make up 20 percent of the catalog but account for only 12.1 percent of sales and 11.3 percent of revenue. | Image: Chakrabarty et al.

The study shows that this view misses the real market dynamics, though. Between Q1 2023 and Q1 2026, the cumulative catalog grew 38.3x while the number of titles selling per quarter grew 19.2x. Quarterly revenue only grew 8.9x. Far more books are now competing for a revenue pool that's growing much more slowly.

Revenue per book is falling even for titles with no detected AI text

In six of eight genres, revenue per book dropped when comparing titles released in 2023 and 2025 over the same post-release window. Looking only at books with no detected AI text, revenue fell in seven of eight genres.

That rules out the explanation that the average is dropping simply because poorly selling AI books are padding the catalog. The authors call this effect "dilution" but stress that their comparisons are observational and associational, not experimental proof of causation.

Four charts on revenue per book: growth curves for catalog, selling titles, sales, and revenue relative to 2023; revenue per title in the launch window by release cohort; and two genre-level comparisons of the 2023 and 2025 cohorts.
The catalog grew 38.3x while revenue grew only 8.9x, and revenue per book fell for titles with no AI text in seven of eight genres. | Image: Chakrabarty et al.

The only exception is Fantasy/Supernatural/Horror, where AI text arrived latest and gained the least traction. There, revenue per book for titles with no detected AI text rose 35 percent. The researchers say this reversal argues against a general market trend as the cause.

The patterns are stronger in genres with high Kindle Unlimited availability, where readers draw from a shared subscription pool. In those genres, the revenue-share lead of books with no detected AI text is 8.4 percentage points smaller than in genres with low Kindle Unlimited availability. The researchers attribute the gap to genre-specific traits and don't draw a causal link to Kindle Unlimited itself.

AI books are breaking into the top ranks

The share of new Top 25 entries with substantial AI content rose from near zero to 31 percent over the study period. The top ranks also turned over faster. The share of books with no detected AI text that stayed in the Top 25 from one quarter to the next dropped as low as about 28 percent at one point before settling around 62 percent by the end of the study.

Four charts on AI books' market position: sales share by AI-text band over time, area chart of Top 25 rank slots, heatmap of Top 25 shares by genre, and curves for bestseller retention and new entrants.
The sales share of books with no AI text fell from nearly 100 percent in early 2023 to about 60 percent in Q2 2026, while AI titles account for up to 31 percent of new Top 25 entrants. | Image: Chakrabarty et al.

Output is concentrated among a handful of prolific producers. Of 385 author identities that published more titles with substantial AI content after their first AI book, 287 increased their monthly output afterward. The highest-grossing pseudonym earned $1.7 million in gross revenue before platform fees across eight titles. The single highest-grossing book with substantial AI content brought in $643,000 on 80,431 copies sold.

The study echoes the New York Times report on "Coral Hart," who reportedly published over 200 romance titles under 21 pen names in a single year and sold about 50,000 copies. AI spam extends beyond books, too. One man scammed millions of dollars through streaming platforms using AI-generated songs.

Six charts on AI exposure and author output: Top 25 share of books with no AI text by catalog exposure, comparison by Kindle Unlimited availability, scatter plot of monthly book output before and after AI adoption, and three bar rankings of output and revenue for the most successful author identities.
Of 385 author identities, 287 ramped up output after their first AI book. The highest-grossing identity earned $1.7 million in gross revenue across eight titles. | Image: Chakrabarty et al.

Top-selling AI books overlap heavily with rare language from existing works

To measure how much successful AI books overlap with rare language from existing works, the researchers used the Allen Institute for AI's infini-gram tool together with the Google Books index.

They identified rare expressions that appear in five or fewer Google Books volumes and are completely absent from a 4.7-trillion-token web snapshot. Phrases like these point to language patterns closely tied to published books.

Three charts on overlap with rare expressions from existing books: bar comparison for Top 50, 100, and 200 by revenue, scatter plot of revenue versus coverage with regression lines, and density curves for AI books, self-published bestsellers, and award-winning literature.
Top-selling AI books draw far more on rare expressions from existing books than award-winning literature, which has a coverage rate of just 19.1 percent. | Image: Chakrabarty et al.

Among the 50 highest-grossing titles with substantial AI content, these rare expressions covered 45 percent of the text, compared to 37.7 percent for the top 50 books with no detected AI text. For award-winning or award-nominated fiction, the figure was just 19.1 percent.

Within AI books, overlap rose 7.6 percentage points for every tenfold increase in revenue. No such correlation existed for books with no detected AI text. The method doesn't trace the origin of individual passages or prove that any specific book was copied. It measures aggregate language overlap.

Chakrabarty told Ai2 that an AI detector only returns "an estimate—a score for how likely a passage is to be synthetic," without pointing to where the language came from. If a suspect text also contains rare expressions that don't appear on the web and show up in only a handful of books, one "can say with some confidence that it was taken from books." That kind of evidence "acts as circumstantial evidence that supports an AI detector score" and at the same time "helps debunk some hackneyed arguments that liken human reading of books to AI being trained on books" – a standard line of defense from AI companies in copyright disputes.

Findings feed directly into ongoing copyright cases

The results have immediate implications for copyright lawsuits against AI companies. In Kadrey v. Meta, Judge Vince Chhabria ruled in Meta's favor in June 2025 but sent a strong warning. He said it was hard to imagine that using copyrighted books to build a product generating billions in revenue while producing a potentially endless flood of competing works would qualify as fair use.

The plaintiffs had presented no empirical evidence of this market dilution, though. The current study now delivers exactly the kind of evidence that was missing, showing that books with no detected AI text earn less as AI titles enter the market in large numbers.

Making the legal case harder is the fact that AI content can't be identified on the platform itself. Authors must disclose AI involvement when publishing through Kindle Direct Publishing, but Amazon doesn't pass that information along to customers. The platform has struggled for years with AI titles that hijack the names and styles of well-known authors and has mainly responded by capping publications at three per day.

The language overlap findings align with a November 2025 study showing that language models can reproduce passages from copyrighted books nearly word for word. Also in fall 2025, another paper showed that just two books are enough to fine-tune a model on an author's style.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: "AI Radar" — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI
Subscribe to The Decoder