Ad
Skip to content

Microsoft faces a lawsuit alleging it used 200,000 pirated books to train AI

Microsoft is being sued by several authors who say their books were used without permission to train a Megatron model. The lawsuit, filed in federal court in New York, claims Microsoft used a dataset of about 200,000 pirated books to build a system that mimics the style, voice, and themes of the original works. The plaintiffs are asking for a ban on further use and up to $150,000 in damages per title.

Courts in similar cases involving Meta and Anthropic have said such use may qualify as "transformative" under fair use rules. But it is still unclear if using pirated books overrides fair use, or if scraping copyrighted content from the internet is considered legal and to which extent, and whether this harms the market for the original books, which could prevent the use from being considered fair use.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: "AI Radar" — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI
Subscribe to The Decoder