The Atlantic's new tool lets you check if your work was used to train AI models
The Atlantic has developed a search tool that lets users check if their work appears in LibGen, a massive archive of pirated books, scientific papers, and articles that was reportedly used to train language models. According to court documents, Meta used the LibGen dataset to train its Llama models. OpenAI told Gizmodo that LibGen content is not included in the current versions of ChatGPT or in OpenAI's API. Other AI companies have not yet commented on whether they used LibGen data in their training. Microsoft recently began offering book licensing deals to publishers.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe nowRead on for the full picture.
Subscribe for hype-free coverage.
- Full access to every article on THE DECODER
- No ads
- Join the comments and community discussions
- A weekly AI news recap via mail
- 6x/year: "AI Radar" — deep dives on the AI topics that matter most
- Daily AI news, always up to date
- Our full ten-year archive
- Covered by a team with 10+ years in AI