Mass downloading and vector conversion of copyrighted text, imagery, and code for analytical ingestion where the end-user query output is non-substitutive.
↳ Statutory Hook: 17 U.S.C. § 107 (Factor 1 Transformativeness)The legal characterization of high-dimensional neural network weights, embedding vectors, and diffusion model parameter checkpoints as potential derivative works.
↳ Statutory Hook: 17 U.S.C. § 101 (Derivative Works) & § 106(2)Jurisdictional challenges when AI training corpora are scraped, embedded, and trained across multinational server clusters with conflicting copyright exceptions.
↳ Statutory Hook: EU DSM Directive Art. 4 & UK CDPAValid 2nd Circuit binding precedent for non-expressive snippet search; caution required when training models produce directly substitutive commercial outputs.
Digitizing complete copyrighted books to build searchable indices and short snippet views is transformative fair use.
Authorized large-scale computational text mining, dataset curation, and search crawlers.
Binding 2nd Circuit Precedent · Heavy National Authority
“Google's mass scanning of books to create a full-text searchable database and display snippets is transformative fair use. Non-expressive computational indexing provides enormous public benefit without creating a market substitute.”
Full-book 100% digital copying without permission was direct reproduction infringement.
Google scanned 20 million library books to create keyword search indexes and 3-line snippets.
Non-Expressive Computational Mining: Copying entire expressive works for computational indexing without market substitution is transformative fair use.
Core legal justification asserted by AI foundation model developers for web scraping.
Google partnered with major research libraries to digitally scan over 20 million copyrighted books without copyright owner authorization. Google created a searchable digital database, allowing users to query words or terms and view verbatim 'snippets' of text, while also providing libraries with digital copies of books in their collections.
On appeal from the United States District Court for the Southern District of New York. District Judge Denny Chin granted summary judgment in favor of Google.
Issue: Whether Google's unauthorized mass scanning of books for searchable digital indexing and snippet display constitutes fair use under 17 U.S.C. § 107.
Authors Guild is the primary legal anchor cited by AI labs (OpenAI, Anthropic, Stability AI) to justify scraping the public internet for dataset pre-training. Penned by Judge Pierre Leval himself, it established that non-expressive computational analysis of copyrighted text is transformative fair use.
Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015).