Five major publishers and Scott Turow sued Meta, alleging repeated copying of copyrighted books and journals for Llama training. My read is blunt: this case is more dangerous than a generic AI-training copyright suit because it attacks the data supply chain, not just model outputs. Macmillan, McGraw Hill, Elsevier, Hachette, Cengage, plus Turow is a serious plaintiff stack. The snippet names LibGen, Anna’s Archive, Sci-Hub, Sci-Mag, and other pirate sites. It does not disclose damages, docket number, venue, Meta’s response, or the exact Llama versions covered. Still, the frame is clear: publishers want to move the story from “learning from public text” to “knowingly ingesting pirated corpora.”
That distinction matters. Earlier author cases against OpenAI and Meta often ran into two hard questions: whether training-time copying is fair use, and whether outputs reproduce protected expression. Defendants usually lean on transformative use, statistical learning, and the claim that outputs do not substitute for the books. Meta has taken that line in cases like Kadrey v. Meta. This complaint appears to push a sharper factual theory. By putting LibGen, Anna’s Archive, and Sci-Hub at the center, the plaintiffs are arguing source contamination, not accidental web-scale collection. Fair use still remains the legal battlefield, but judges treat “we scraped the web” differently from “we knowingly used notorious pirate libraries.”
The part I don’t buy from Meta’s broader narrative is the clean split between “open model public good” and “training data as trade secret.” Llama 2, Llama 3, and Llama 3.1 model materials have talked about token counts, language coverage, safety testing, and release policy. They have not given a clean copyright provenance map. Meta can say that full dataset disclosure is commercially sensitive. It can also say web-scale training makes item-level licensing impractical. The problem is that the alleged conduct here is not mere accidental inclusion. The complaint, at least as summarized, says Meta knowingly ripped copyrighted work from pirate sites. If discovery produces internal emails, download logs, dataset manifests, deduplication records, or data-quality notes praising LibGen-style corpora, this becomes an engineering-records case. AI companies fear that more than screenshots of memorized text.
There is useful context from the rest of the field. The New York Times case against OpenAI and Microsoft focused on news content used in training and alleged output reproduction. OpenAI then kept signing publisher deals with groups like Axel Springer, AP, Financial Times, and News Corp. Anthropic has also faced copyright pressure from authors and music publishers. I have not verified every latest procedural update across those cases, but the direction is stable: rights holders want discovery into training sets. They no longer rely only on prompts that make a model quote text. Publishers are better positioned than individual authors here. Elsevier, for example, runs massive journal and database businesses and understands licensing economics better than almost anyone in this fight.
The open-model angle makes this messier. OpenAI and Anthropic can absorb licensing costs into API margins and enterprise contracts. Customers rarely ask what every token in the pretraining set cost. Meta’s Llama strategy is different: free weights, broad developer distribution, cloud hosting, enterprise self-deployment, and ecosystem pressure against closed labs. That distribution model amplifies the upside, but it also leaves Meta holding the training-data liability. If publishers seek an injunction or deletion remedy, enforcement would be complicated; the snippet does not say what relief they requested. Even a settlement, though, raises the cost of future Llama runs. For Llama 4 or later models, Meta then has three unpleasant options: license more publishing catalogs, rely more heavily on synthetic data, or lean harder on licensed web and first-party platform data. Each option carries either performance cost, data-skew cost, or reputational cost.
I also would not overstate the publishers’ position. US law still has no Supreme Court rule written specifically for AI pretraining. Search indexing, text mining, and book scanning have fair-use precedents; Google Books is the obvious historical comparison. The difference is that a generative model is not just a search index returning snippets. It can produce substitute text, power commercial writing products, and compress value from copyrighted catalogs into a broadly distributed model. But “copying occurred” alone may not win. The publishers still need a coherent market-harm theory. McGraw Hill and Cengage textbook markets are not the same as Hachette trade books or Elsevier journals. Those damages models should not be collapsed into one angry number.
Honestly, I would watch discovery, not statements. If plaintiffs obtain Meta records showing that LibGen or Sci-Hub data was selected because it was high quality, broad, and useful for training, Meta is in a bad position. If the evidence only shows that pirated text appeared somewhere inside a huge web-derived corpus, Meta has room to argue third-party processing, Common Crawl contamination, deduplication limits, and fair use. The body here lacks docket details, damages, exact model versions, and training windows. That limits any call on maximum exposure. For practitioners, the operational lesson is already visible: copyright litigation is moving from “does the model quote the book?” to “show me the pipeline.” Dataset manifests, scraper configs, filtering rules, and internal Slack threads may matter more than benchmark charts.