This week saw two big decisions in the US courts relating to whether the use of works protected by copyright to train AI systems is legal or not.
Both cases related to AI companies who had been alleged to have trained their LLMs (Large Language Models) on large quantities of copyright protected works and as to whether this is ‘fair use’ in the US. Wherein the ‘fair use’ doctrine is an allowable exemption from copyright infringement and these lawsuits have both pivoted on how the respective judges made their ‘fair use’ assessment using the following four established factors[1]:
- the purpose and character of the use,
- the nature of the copyrighted work,
- the amount and substantiality of the portion used, and
- the effect of the use on the potential market for the original work (‘market harm’).
These factors are then weighed together to determine if the use of a copyrighted work is considered a fair use, or an infringement.
Case 1 – Bartz vs Anthropic
This case centred on Anthropic, an AI company with a LLM chatbot named Claude. The plaintiff, Bartz, had alleged that Anthropic had copied their books without permission as part of its model training process. It was alleged that Anthropic had downloaded over seven million books from pirate sites including Books3, LibGen and PiLiMi. That Anthropic also bought millions of print books, stripped pages from their bindings, scanned them and fed them into a central ‘data library’ to train successive versions of Claude.
In this case, Judge William Alsup ruled that using copyrighted books to train a large language model can be fair use under U.S. copyright law, describing the process as “spectacularly transformative,” comparing it to a human learning how to write by reading other people’s work. The Claude model, in his view, was not simply reproducing or substituting the original books. It was learning from them in order to generate original responses.
Importantly, the court took care to say this case did not involve Claude generating 100% ‘knockoff’ versions of the plaintiffs’ work. That, Alsup said, would be a very different outcome.
The court also heavily criticised the company’s decision to download books from pirate sites and held that building a permanent digital library of unlawfully acquired material, even if only used internally, was not fair use. So Anthropic will still have to face Bartz in court at a future date for using pirated copies of their books.
Case 2 – Meta vs Kadrey
Meta has also won, in part, against Kadrey et al — an authors’ case brought against Meta for allegedly training their ‘Llama’ LLM on the authors’ books, including on pirated copies also.
Judge Vince Chhabria has also ruled that Meta’s training was squarely fair use in this case. However, in this decision, Chhabria was at pains to say that the authors lost against Meta because they didn’t present strong enough arguments with regards to any harmful effect (test 4):
“this ruling does not stand for the proposition that Meta’s use of copyrighted materials to train its language models is lawful. It stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one”.
Chhabria also started that “companies, to avoid liability for copyright infringement, will generally need to pay copyright holders for the right to use their materials.” Claiming fair use is specific, not general. As described above, fair use under US law depends on the specific facts and circumstances of a particular case, considering those important four factors in the balance.
It seems where Kadrey failed was in explaining how what Meta was doing had had a harmful effect on the market. Chhabria even mentioned Judge William Alsup’s ruling in Anthropic from the day before and disagreed with Alsup’s dismissal of market effects as much as he did. Chhabria further distinguished his view from Alsup’s by stressing that Alsup was “brushing aside” the importance of ‘market harm’ in his fair-use ruling, by focusing too much on whether the use of the work was “transformative.”
So even with these two judgements, it’s still not clear what weighting should be given to each of the 4 points of any fair use assessment. Further, these recent rulings on fair use follow not just from rulings over the past year, but from Author’s Guild versus Google[2], where authors sued Google over scanning their books for Google Books — and Google’s scanning was ruled solidly ‘transformative’ and was upheld on appeal.
Conclusions
So, it looks like training AI on books is likely to be allowed as fair use now, whether or not that ‘feels right’ morally or ethically. However, both the Meta and Anthropic rulings are district court decisions, and so they’re appealable and as mentioned, there are certain discrepancies in the way that each the decision was reached. Earlier this week also, a new class action suit has been filed, Bird versus Microsoft [3] – which is specifically about Microsoft training their in-house AI models on pirated copies. The class action complaint filed by several authors and professors, including Pulitzer Prize winner Kai Bird, and Whiting award winner Victor LaVelle, and claims that Microsoft ignored the law by downloading around 200,000 copyrighted works and feeding it to the company’s Megatron-Turing model.
Conversely, ‘legit’ licensing deals are being done, but at the moment they are tending to be from major corp to major corp, with Bloomberg reporting last year that Microsoft and publishing giant HarperCollins signed a content licensing deal where the tech giant could use some of HarperCollins’ books for AI training. While AI search engine Perplexity, has also launched a revenue sharing platform with publishers after receiving backlash. Meanwhile OpenAI has a content-sharing deal for ChatGPT with more than 160 outlets in several languages.
A trend is certainly emerging therefore, and unfortunately, it’s not a good trend for small or individual authors of original copyrighted works. Although in both cases reported on, the potential for future action and significant damages for authors remains over the use of pirated works.
So, watch this space and if you wish to understand how developments in this area of law might impact your business, please do feel free to reach out and contact us.
[1] https://www.copyright.gov/fair-use/
[2] Authors Guild, Inc. v. Google, Inc. – Wikipedia
[3] https://www.reuters.com/sustainability/boards-policy-regulation/microsoft-sued-by-authors-over-use-books-ai-training-2025-06-25/
