Training non-generative AI on copyrighted material is not fair use

While I usually focus this blog on trade secrets and restrictive covenants, given the significance of artificial intelligence and its overlap with trade secrets (as well as my general interest in AI1), I plan to cover some of the significant legal developments involving AI moving forward, even if they do not involve trade secrets.

Today’s is Thomson Reuters Enterprise Center GmbH v. Ross Intelligence Inc., 2025WL458520 (D. Del. Feb. 11, 2025). The decision, authored by Circuit Judge Bibas, sitting by designation, is about the intersection of non-generative AI, copyright, and fair use. 

There is a lot here, so go ahead and jump straight to the takeaways at the end if that’s all you need.  

Background

The court’s opening quote is fantastic:

A smart man knows when he is right; a wise man knows when he is wrong. Wisdom does not always find me, so I try to embrace it when it does — even if it comes late, as it did here.

Gotta love judges like that! (The context was that the judge had issued a prior decision that he was now fixing.) 

The judge set the stage as follows:

The law is no longer a brooding omnipresence in the sky; it now dwells in legal-research platforms. Thomson Reuters owns one of the biggest of those platforms: Westlaw. Users can pay to access its contents, including “case law, state and federal statutes, state and federal regulations, law journals, and treatises.” “Westlaw also contains editorial content and annotations,” like the headnotes here. Those headnotes summarize key points of law and case holdings. Westlaw organizes its content using the Key Number System, a numerical taxonomy. Thomson Reuters owns copyrights in Westlaw’s copyrightable material.

Ross, a new competitor to Westlaw, made a legal-research search engine that uses artificial intelligence. To train its AI search tool, Ross needed a database of legal questions and answers. So Ross asked to license Westlaw’s content. But because Ross was its competitor, Thomson Reuters refused.

So to train its AI, Ross made a deal with LegalEase to get training data in the form of “Bulk Memos.” Bulk Memos are lawyers’ compilations of legal questions with good and bad answers. LegalEase gave those lawyers a guide explaining how to create those questions using Westlaw headnotes, while clarifying that the lawyers should not just copy and paste headnotes directly into the questions. LegalEase sold Ross roughly 25,000 Bulk Memos, which Ross used to train its AI search tool. In other words, Ross built its competing product using Bulk Memos, which in turn were built from Westlaw headnotes. When Thomson Reuters found out, it sued Ross for copyright infringement.

Recognizing that it may have erred in an earlier summary judgment decision, the court allowed the parties to move for summary judgment a second time. The court explained that the “dispute boils down to whether the LegalEase Bulk Memo questions copied Thomson Reuters’s headnotes or were instead taken from uncopyrightable judicial opinions.”

In deciding the issue, the court noted that it needed to “compare the Bulk Memo questions, headnotes, and opinions side by side.” So, to avoid revealing confidential information about the questions in a public filing, the court created its own example: 

Copyright infringement

On the issue of copyright infringement, the court explained several copyright principles applicable to the claim: 

  • A plaintiff “must show both that (1) it owned a valid copyright and (2) [defendant] copied protectable elements of the copyrighted work. The second element requires showing both that (2a) [defendant] actually copied the work and that (2b) its copy was substantially similar to the work.” 
  • “Copyright registrations are ‘prima facie evidence of the validity of the copyright’ if ‘made before or within five years after first publication of the work.’” 
  • “Originality is central to copyright,” and “originality threshold is ‘extremely low,’ requiring only ‘some minimal degree of creativity . . . some creative spark.’ The key question, then, is whether a work is original, not how much effort went into developing it.”  
  • “The text of judicial opinions is not copyrightable.”
  • “‘Factual compilations’ are original if the compiler makes ‘choices as to selection and arrangement’ using ‘a minimal degree of creativity.’” 

Based on the registration of the headnotes, the court held that Thomson Reuters has a valid compilation copyright. The headnotes (both their selection and arrangement) have sufficient originality to qualify for copyright protection as a compilation. They also qualify for protection individually, “even any that quote judicial opinions verbatim . . . .”2  

The court also held that the key note system itself was protectable as a copyright. (I thought that had been decided years ago.) Instructively, the court explained, “Even if ‘most of the organization decisions are made by a rote computer program and the high-level topics largely track common doctrinal topics taught as law school courses,’ it still has the minimum ‘spark’ of originality. The question is whether the system is original, not how hard Thomas Reuters worked to create it. So whether a rote computer program did the work is not dispositive.” 

As for the copying of the work, the court concluded there were factual issues that needed to be decided by the jury. As the court noted, actual “copying means that ‘the defendant did, in fact, use the copyrighted work in creating his own.’ One can prove this directly, with evidence that the defendant copied the work, or indirectly, by showing that the defendant had access to it and produced something similar (‘probative similarity’).” “Access alone is not proof.” Based on that, the court “grant[ed] summary judgment only on the headnotes for which actual copying is so obvious that no reasonable jury could find otherwise.” 

The court then turned to “substantial similarity,” which “requires evaluating whether ‘the later work materially appropriates the copyrighted work.’” The court summarized the law on substantial similarity as follows: “The less protectable expression a work contains, the more similar the allegedly infringing work must be to it.”

Ultimately, though finding infringement of 2,243 of the headnotes, the court determined that there was still a factual question about whether any of those 2,243 headnotes entered the public domain.

Defenses to infringement 

The court made quick work of the defendant’s defenses. 

There could be no innocent infringement, given the copyright notice on Westlaw’s headnotes. There was no copyright misuse because the copyright was not weaponized to stifle competition. There was no merger of the ideas with their expression. And “the scenes à faire defense does not fit” because “nothing about a judicial opinion requires it to be slimmed down to Thomson Reuters’s headnotes or categorized by key numbers.” 

Fair use

This all brings us to the key question: Whether training defendant’s large language model on the copyrighted material was a fair use? 

Here is the court’s reasoning in a nutshell (all quotation marks and other signals of alterations omitted, but substance not altered):

Fair use is an affirmative defense, so defendant bears the burden of proof.

The court must consider at least four fair-use factors: (1) the use’s purpose and character, including whether it is commercial or nonprofit; (2) the copyrighted work’s nature; (3) how much of the work was used and how substantial a part it was relative to the copyrighted work’s whole; and (4) how defendant’s use affected the copyrighted work’s value or potential market. The first and fourth factors weigh most heavily in the analysis.

First, the court considers the purpose and character of defendant’s use, looking mainly at whether it was commercial and whether it was transformative. If Ross and Thomson Reuters use copyrighted material like the headnotes for very similar purposes and Ross’s use is commercial, this factor likely disfavors fair use.

Ross’s use is commercial. Ross admits as much. But commerciality is not dispositive. I must balance it against how different this work’s purpose or character is.

Ross’s use is not transformative. Transformativeness is about the purpose of the use. If an original work and a secondary use share the same or highly similar purposes, and the second use is of a commercial nature, the first factor is likely to weigh against fair use, absent some other justification for copying. Ross’s use is not transformative because it does not have a further purpose or different character from Thomson Reuters’s use.

Ross was using Thomson Reuters’s headnotes as AI data to create a legal research tool to compete with Westlaw. It is undisputed that Ross’s AI is not generative AI (AI that writes new content itself). Rather, when a user enters a legal question, Ross spits back relevant judicial opinions that have already been written. That process resembles how Westlaw uses headnotes and key numbers to return a list of cases with fitting headnotes. Thomson Reuters uses its headnotes and Key Number System primarily to help legal researchers navigate Westlaw and (possibly, as the parties dispute this) to improve Westlaw’s internal search tool. The parties agree that Ross and Westlaw are competitors. So at first glance, this factor looks simple. But, as Ross argues, the headnotes do not appear as part of the final product that Ross put forward to consumers. The copying occurred at an intermediate step: Ross turned the headnotes into numerical data about the relationships among legal words to feed into its AI. That makes this factor much trickier.

Ross is right that intermediate copying has been permitted under fair use factor one in analyzing computer programs. But those cases are inapt. First and foremost, those cases are all about copying computer code. This case is not. (Though Ross did computer programming, the material it allegedly copied from Thomson Reuters was not computer code.) In copyright, computer programs differ from books, films, and many other literary works in that such programs almost always serve functional purposes. So the fair-use considerations for these programs do not always apply to cases about copying written words.

Second and relatedly, these computer-programming cases about intermediate copying rely on a factor absent here: The copying was necessary for competitors to innovate (e.g., necessary to reverse engineer access to unprotected functional elements within a program or for the purpose of gaining access to the unprotected elements of software or solely in order to discover the functional requirements for compatibility). Here, though, there is no computer code whose underlying ideas can be reached only by copying their expression. The copying is not reasonably necessary to achieve the user’s new purpose.

I thus look to the broad purpose and character of Ross’s use. Ross took the headnotes to make it easier to develop a competing legal research tool. So Ross’s use is not transformative. Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today.

The nature of the original work involves focusing on the degree of creativity inherent to the work. More creative works get more protection. Westlaw’s material has more than the minimal spark of originality required for copyright validity. But the material is not that creative. Though the headnotes required editorial creativity and judgment, that creativity is less than that of a novelist or artist drafting a work from scratch. And the Key Number System is a factual compilation, so its creativity is limited. This factor favors defendant, though it has rarely played a significant role in the determination of a fair use dispute.

Third, I focus on how much of the work was used and how substantial a part it was relative to the whole. I ask whether that usage was reasonable in relation to the purpose of the copying. Courts consider both the quantity of the materials used and their quality and importance. To win on this factor, the alleged copier must not take the “heart” of the work.

Ross’s output to an end user does not include a West headnote. What matters is not the amount and substantiality of the portion used in making a copy, but rather the amount and substantiality of what is thereby made accessible to a public for which it may serve as a competing substitute. Because Ross did not make West headnotes available to the public, Ross benefits from factor three.

Factor four is undoubtedly the single most important element of fair use. I consider the likely effect of Ross’s copying on the market for the original. I must consider not only current markets but also potential derivative ones that creators of original works would in general develop or license others to develop. I also consider any public benefits the copying will likely produce. The original market is obvious: legal-research platforms. And at least one potential derivative market is also obvious: data to train legal AIs.

It does not matter whether Thomson Reuters has used the data to train its own legal search tools; the effect on a potential market for AI training data is enough. Ross bears the burden of proof. It has not put forward enough facts to show that these markets do not exist and would not be affected.

Nor does a possible benefit to the public save Ross. Yes, there is a public interest in accessing the law. But legal opinions are freely available, and the public’s interest in the subject matter alone is not enough. The public has no right to Thomson Reuters’s parsing of the law. Copyrights encourage people to develop things that help society, like good legal-research tools. Their builders earn the right to be paid accordingly. There is nothing that Thomson Reuters created that Ross could not have created for itself or hired LegalEase to create for it without infringing Thomson Reuters’s copyrights.

Based on that analysis, the court concluded that Ross’s use was not a fair use. 

Takeaways

Training non-generative AI (an intermediate use) on copyrighted material without a license is not fair use, especially when the purpose is to develop a competing product and it involves non-functional, copyright-protected text.

The decision is limited to non-generative AI, meaning that its implications may not apply to AI models that generate new content rather than retrieve and organize existing information.

And, my favorite takeaway: Those first two takeaways were created in the first instance by ChatGPT based on my draft blog post. They needed work to get them finalized, but what it provided was a good start and AI output will only improve over time. I cannot wait to see how long it takes before we are no longer needed!

_____

[1]  I was a computer science major in college and have maintained an interest in computers and technology, and my son is wrapping up a PhD in a type of machine learning, which keeps me learning.

[2]  The court “analogized the lawyer’s editorial judgment [in creating the headnotes] to that of a sculptor. A block of raw marble, like a judicial opinion, is not copyrightable. Yet a sculptor creates a sculpture by choosing what to cut away and what to leave in place. That sculpture is copyrightable. 17 U.S.C. § 102(a)(5). So too, even a headnote taken verbatim from an opinion is a carefully chosen fraction of the whole. Identifying which words matter and chiseling away the surrounding mass expresses the editor’s idea about what the important point of law from the opinion is. That editorial expression has enough ‘creative spark’ to be original. Feist, 499 U.S. at 345, 111 S.Ct. 1282. So all headnotes, even any that quote judicial opinions verbatim, have original value as individual works.” I am not sure I agree with this. While as a compilation, I would agree, I do not agree insofar as the copyright is for the individual quote. First, there would be no creativity whatsoever. And second, the result would be (in theory at least) that Thomson Reuters could sue someone who copied an individual headnote that was verbatim from the decision, which is not itself protectable. Anyway, I’m not the judge. We’ll see how this proceeds, as the court did not grant summary judgment on the verbatim headnotes.