Close Menu
NCIJ Network NCIJ Network
    What's Hot

    Live: Ukraine backers gather as Russian attacks escalate, air defence runs low

    August 24, 2026

    The shipping industry profits from conflict and looks East – POLITICO

    August 24, 2026

    Photo of Disraeli among missing parliamentary works of art, FoI request finds | Politics

    August 24, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Live: Ukraine backers gather as Russian attacks escalate, air defence runs low
    • The shipping industry profits from conflict and looks East – POLITICO
    • Photo of Disraeli among missing parliamentary works of art, FoI request finds | Politics
    • AIPAC Wants to Back This Republican, but He Has Some Reservations
    • US battery startups have found a lifeline in defense
    • ZK International’s $20M AWA tokens leave cash under $83K
    • US territories hit out at Trump plan to explore deep sea mining in the Pacific | Deep-sea mining
    • Der Herbst des Friedrich Merz – POLITICO
    • About
      • Our Team
      • Editorial Policy
      • Editorial Independence
      • International Support
    • Trust & Standards
      • AI Usage Policy
      • Conflict of Interest Policy
      • Corrections Policy
      • Ethics Policy
      • Fact-Checking Policy
      • Source Protection
    • Get Involved
      • Guide for Sources
      • Support Independent Journalism
    • Legal
      • Cookie Policy
      • Privacy Policy
      • Terms of Use
    Facebook X (Twitter) Instagram
    NCIJ Network NCIJ Network
    Monday, August 24
    • Home
    • World
    • Ai
    • Business
    • Politics
    • Health
    • Crypto
    • Science
    • Technology
    • Cybersecurity
    • Defense & Security
    • Economy
    • Energy
    • Europe
    • More
      • Fact Check
      • Investigations
      • Opinion & Analysis
      • Environment
    NCIJ Network NCIJ Network
    Home»Technology

    Is it legal to train AI models on copyrighted books? It’s complicated

    NCIJ NETWNCIJ NETWORKBy NCIJ NETWNCIJ NETWORKAugust 24, 2026 Technology No Comments5 Mins Read
    Share
    Facebook Twitter LinkedIn Pinterest Email

    You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right?

    The reality isn’t that simple. 

    “I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.”

    Last year, in one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay a mammoth $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models. At face value, this seemed like a moral victory favoring authors, but Judge Alsup actually ruled that Anthropic’s AI training was lawful. What Alsup penalized Anthropic for was pirating these books from illegal online shadow libraries.

    “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” the judge wrote, comparing the way an LLM ingests trillions of words to a writer’s study of literature.

    Gellis thinks the ruling is more advantageous for AI companies. What’s a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028?

    “I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work,” Gellis said. “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work.”

    Copyright law hasn’t been updated since 1976, which means that judges have to figure out how to interpret guidelines from 50 years ago when confronting legal questions that have the potential to shape the future of the AI industry.

    “Everybody is very worried right now because the law is all over the place, and it’s because of this question,” Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch. “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.”

    These questions often hinge on fair use law — namely, whether use of a copyrighted work is “transformative” enough to be considered legally permissible.

    Fair use is a carve out of copyright law that allows for the use of copyrighted materials without explicit permission, protecting the ability to comment and iterate on copyrighted works through criticism, parody, education, and other means. Judges consider specific factors when deciding if something is fair use, including the purpose and nature of the work, the amount used, and its impact on the market.

    “Copyright is always about protecting and growing the market,” Henderson noted. “The courts are kind of all over the place in their reasoning [in AI cases]. What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.” 

    Henderson is referencing a case in which the media and technology company Thomson Reuters sued the research firm Ross Intelligence for copying its content in order to build a competing, AI-based legal platform. 

    “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s,” Judge Stephanos Bibas wrote last year. 

    In that case, Judge Bibas decided that it was not fair use to train on Reuters’ content to make a new platform that would directly compete with it. While authors could potentially argue that chatbots are competing with them by using their works to generate new, synthetic books, that argument has not yet prevailed in court.

    When it comes to the relationship between AI and copyright, Gellis finds it helpful to narrow down what we’re actually talking about – the way we think about copyright in terms of AI training is quite different from how we think about copyrighting AI-generated content. 

    In one case, Thaler v. Perlmutter, the court ruled that if a work is 100% AI-generated, it’s not copyrightable, which opens a whole new can of worms – how can we definitively prove whether or not a work was generated using AI, and if so, how do we know what percentage of it was created or assisted with AI?

    “If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,” Gellis said. “[AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.”

    Most AI companies are still lodged in pending litigation over these issues, which means that we won’t have a definitive solution to these problems any time soon.

    “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail,” Gellis said. “But in the meantime, all these decisions are shaping everything that’s happening. It would be kind of foolish for the AI companies to ignore them.”

    When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

    books complicated copyrighted legal models train
    NCIJ NETWNCIJ NETWORK
    • Website

    Keep Reading

    US battery startups have found a lifeline in defense

    7 Basic iPhone Tricks I Built With iOS 27’s Revamped Shortcuts App

    Flock CEO calls for ‘compromise’ as surveillance company faces growing backlash

    Why students are being paid £2,000 to play computer games

    The 6 Best Laptop Docking Stations to Unlock the Full Desktop Experience (2026)

    TechCrunch Mobility: The custom chip driving Waymo’s robotaxi ambitions

    Add A Comment
    Leave A Reply Cancel Reply

    Editors Picks

    Live: Ukraine backers gather as Russian attacks escalate, air defence runs low

    August 24, 2026

    The shipping industry profits from conflict and looks East – POLITICO

    August 24, 2026

    Photo of Disraeli among missing parliamentary works of art, FoI request finds | Politics

    August 24, 2026

    AIPAC Wants to Back This Republican, but He Has Some Reservations

    August 24, 2026
    Latest Posts

    Little Italy group, city of San Diego at ‘stalemate’ over bike lane

    July 28, 2026

    Blazing like 10 billion suns: NASA’s Swift sees a wandering black hole devouring a star

    July 28, 2026

    Apple’s App Store promoted fake Bitcoin wallet that stole $1.8M after developer spent a year warning them

    July 28, 2026

    Subscribe to News

    Get the latest sports news from NewsSite about world, sports and politics.

    NCIJ Network is an independent digital news platform delivering trusted investigative journalism, European and global news, in-depth analysis, and fact-based reporting with accuracy, transparency, and integrity.

    Facebook X (Twitter) Instagram Pinterest YouTube

    Live: Ukraine backers gather as Russian attacks escalate, air defence runs low

    August 24, 2026

    The shipping industry profits from conflict and looks East – POLITICO

    August 24, 2026

    Photo of Disraeli among missing parliamentary works of art, FoI request finds | Politics

    August 24, 2026

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Type above and press Enter to search. Press Esc to cancel.