[News] Anthropic settles $1.5 billion in lawsuit over book use on pirate site

economy

U.S. AI startup Anthropic has agreed to a $1.5 billion settlement in a class-action lawsuit accused of using book data obtained from pirated sites for AI training. This settlement marks a crucial turning point in the legal risks of AI training and removes barriers to the company’s goal of a large-scale IPO.

.5 billion settlement and compensation for 500,000 writers

On July 21, 2026, the U.S. District Court for the Northern District of California finally approved a $1.5 billion (approximately 240 billion yen) settlement agreement between Anthropic and the author group. This amount is historically record-breaking, setting the largest record in U.S. copyright litigation. More than 480,000 works are subject to settlement, and more than 500,000 authors and publishers, including Andrea Bartz, are entitled to compensation. Specifically, compensation is expected to be paid between about $3,000 and $3,100 per eligible book, which is expected to become an international benchmark for copyright in the future AI industry.

In this settlement, Anthropic did not admit any wrongdoing, but under the terms of the settlement, it was required to destroy all file copies obtained from the pirated book repository. This large-scale financial settlement allowed the company to avoid potential damages risks that could have reached up to $1 trillion (approximately 154 trillion yen). Please refer to the timeline below.

Figure 1

Litigation History and Actual Use of Pirate Sites

This lawsuit follows the 2024 lawsuit filed by Andrea Bartz and two other authors in the case of “Bartz v. It originated from a class action lawsuit filed under the name “Anthropic.” The plaintiffs claimed that when Anthropic trained its large language model “Claude,” it downloaded and used over 7 million books, including their own copyrighted works, from pirated book sites such as “Library Genesis (LibGen)” and “Pirate Library Mirror (PiLiMi)” without permission. The investigation revealed that the company had been scanning and digitizing books worldwide on a large scale under an internal initiative called “Project Panama.”

A particular issue was that the company not only used the data obtained from the “Shadow Library” for training but also permanently stored and stored it as the company’s “central library.” According to court documents, Anthropic also used a method of mass-scanning physical books purchased from secondhand bookstores through cutting machines, constructing an extensive digital archive without obtaining copyright holders’ permissions. The illegality of these “data acquisition routes and storage methods” was the decisive factor leading to this massive settlement.

The Intersection of Judicial Decisions and Corporate Strategy

The legal boundary of “learning is legal, obtaining is illegal”

A major factor in this case was the division judgment delivered by Judge William Alsap in June 2025. This ruling indicated that the act of using copyrighted works to train AI models falls under “transformative use” and is highly likely to be recognized as fair use under U.S. copyright law. This was seen as a major step forward for the AI industry in ensuring the legality of future developments.

However, on the other hand, Judge Alsup ruled that the act of unauthorized storage and accumulation of data obtained from pirate sites constituted copyright infringement that exceeded the scope of fair use. In other words, while it is permissible for AI to read books and “learn,” a clear line has been drawn where legal responsibility is held in the process of “where and how to obtain that learning material.” This judicial ruling serves as a strong warning to AI companies about the importance of ensuring legitimate channels for obtaining data.

Trillion-dollar risk avoidance targeting an IPO

Anthropic’s acceptance of a massive settlement is driven by strategic intentions toward an initial public offering (IPO) planned for late 2026. The company expects its market capitalization at IPO to be around $965 billion, and to gain investor trust, it needed to quickly resolve the opaque legal risks related to past copyright infringement. The $1.5 billion settlement amounts to about 3.4% of the company’s annual sales (approximately $44 billion), but compared to statutory damages that could have ballooned from hundreds of billions to $1 trillion, it is considered a reasonable choice to ensure business continuity.

With this settlement, Anthropic has successfully eliminated its biggest concerns stemming from past download and training activities. Market attention has shifted from resolving legal troubles to the fundamental question of whether the company’s technology roadmap and revenue growth potential can justify its massive market capitalization of $965 billion. However, it should be noted that this settlement is merely a disclaimer for “past acts,” and regarding the use of copyright rights in future model development, the rights holder still retains the right to file individual lawsuits.

Changes in the Generative AI Industry and Future Prospects

Spillover to Other AI Companies and Formation of the Licensing Market

The settlement with Anthropic has become an important benchmark for other major AI companies such as OpenAI, Meta, and Microsoft, which are facing similar copyright infringement lawsuits. The current “$3,000 per book” compensation level is expected to set a benchmark for settlement negotiations in similar lawsuits going forward, prompting a reconsideration of compliance strategies across the industry. As it has become clear that obtaining data from pirate sites carries extremely high legal risks, companies are shifting toward sourcing data through formal licensing agreements with rights holder groups and publishers.

A new content trading market for the generative AI era is already taking shape, with OpenAI partnering with the Associated Press and Meta signing data usage agreements with Reuters. Going forward, the “brutal growth phase” of collecting data without permission and responding afterward will end, and a business model based on licensing, coexistence, and mutual prosperity will become an essential industry standard. This is expected to establish a framework for creators to receive compensation when their works are used by AI, leading to a new phase that balances technological innovation and rights protection.

New challenges in strengthening regulations and ensuring safety

While copyright issues are being resolved, the regulatory environment surrounding AI safety and transparency is becoming increasingly stringent worldwide. In the United States, executive orders require government submission of evaluation reports before the release of powerful AI models, thereby strengthening monitoring of the development process. At congressional hearings, there is growing demand for the disclosure of specific sources for training data and model evaluation criteria, and if these regulatory proposals pass, they could become a substantial constraint for all AI companies.

Anthropic itself has pointed out the risk of humanity losing control and suggested that development should be paused for “Recursive Self-Improvement,” where AI autonomously designs and develops its successor models. As technology accelerates, alongside solving the economic challenge of copyright protection, security debates on how to manage systemic risks posed by AI will remain the top focus for industry and policymakers worldwide going forward.

Reference Page

  • 【Bartz v. Anthropic PBC – Order on Fair Use】https://copyrightalliance.org/court-case/bartz-v-anthropic/

  • 【When AI builds itself – Anthropic Official Blog】https://www.anthropic.com/institute/recursive-self-improvement

  • 【AI and Copyright – U.S. Copyright Office】https://www.copyright.gov/ai/

  • 【H.R.7913 – Generative AI Copyright Disclosure Act】https://www.congress.gov/bill/118th-congress/house-bill/7913

[#AI #著作権 #Anthropic #IPO #ビジネス #テクノロジー #法務 #生成AI]

コメント

Copied title and URL