Can Printed Book Warnings Prevent AI Training? Can a Line on the Copyright Page Block Large Language Models?
In March 2026, New Translation of Huang‑Ting‑Jing and Yin‑Fu‑Jing, published by Huaxia Publishing House, carried an unprecedented warning on its copyright page: “Reproduction of contents in this book for artificial‑intelligence training is prohibited. Violators will be held legally liable.” This is not an isolated case. A growing number of new publications, especially imported foreign titles, carry similar notices on their copyright pages.
Publishing‑house executives are candid: it is “very hard to enforce rights” if book contents are used for AI training without authorization. Even so, publishers keep printing such notices. “We need to make a statement, to raise public awareness that books are protected by copyright.”
Validity Analysis: Does a Printed Warning Carry Legal Weight?
Short answer: It carries certain significance yet has limited practical effect.
Liang Fei, Deputy Secretary‑General of Chinese Written Works Copyright Society, points out that warnings such as “prohibited for AI training” represent unilateral terms of use set forth by copyright holders to govern how their works may be exploited. In copyright‑law adjudication, whether a right‑holder has explicitly objected to certain uses will influence courts’ assessment of the “fair use” defence.
More notably, if such usage‑restriction notices are digitised and added to training datasets, and AI firms deliberately erase or alter those electronic rights‑management information, such conduct constitutes separate infringement under the Copyright Law of the People’s Republic of China. In that scenario, enterprises bear liability both for unauthorised training and for tampering with electronic rights‑management information.
Nevertheless, printed paper‑based notices cannot technically block web crawlers. As commentators observe: once a physical book is purchased, scanning and extracting text data is technically trivial. Warnings printed on copyright pages function more like a gentleman’s agreement.
Overseas Perspective: Three Distinct Global Governance Approaches
Countries and regions have taken divergent stances on copyright disputes stemming from AI “book‑feeding” training.
United States: Fair‑use doctrine with blurred boundaries. In June 2025, a U.S. federal court held in Bartz v. Anthropic that using lawfully purchased books for AI‑model training qualifies as fair use. Meanwhile, downloading books from pirate websites was ruled infringing. Anthropic ultimately reached a USD 1.5‑billion settlement with authors. A bright dividing line is drawn: lawful source = potential fair use; pirated source = infringement.
Japan: Permissive statutory route; no prior‑licence requirement in principle. Article 30‑4 of the Japanese Copyright Act permits exploitation of works “not for the purpose of enjoying the thoughts or sentiments expressed in works”, in principle without copyright‑holder consent. AI training falls within this category of “non‑enjoyment‑oriented use”.
European Union: Mandatory transparency for right‑holder visibility. Article 53(1)(d) of the EU AI Act, which entered into force in August 2024, requires providers of general‑purpose AI models to publish sufficiently detailed summaries of training data. The European Commission released an official template in July 2025. Relevant enforcement provisions became effective on 2 August 2026. Violations may attract fines up to 3 % of global annual turnover.
China remains in an exploratory phase. The Interim Measures for the Administration of Generative‑AI Services require training data to have legal origins and not to infringe intellectual‑property rights. The 15th Five‑Year‑Plan Outline explicitly calls for establishing a rational‑use regime for AI‑training data. However, current Chinese legislation contains no direct, concrete provisions addressing copyright issues of AI training datasets.
Industry Dilemma: From Unilateral Notices toward Institutional Arrangements
Huaxia Publishing House’s warning sparked public discussion precisely because it highlights a core governance challenge: legislation lags behind technology.
During AI training, texts are tokenised and converted into machine‑oriented digital code; traces of the original complete text largely disappear. This makes evidence collection and trace‑back extremely difficult for rights‑holders. Meanwhile, some AI companies source “clean corpora” by bulk‑buying second‑hand physical books, cutting spines for high‑speed scanning and then destroying the original volumes. Media reports estimate around two million physical books have been consumed as one‑off consumables for model training.
Some argue copyright protection will restrain AI technological innovation. A healthy industry ecosystem, however, requires continuous output of original works as well as compliant, remunerated access by AI developers. Printed “no‑AI‑training” notices express the publishing sector’s stance and underline the baseline that AI‑training datasets ought to obtain proper authorisation.
Moving from “expressing an attitude” to “building institutions” requires a complete governance system.