A coalition of major publishers and authors has filed a class action lawsuit against Google, accusing the tech giant of using copyrighted books without permission to train its Gemini artificial intelligence models.
The lawsuit represents the latest escalation in the growing legal battle between the publishing industry and AI developers over the use of copyrighted content for training generative AI systems.
Major Publishers Join the Legal Action
The plaintiffs include Hachette Book Group, Cengage Learning, Elsevier, and bestselling author Scott Turow. The lawsuit was filed in the U.S. District Court for the Southern District of New York on behalf of a proposed class of authors and publishers.
The group is seeking an injunction to stop the alleged practices, along with statutory damages for what it describes as widespread copyright infringement.
Google Books at the Center of the Dispute
According to the complaint, Google allegedly trained its Gemini AI models using books that publishers had previously provided for the Google Books project.
The publishers argue that the original agreement only allowed Google to digitize books and display limited snippets in search results. They claim it did not authorize the company to repurpose those works as training data for commercial AI products.
The lawsuit also alleges that Google was aware of the potential legal consequences. An internal company document cited in the filing reportedly estimated that Google could face fines ranging from tens to hundreds of billions of dollars if its practices were challenged.
Publishers Claim AI Training Harms the Content Market
The plaintiffs argue that Google’s alleged use of copyrighted books threatens the publishing industry by reducing legitimate book and journal sales, enabling low-cost AI-generated alternatives, and weakening the rapidly growing market for licensed content.
They contend that AI companies should compensate rights holders rather than relying on copyrighted material without permission.
Part of a Broader Legal Campaign
The case against Google follows a similar lawsuit filed in May against Meta by many of the same publishers, including Elsevier, Cengage, Hachette, Macmillan, and McGraw Hill.
That lawsuit accuses Meta of using millions of pirated books to train its Llama AI models. Meta has denied the allegations, maintaining that training AI models on copyrighted material can qualify as fair use under U.S. law.
The parallel lawsuits indicate that publishers are adopting a coordinated legal strategy as courts continue to issue mixed rulings on AI copyright disputes.
AI Copyright Cases Continue to Grow
The publishing industry’s legal campaign comes as several high-profile copyright cases involving AI companies remain active.
Last year, Anthropic agreed to pay $1.5 billion to settle a class action lawsuit over allegations that it used pirated books to train its AI systems.
Meanwhile, The New York Times’ copyright lawsuit against OpenAI and Microsoft is still ongoing, making it one of the industry’s most closely watched legal battles.
Google Faces Growing Pressure Over Content Licensing
Unlike companies such as Microsoft, Amazon, OpenAI, Meta, and Anthropic, Google has not signed broad licensing agreements with digital publishers for AI training content.
The company does maintain a licensing partnership with The Associated Press, but publishers argue that broader agreements are necessary as AI systems increasingly rely on copyrighted material.
The dispute has intensified as media organizations search for ways to protect their intellectual property while preserving one of their largest traffic sources—Google Search.
Cloudflare Tightens Controls on AI Crawlers
Beginning September 15, Cloudflare will automatically block multi-purpose web crawlers on ad-supported pages for new and free-tier customers.
The move is widely seen as targeting Google’s crawler, which publishers say serves both traditional search indexing and AI data collection.
Adding to the pressure, USA Today Inc. CEO Mike Reed recently stated that the company could consider removing its content from Google Search within six to twelve months if no licensing agreement is reached.
A Defining Moment for AI and Copyright
After years of negotiations, crawler restrictions, and industry pressure, publishers are increasingly turning to the courts to determine how copyrighted works can be used in AI development.
The outcome of the lawsuit could set an important legal precedent for whether technology companies must obtain licenses before using books and other protected content to train future AI models.
Read the article in












