Back to blogPrivacy & Security

When Big Tech Stole the Library — And What's Coming Next

19 June 2026|Hynt Digital
When Big Tech Stole the Library — And What's Coming Next

If further illustration were needed of why AI ethics are not an academic exercise, it is already in the court record.

In early 2025, unredacted court documents revealed that Meta — parent company of Facebook and Instagram — had knowingly used LibGen, one of the world's largest repositories of pirated books, to train its Llama AI models. LibGen hosts more than 7.5 million books and 81 million research papers, the overwhelming majority obtained without authorisation. According to documents filed in the Kadrey v. Meta lawsuit, CEO Mark Zuckerberg approved the use of the dataset knowing it contained pirated material.

One internal Meta employee message, cited in court, put the ethical awareness plainly: "Torrenting from a corporate laptop doesn't feel right." Another noted that if media coverage suggested Meta had used a dataset it knew to be pirated, "this may undermine our negotiating position with regulators." The concern was not the theft. It was the optics.

The Authors Who Pushed Back

Authors including Ta-Nehisi Coates, Sarah Silverman, and Jacqueline Woodson were among those who brought legal action. French publishers and authors filed separate proceedings in a Paris court. The Authors Guild described the practice as "Meta's massive AI training book heist." One bestselling author, on discovering her works had been taken, described it simply as: "It's not just theft of the work — it's theft of the work to create something to replace us."

The judge who eventually ruled on the case was candid even in dismissing it on procedural grounds: the ruling, he wrote, "does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful." He indicated that the record suggested Meta and other AI companies had become serial copyright infringers in training their technology — and appeared to invite better-constructed cases to his court.

The Regulatory Direction

This matters beyond the publishing industry. If one of the world's most valuable technology corporations was prepared to torrent millions of copyrighted works to train its models — while its internal compliance team made precisely the objections you would expect — the question of what any AI vendor has used to train their systems is not a paranoid question. It is basic due diligence.

The regulatory framework is now catching up. The EU AI Act (Regulation 2024/1689), which entered full force in 2025, imposes significant obligations on providers of general-purpose AI models, including disclosure of training data sources under Article 53. Providers must maintain detailed technical documentation and make it available to the AI Office. The Act specifically addresses copyright compliance, requiring providers to operate a policy respecting EU copyright law — the same law Meta was found to have violated.

In the UK, the Intellectual Property Office's 2022 AI and IP consultation concluded that existing copyright law applies to AI training, meaning Meta's actions would likely be unlawful under UK law as well. The UK government's pro-innovation approach to AI regulation (published February 2024) places responsibility on existing sector regulators — the ICO, FCA, and others — to enforce data protection and IP rules within their domains.

The Biden administration's Executive Order on AI (October 2023), while partially rolled back by the current administration, established a precedent that AI safety testing and transparency are legitimate regulatory concerns. Several US states — including California, Colorado, and Illinois — have introduced their own AI transparency legislation requiring disclosure of training data sources.

For Welsh SMEs across West Wales — from Carmarthenshire to Ceredigion, Pembrokeshire, and Swansea — the implication is clear. The regulatory environment is shifting toward requiring organisations to know where their AI tools come from and what data they were trained on. That shift is happening faster than most businesses realise.

The practical implication is equally clear. An AI vendor who cannot tell you where their training data came from is not just an ethical risk. They are a regulatory liability waiting to crystallise.

Due diligence on AI suppliers is not yet standard practice for most Welsh SMEs. It should be. The question is not whether the regulation is coming. It is whether your business will be caught unprepared when it arrives.


Concerned about your AI supply chain? We help Welsh businesses audit their AI vendors and build privacy-first infrastructure. Get in touch or book a Discovery Audit.

Got a project in mind? Let's talk.

Whether you are just starting to explore AI or ready to build something specific, we can help. Start with a free 15-minute call to scope your idea.