NVDA 225.30 ▲0.54%GOOGL 346.36 ▲0.82%MSFT 496.88 ▲0.90%AMD 483.01 ▲0.02%INTC 104.56 ▲3.58%TSMC 430.49 ▲0.31%AMZN 265.13 ▼0.80%META 594.97 ▲2.78%AAPL 305.26 ▲1.00%PLTR 179.01 ▲4.66%
Markets at last close

Policy

Generative AI strains copyright law across training and retrieval

·1 min read

A UCL workshop on generative AI and copyright brought together legal scholars, computer scientists, and practitioners to examine data collection, model training, and post-training systems such as RAG. Participants identified a widening gap between copyright concepts and technical practice, with uncertainty around reproduction, lawful access, originality, and liability.

Lawful access emerged as a central concern under Articles 3 and 4 of the EU Copyright in the Digital Single Market Directive and section 29A of the UK Copyright Designs and Patents Act 1988. The concept remains underdefined, especially when AI developers scrape freely available online material that may itself be infringing. Infrastructure controls such as bot blocking and services including Cloudflare further complicate whether data is legally accessible.

Territorial copyright rules also sit uneasily with AI development, where training, hosting, and deployment can occur in different jurisdictions. Participants noted that models encode statistical relationships rather than literally store works, but outputs that reproduce recognisable material may still raise infringement questions. RAG adds distinct risks around caching, vectorisation, chunking, indexing, and leakage, while proposed governance tools such as filtering, machine unlearning, licensing, and collective bargaining all carry trade-offs for innovation, competition, and expression.

Originally reported by legalblogs.wolterskluwer.comRead the source →
Related coverage