AI Ethics in Data Training and Copyright Issues

Can OpenAI, Meta, and Anthropic legally train on your work without permission? It's yet another flashpoint in the battle between innovation and regulation.
Large Language Models (LLMs) like GPT-4 are trained on vast datasets from the internet, blogs, written work, photos, and more. Chances are, some of your content is included. As GenAI grows, so do the legal and ethical dilemmas. Who owns the training data? Can copyrighted material be used without consent? Do current laws protect creators?
Much of the training data comes from scraping, automated tools that pull content from websites at scale. While it's a cheap, scalable way to collect diverse data, it triggered serious concerns about:
- Copyright infringement in AI-generated content
- Fair Use and Text and Data Mining (TDM) exemptions
- Violations of Rights Management Information under treaties like the WIPO Copyright Treaty
Artists, musicians, journalists, and other content creators argue that their work is being taken without consent, compensation, or credit. Common concerns include:
- Style mimicry: AI replicates an artist's style without infringing on specific works, but still undermines originality
- Economic harm: AI outputs displace human-created work, particularly in AI content generation for blogs, marketing materials, and creative industries
- Lack of attribution: most training datasets erase metadata, making credit or royalties impossible.
Some EU jurisdictions like France and Germany strongly protect moral rights, including the right to attribution and integrity, even after copyright is transferred.
As businesses increasingly deploy AI Agents for marketing automation workflows, understanding these IP risks becomes crucial for avoiding costly legal disputes. In any case, whether you're a developer, policymaker, creator, or user, one thing is clear: IP in the age of AI is no longer an abstract legal issue, it's a business, ethical, and societal imperative.
High-Profile Legal Cases

Authors Guild vs. OpenAI & Microsoft
A group of authors sued OpenAI for allegedly training ChatGPT on their copyrighted books without permission. The lawsuit claims copyright infringement and demands compensation. In November 2024, Microsoft was dismissed from the case and in Jan 2025, the case was consolidated with The New York Times v. Microsoft & OpenAI (2023) due to overlapping issues.
In Jan 2023, Getty Images sued Stability AI for using 12 million copyrighted images to train Stable Diffusion. Getty claimed the model even generated distorted images with its watermark, showing direct copying and it seeks $1.7B in damages. On Jan 14, 2025, the UK High Court allowed key claims to proceed, signalling major copyright implications for AI and content use.In Jan 2023, Getty Images sued Stability AI for using 12 million copyrighted images to train Stable Diffusion. Getty claimed the model even generated distorted images with its watermark, showing direct copying and it seeks $1.7B in damages. On Jan 14, 2025, the UK High Court allowed key claims to proceed, signalling major copyright implications for AI and content use.
Legal Case: Thomson Reuters Enterprise Centre GMBH v. ROSS Intelligence Inc.
On February 11, 2025, a Delaware federal court issued the first major decision concerning the use of copyrighted material to train AI. Thomson Reuters, the owner of Westlaw, sued Ross for using Westlaw headnotes—summaries of key points of law and case holdings, to train a competing, AI-driven legal research search engine. The court granted Thomson Reuters’s partial motion for summary judgement on its direct infringement claim and rejected Ross’s "fair use" defense.
Legal Fragmentation
While some data scraping is legally allowed, large-scale scraping for AI model Training often uses copyrighted or proprietary material without permission. Legal approaches differ, creating risk and uncertainty for developers. The EU AI Act requires developers to disclose training data and follow copyright rules, even if models are trained outside the EU.
A step in the right direction would be to release full Open Source models that give access to the model parameters, the code and the training data. DeepSeek R1 and Meta models are currently open weights only. Intermediaries like Common Crawl, which provides scraped web data used by OpenAI, Google, and Meta add to the complexity.
The 'Fair Use' Doctrine
Fair use allows limited use of copyrighted material without permission under certain conditions, considering purpose, nature, amount used, and market impact.
The EU doesn't have a broad, flexible "fair use" doctrine like the US, but it is moving toward supporting data-driven innovation. It has introduced mandatory Text and Data Mining (TDM) rights for research and optional rights for commercial use, giving researchers a solid legal foundation.
Singapore takes a more conservative approach. Its 'fair dealing' exception is narrower and only applies to specific purposes such as research, private study, criticism, review, news reporting, parody, satire, or education.
Text and Data Mining
TDM refers to the automated analysis of large volumes of text or data to identify patterns, trends, or relationships. Some jurisdictions like the EU and Singapore provide explicit TDM exceptions to copyright law, often under specific conditions.
Singapore
Section 243 of Singapore's Copyright Act 2021 permits use of copyrighted works for Computational Data Analysis (CDA), including AI model Training, if conditions are met:
- Purpose: must be CDA
- Lawful Access: content must be accessed legally (e.g. purchase, subscription, open access)
- Non-Commercial Use: use must be non-commercial, unless for research or news under fair dealing
- No Redistribution: content can't be reused in ways that infringe copyright
- Security: reasonable steps must prevent misuse or unlawful sharing
Additionally, the law permits contracts to override the TDM exception, meaning rights holders can still restrict TDM through licensing terms. Finally, the TDM exception only applies to copyrighted works. It does not cover other rights such as those related to Rights Management Information (RMI).
Example: A Singapore-based AI startup developing a chatbot uses TDM to analyse medical research papers. As long as they have lawful access to the papers (e.g., through subscriptions or open-access platforms) and do not bypass technical protection measures, they are legally permitted to use the data for training their AI model.
EU
Under the EU DSM Directive (2019), the EU added specific copyright exemptions for TDM, which are particularly relevant for AI developers:
Article 3: TDM for Research: law is mandatory in all EU countries but must have lawful access to the content. It allows non-commercial TDM by research organizations and cultural heritage institutions
Article 4: TDM for Any Purpose: A legal provision that allows TDM by anyone (not just researchers), including for commercial purposes like AI model Training. Article 4 is optional for member states to implement. Rights holders can refuse permission for TDM by clearly stating so.
Rights Management Information (RMI)
Rights Management Information (RMI) refers to the metadata embedded in digital content that identifies the author, copyright owner, licensing terms, and usage restrictions. RMI helps ensure that creators receive proper attribution and enables enforcement of IP rights.
Removing or altering RMI is prohibited under international IP treaties such as the WIPO Copyright Treaty and national laws in many countries. Doing so can:
- Obscure ownership or licensing information
- Prevent proper attribution
- Facilitate unauthorized or infringing use of content
Example: If an AI developer scrapes images from a stock photo website and strips the embedded metadata that names the photographer and usage restrictions, it constitutes an RMI violation. Such acts hinder rights holders from tracking usage and enforcing their rights.
For AI systems, maintaining RMI in training data is crucial for transparency, accountability, and lawful use, especially when datasets are shared or reused across projects.
Organizations implementing AI systems must navigate these complex legal waters while deploying technology effectively. The challenges extend beyond technical implementation to include comprehensive risk management around AI governance and compliance. Ideally, models would disclose their model training data details, so that it becomes a parameter for model selection. Unfortunately, none of the AI vendors provide this information, even open-weights model like DeepSeek.
Conclusion
Scraped data powers today's LLMs, but its legal status remains shaky. As lawsuits unfold, policymakers and industry are exploring softer solutions, like voluntary codes, technical tools, and dataset disclosure, to address IP concerns.
Organizations must navigate these complex legal landscapes while deploying AI systems responsibly. Handled poorly, this could lead to restricted data access, eroded trust, and harm to creators. But done right, it's a chance to build a transparent, fair AI ecosystem.
If AI models can generate profit from scraped data, should creators be compensated, even when their content was 'publicly and freely available'?
Rainmakers SG helps small and medium businesses design safe and scalable Agentic AI systems that provide immediate ROI!



