The Unseen Labor of AI: A Theft in Plain Sight The unredacted filings in the ongoing copyright lawsuit against OpenAI and Microsoft reveal a stark reality: the use of copyrighted material to train generative AI models is not just an issue of fair use, but a brazen act of theft.
Top executives at both companies have made admissions that paint a picture of a complex web of scraping, bypassing paywalls, and stripping copyright notices – all in pursuit of training data.
The scale of the copying is astonishing, with over 91,692 copies of works published by The New York Times alone contained within OpenAI's mid training datasets.