The Bill Comes Due for AI's Data Feast
What makes an artificial intelligence model "smart"? The secret ingredient isn't just advanced mathematics or massive computing power—it is the sheer volume of...

What makes an artificial intelligence model "smart"? The secret ingredient isn't just advanced mathematics or massive computing power—it is the sheer volume of human-written text the system consumes. But what happens when the humans who wrote that text demand to be compensated?
OpenAI and Microsoft are facing yet another major legal challenge that strikes at the very foundation of how generative AI is built. The Seattle Times and Newsday have filed lawsuits against the tech giants, alleging severe copyright infringement. According to the publishers, OpenAI ingested their copyrighted journalism to train its AI models without seeking permission or offering compensation.
Crucially, the lawsuit alleges that this isn't just about AI learning abstract concepts. The publishers claim that chatbots powered by OpenAI's technology frequently reproduce exact passages from their reporting in response to user prompts. Because Microsoft's Copilot assistant is deeply intertwined with OpenAI's underlying models, Microsoft has been named as a co-defendant.
This legal action represents a significant escalation in the ongoing battle between content creators and tech platforms. The Seattle Times and Newsday are not fighting alone; they join a growing coalition of nearly 400 local newspapers that have recently sued the two companies. This wave of litigation follows high-profile lawsuits from national heavyweights like The New York Times, as well as reference giants like Merriam-Webster, Ziff Davis, and Encyclopedia Britannica.
The core of the dispute revolves around the legal concept of "fair use." AI companies generally argue that training models on publicly accessible internet data is transformative and legally permissible—akin to a human student reading hundreds of books at a library to learn how to write. Publishers, however, view it as commercial exploitation. They argue that AI companies are building multi-billion-dollar businesses on the backs of journalists, creating products that directly compete with the original news sources by answering user queries without driving traffic back to the publishers' websites.
For local journalism, which has already spent two decades battling declining revenues in the digital age, this dynamic poses an existential threat. If AI tools effectively replace the need for users to visit news sites, local newsrooms may lose the advertising and subscription revenue necessary to fund their investigations.
Ironically, the AI industry relies heavily on this very ecosystem. Without a continuous stream of high-quality, fact-checked reporting produced by human journalists, AI models risk degrading in quality. How the courts ultimately interpret copyright law in the age of generative AI will not only determine the financial liabilities of tech giants, but also shape the future economic model of journalism itself.
Key Points
- The Seattle Times and Newsday are suing OpenAI and Microsoft over alleged copyright infringement.
- The publishers claim AI models were trained on their articles without permission and sometimes regurgitate exact passages.
- Nearly 400 local newspapers have joined the legal pushback against AI data scraping.
- The lawsuits challenge the tech industry's reliance on the "fair use" doctrine for training generative AI models.
Why It Matters
The outcome of these lawsuits will define the legal boundaries of AI training data, potentially forcing tech companies to fundamentally change how they source information and compensate creators.
Sources:
- Seattle Times and Newsday sue OpenAI and Microsoft for infringement — The Verge - AI
更多专栏

Your Next Coworker is a Blob That Orders Burritos
For decades, enterprise software has been synonymous with sterile dashboards, en...

The Midnight Bill: Why AI Agents Demand Hard Budget Caps
The dream of artificial intelligence is to have a tireless digital assistant wor...

Beyond Transformers: How Mamba is Rewriting the Rules of AI Memory
Think about how a human reads a sprawling, thousand-page fantasy series. You don...