Are Chatbots Harvesting Unpublished Math?
Every time you type a prompt into an artificial intelligence chatbot, you are not just asking for help—you might also be providing free tutoring. For most...

Every time you type a prompt into an artificial intelligence chatbot, you are not just asking for help—you might also be providing free tutoring. For most everyday tasks, this invisible exchange goes unnoticed. But when the user is a world-class mathematician inputting unpublished research, the stakes change dramatically.
A brewing conflict between the academic mathematics community and OpenAI highlights a critical vulnerability in how modern AI models are trained. Recently, mathematician Andreas Thom took to the social platform Mastodon to voice a serious concern: he suspects that OpenAI’s recently touted breakthroughs in mathematical reasoning may have been fueled by the unpublished work he and his colleagues fed into ChatGPT during casual interactions.
Thom is not the first to raise this alarm; he is the second researcher in a matter of days to accuse the AI giant of lacking transparency and potentially crossing ethical lines regarding its data sources. The core of the dispute lies in the opaque nature of AI training pipelines. Historically, companies developing large language models have used interactions from their users to refine and train future iterations of their software. While users can often opt out, the default settings and the sheer complexity of data harvesting leave many unaware that their "private" brainstorming sessions might become part of a global neural network.
For a discipline like mathematics, where progress relies on years of solitary or small-group labor culminating in a published proof, the idea that an AI might absorb and regurgitate uncredited, unpublished theories is alarming. The researchers are demanding proof that OpenAI did not scrape their intellectual property to build its latest mathematical models. However, because the exact datasets used to train these models remain a closely guarded corporate secret, verifying these claims is nearly impossible for outsiders.
This controversy extends far beyond a few disgruntled academics. It touches on the fundamental sustainability of AI development. Artificial intelligence requires a constant diet of high-quality, human-generated data to improve. If the creators of that data—whether they are mathematicians, writers, or software engineers—feel that their intellectual property is being strip-mined without consent or credit, they will simply stop engaging with these platforms.
To maintain the rapid pace of innovation, the tech industry will eventually need to replace its "black box" approach with verifiable transparency. Until then, the smartest minds in the world might decide that the safest place for their best ideas is offline.
Key Points
- OpenAI is facing backlash from mathematicians over the data used to train its mathematically capable AI models.
- Researchers, including Andreas Thom, suspect their prompt interactions containing unpublished work were harvested by OpenAI.
- The lack of transparency regarding AI training datasets makes it impossible for academics to verify if their IP was used.
- This dispute highlights a broader tension between AI companies' need for high-quality data and users' right to protect their intellectual property.
Why It Matters
The "black box" nature of AI training threatens to alienate the very experts whose knowledge is needed to advance the technology. If academics cannot trust chatbots with unpolished ideas, the pace of both human and artificial innovation could suffer.
Sources:
- Mathematicians want proof OpenAI didn’t use their work — The Verge - AI
更多专栏

Your Next Coworker is a Blob That Orders Burritos
For decades, enterprise software has been synonymous with sterile dashboards, en...

The Midnight Bill: Why AI Agents Demand Hard Budget Caps
The dream of artificial intelligence is to have a tireless digital assistant wor...

Beyond Transformers: How Mamba is Rewriting the Rules of AI Memory
Think about how a human reads a sprawling, thousand-page fantasy series. You don...