← 深度专栏/原创观点
原创观点

The Physical Cost of AI: Scanning and Destroying Books

When we imagine how artificial intelligence learns, the process usually seems entirely virtual. We picture algorithms silently scraping billions of web pages,...

潜
作者
潜龙编辑部
关注 AI 与社会议题
发布于
2026/10/5
READ
长读
The Physical Cost of AI: Scanning and Destroying Books
illustration · QianLong editorial

When we imagine how artificial intelligence learns, the process usually seems entirely virtual. We picture algorithms silently scraping billions of web pages, absorbing text from forums, blogs, and digital archives stored in invisible data centers. But the quest for high-quality training data has a surprisingly physical—and destructive—footprint.

A recent investigation by tech outlet 404 Media brought this physical reality into sharp focus. According to a worker at an Amazon warehouse, the company has been running an operation where physical books are systematically scanned to train Amazon’s artificial intelligence products. Once the text is successfully digitized and fed into the data pipeline, the physical books are not repacked, resold, or donated. Instead, they are destroyed.

This revelation highlights a jarring intersection between traditional media and cutting-edge technology. For centuries, physical books have been revered as enduring vessels of human knowledge and culture. Yet, in the race to build smarter and more capable large language models, these objects are being reduced to mere raw material. Once their informational value is extracted by the scanner, their physical form is treated as industrial waste.

The practice sheds light on a growing challenge in the AI industry: the looming "data wall." As tech companies exhaust the supply of freely available, high-quality text on the open internet, they are increasingly desperate for professionally edited, long-form content. Books offer the perfect linguistic complexity to teach AI how to reason, structure arguments, and mimic human storytelling. By processing and discarding physical copies, companies can rapidly ingest vast libraries of material, transforming literature into algorithmic fuel.

Beyond the obvious environmental concerns of pulping perfectly good books, this operation raises significant ethical questions. It forces us to reconsider the hidden costs of the AI tools we use every day. The intelligence of these systems is not conjured out of thin air; it is built on immense physical resources, from the staggering amounts of electricity and water required to cool servers, to the manual labor of warehouse workers dismantling physical media.

As we continue to integrate AI into our daily lives, it is crucial to recognize the tangible sacrifices made to create it. The smooth, conversational output on our screens often obscures a messy, resource-intensive supply chain. Understanding this reality helps us look past the magic of the technology and ask harder questions about sustainability, the value of physical media, and what we are willing to consume to make our machines a little bit smarter.

Key Points

  • An Amazon warehouse worker revealed that physical books are being scanned to train the company's AI models.
  • After the data is extracted, the physical books are systematically destroyed.
  • The practice highlights the tech industry's desperate need for high-quality, long-form training data.
  • This reveals the hidden physical and environmental footprint behind seemingly invisible digital AI tools.

Why It Matters

It shatters the illusion that AI is an entirely virtual technology, exposing the real-world material waste and ethical compromises involved in training large language models.


Sources:

潛
本文完
潜龙编辑部 · 2026/10/5
潜龙 QianLong · 中文 AI 内容与工具平台