Amazon is drawing attention for an unusual approach to AI development: sourcing rare printed books and converting them into digital text for model training. According to reporting from 404 Media, a tracked rare book was ultimately found at Amazon's VGT3 facility in Las Vegas, a site associated with AI data work.
Amazon said it buys books through commercial channels to help improve the products and services used by customers. The broader context is clear: large language models require enormous volumes of text, and rare books can offer material that is hard to find online or in modern digital archives.
That makes older publications especially valuable for training systems that need diverse, high-quality language data. Texts published before 2022 also carry another advantage: they were created before the rise of mainstream generative AI, which helps reduce the risk of models learning from machine-written content.
As AI systems expand, the search for trustworthy, original training data is becoming a defining challenge. This shift suggests that the future of AI may depend not only on computing power, but also on how carefully knowledge is preserved, digitized, and used.