· 1 views

Amazon joins the list of companies destroying rare books to feed to the AI machine

Amazon joins the list of companies destroying rare books to feed to the AI machine

An AirTag hidden in a rare book followed its journey to a Las Vegas warehouse, where it was likely destroyed in an effort to supply fresh data to hungry LLMs.

I don't know if you've noticed, but everything is becoming increasingly AI-generated. Whether it's news, music, or social media, AI has seemingly stuck all of its nasty little fingers into any pie it can find.

Experts have long warned that eventually we'd get to the point where AI was training itself on data that AI had created, and we're pretty sure that point is behind us. And that's not actually particularly useful, it turns out.

It's harder to pass the Turing test when everything sounds exactly like it fell out of the back of a ChatGPT prompt. And that's why AI companies don't really want the AI version of the human centipede.

They want to train AI off of human-produced content. Which, as a human who produces human-produced content, you'd assume I'd be relieved to hear that I at least have a spot in the pipeline.

Unfortunately, those same companies are telling other companies that they can get rid of the human in favor of their hyper-efficient, human-like product.

So, in the human centipede analogy, where does that leave me, specifically? Am I at the front, or did they try to find a space for me somewhere in the middle?

Just kidding, they really don't want humans involved in the process at all.

That's why they've gone straight to destroying books.

In a warehouse in northeast Las Vegas...

Rare booksellers have seen historical spikes in sales in the past year. So 404 Media reached out to one to see if they could track who was buying these books.

The seller agreed to hide an AirTag in a rare book that was part of a larger 1,000-book order. The book traveled from California to Milwaukee, then from Milwaukee to Grand Junction, Colorado.

Stack of four round white tracking disks on a smooth light surface, softly lit with purple and blue background tones, emphasizing their clean, minimal tech design

An AirTag was used to track the book to its final destination

After a few weeks, it arrived at its final destination: Las Vegas, Nevada.

Specifically, it wound up at an Amazon warehouse known as "LAS8." Typically, LAS8 operates as a print-on-demand book-selling operation.

But 404 learned that the north end of the warehouse doesn't create books. Instead, it destroys them.

The location in question is known as VGT3. It is there that books are received, sorted, and then processed. Processing requires the spines to be cut off and the pages to be fed through scanners that convert the scans into data.

The data is not available anywhere, they are not being scanned into PDFs to read. It is used exclusively to train AI.

The remnants of the books are discarded. The contents are never meant to be viewed by humans, only by AI.

If this were a movie, I would have called the writer out for being a hack. But unfortunately, this is just the reality we're living in now.

Welcome to the future, dear reader. Is it as exciting as you'd hoped it would be?

21st Century book burning

Data is finite, unfortunately. Especially as far as large-language models are concerned.

There are lots of reasons for this, but two major ones that seemingly have AI companies and researchers on edge.

The first is model collapse. When AI trains on its own data repeatedly, eventually the data condenses down so much that it becomes useless.

When this happens, AI will produce generic, unusable information. Effectively, it's the same thing that happens when you photocopy a photocopy of a photocopy.

The second is that AIs hallucinate. It hallucinates a lot, actually. So if one model trains off of another model's hallucination, the hallucination is treated as the baseline.

Now imagine that happening over and over and over again. Eventually you arrive at the same point as model collapse, but this time with the added danger of AI citing hallucinations as fact before they eventually become unusable.

All roads lead to the same location. And it's not a location anyone is really interested in going to, or alternatively, staying there, like my editor thinks we are now.

Nearly all of the human-produced training data available on the internet has already been slurped up by LLMs. And, as mentioned before, new data is overwhelmingly being produced by AI.

So AI companies need to figure out how to circumvent the digital ouroboros effect. But if humans can't create fresh data fast enough, where do you find new information to feed into the machine?

Well, as we learned above, it's apparently by destroying books. Books are a goldmine where AI training data is concerned.

Physical books are even better. Much of the information contained within hasn't been digitized, which means it's raw, virgin data.

Open book on wooden table with reading glasses resting on its pages, a smartphone lying beside it, and another closed book stacked in the background

Image credit: DariuszSankowski on Pixabay

And the data doesn't really need to be "factual"; it just needs to be unique, unsullied by AI. It just needs to train the math behind the algorithm.

This is all is why rare booksellers are seeing an increase in sales. And judges have ruled that it's fine for a company to do so, provided they're not selling the digitized copy afterward.

Amazon isn't the first company to do this. Anthropic has also done this. In fact, I'd be willing to bet that most big commercial AI outfits have been doing this for a while now.

Eventually, the companies will run out of books. And likely, some exceedingly rare books will be entirely lost in the process.

At some point, the stone will cease to yield blood, but I severely doubt that the AI powers-that-be will ever stop squeezing.