Almost all these books were either headed to the dump or the recycling center.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
Are you sure they’re digitizing Book By the Yard quality material?
Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.
What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.
There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.
But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.
Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.
My wife works at a law firm that contracts with “Books By The Yard”, which provides books purely for office aesthetics. For a few hundred bucks you can plaster a bookshelf full of material nobody will ever read, because they’re such a commodity.
It’s so crazy to see people ingesting this news from the conspiracy-brain perspective of “The AI companies are stealing all the knowledge!” without recognizing the more pressing reality of decade upon decade of publisher overproduction, resulting in a total devaluing of physical media for its academic importance.
Imagine going into hysterics because a warehouse full of shitty airport books went up in smoke, like it’s the Library of Alexandra that just burned down.
Are you sure they’re digitizing Book By the Yard quality material?
Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.
What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.
There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.
But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.
Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.