• isleepinahammock@lemmy.blahaj.zone
    link
    fedilink
    English
    arrow-up
    3
    ·
    5 hours ago

    Are you sure they’re digitizing Book By the Yard quality material?

    Consider this. Purchasing, shipping, and physically scanning a book is the single most expensive way for an AI company to acquire the text of a work. It’s been widely documented in court cases that these companies engage in mass IP theft. They’re literally running massive torrenting farms, grabbing copies of every film, song, book, etc. that they can get their hands on.

    What kinds of texts are most likely to be found pirated on the internet? It’s the mass market stuff. The common stuff was digitized long ago. They can just download that. They can pirate it. They can buy the ebook. There’s no need for OpenAI to purchase and destructively scan the works of Steven King. No shade on the man, but his works aren’t exactly hard to find. I’m sure they can just find a torrent.

    There’s little value in scanning the mass market books that are produced in enormous quantities. They probably don’t buy “Book by the Yard” books, because they already have digital copies of those.

    But the rare, long out-of-print stuff? The stuff that you actually cannot find a legal or illegal copy of online? That’s only stuff worth paying money to buy, ship, and scan.

    Doing anything physical, especially at scale, is slow and expensive. I would guess a good portion of these books, perhaps an outright majority, have simply never been digitized.