- cross-posted to:
- technology@lemmy.world
- technology@beehaw.org
- cross-posted to:
- technology@lemmy.world
- technology@beehaw.org
The real problem here is copyright law. We need a system where registration is required to extend the copyright term beyond an initial relatively short term. That way we can identify orphaned works and make them available to the public.
It’s not simply that they don’t care about the value of the books.
The tech oligarchs are trying to convert information into property, with themselves as sole owners
These dumb mother effers are probably training the next model on the Voynich Manuscript.

Rare Books
I’m still waiting to see what an example of a “Rare Book” is supposed to be. Are these first edition copies of To Kill A Mockingbird and Ulysses? Or are we just talking about books by new authors that were never widely distributed or reprinted, because their sales numbers were no good.
This post is for paid members only
I guess I’ll never know.
Rare meaning uncommon, hard to find, not many copies are available.
But that doesn’t necessarily mean valuable. These have been sitting on shelves in warehouses unsold for a long time. No one else wanted them.
But let me be clear, I’m against their destruction for private use. The digitized copies should be made freely available, if it all possible.
There are so many niche instructional books that are now out of print. If these books were scanned (yes, destroyed in the process) and then added to a free and open book archive for all to access, I wouldn’t have a problem with it. I have a problem with these AI training centers because all that knowledge is just going into a back hole.
There are music theory books and penmanship books I’m actively hunting for through piracy and ebay because they are out of print and rapidly getting lost to time. If I find them, I plan to scan and share them, not for piracy but because the original authors works do not deserve to disappear just because they no longer generated profit.
That’s not even touching in the price. Sometimes they sit there untouched simply because whoever happens to have one of the few copies put a ridiculous price on it.
I’m not saying it needs to be free (I mean I do but that’s a separate argument). I’m saying many of these sellers want as much money as physically possibly with no real world basis for the prices, which is why they sit there.
There’s several old Welsh poetry books I’ve gone looking for only to find copies marked for a thousand bucks pricing out any regular person and leaving only companies with stupid money to burn or maybe one day a collector.
One of the books in my hunt list is the one below. It originally sold for $16 but is now out of print. Notice the reseller’s price.

Yup, and it’s just gonna sit there and go up in price as no one buys it because it’s “rare” and out of print. I feel ya
The big tech companies will drive up the price of second-hand books, just as they drove up the price of computer hardware and electricity.
At this point, I’m just hoping that when we are sifting through the smoking wreckage of the AI crash we’ll find a hundred million books that we can add to Anna’s Archive.
And some RAM. RAM would be nice.
They aren’t never seen before unique rare books, they’re just out of print and not lots of copies. They’re destroying them to sidestep copyright concerns.
They are ripping them from their spines and pulping them once the have what they want.
They are ripping them from their spines and pulping them once the have what they want.
They treat their books like they treat their employees.
Also like they treat nature.
They mean the PDFs of them or ePub or whatever they’re scanning them into. Yeah, the original is gone but imagine if the Library of Alexandria went up in flames but every ash contained a complete work for free to everyone to read at their leisure.
Oh, and Ram.
Storage space is valuable for their AI stuff. Why would they keep around copies like that on valuable storage space if they don’t think they’re going to be using them again once they’ve trained their model? I really hope they are scanning them to a format or something like that that. But I doubt it. These aren’t the most forward looking or intelligent people you’ll find. I wouldn’t be at all surprised to find out that they never did anything more than scan it into train the model and then flush it from the system. It’s 100% on brand for them.
You always keep training data because you never know what you need for the next generation. Ironically they might be preserved (privately though unfortunately) pretty well






