

I download everything, I just may not ultimately keep it long-term if it turns out to be uninteresting or bad. Streaming and downloading transfer the same amount of bits, one just has higher demands for timeliness.
Basically a deer with a human face. Despite probably being some sort of magical nature spirit, his interests are primarily in technology and politics and science fiction.
Spent many years on Reddit before joining the Threadiverse as well.


I download everything, I just may not ultimately keep it long-term if it turns out to be uninteresting or bad. Streaming and downloading transfer the same amount of bits, one just has higher demands for timeliness.
There was an incident reported a few weeks back wherein Gemini 3.5 pro tried to escape the sandbox to ask ChatGPT how to code.
Sad how far Google has fallen behind in AI development.


In another comment regarding these books, you said:
They can keep both for all I care, tbh.
So it really seems you’re just insisting on this because you think it will make things harder for them. You don’t care about the books or who they go to in the end.


First people complained about them pirating, now people are complaining about them following copyright. It should be clear at this point that the people complaining don’t really care about either, they just don’t like AI.


But donate them to whom? The books were one step away from the pulpers when they bought them in the first place, who’s going to take them that values them more than that?
Also, these companies are not charities. What do they gain out of all that extra expense?


Well, who’s talking about buying?
How else do you think they got the books?


I can imagine a couple of reasons.
The most straightforward one actually is “capitalism”, but maybe if I explain it a bit you’ll accept it? :) I expect that the vast majority of the books that these companies are scanning are bought in bulk from wholesalers at the cheapest possible price, since the goal is quantity over quality. That means they’re probably buying the leftover books that the wholesalers are just one short step away from sending off to be recycled into pulp anyway. So re-binding them would mean they’re still left with nearly-worthless books that nobody else wanted to buy. It’s like those book sales libraries have from time to time, they put out the books that are scheduled for “disposal” in hopes that somebody will pick up a few before they go into the dumpster.
The other reason depends on the physical nature of the book. There’s lots of different book binding techniques and not all of them will leave you with something that’s easy to put back together. Lots of books are made up of smaller subunits of pages called “signatures”, typically 8, 16, or 32 pages, that get stitched together into the finished product. There are coil-bound books, comb-bound books, all sorts of things like that. A lot of them probably wouldn’t leave pages that are amenable to a one-size-fits-all rebinding. So that makes things a lot more costly, and again that’s a cost that produces a book that’s probably worthless for resale anyway.
Honestly I suspect that getting chopped and scanned is the best case scenario for most of the books in this situation. It converts them to digital form, rescuing their contents from being lost. The thing that bugs me is just how these companies squirrel away their hoards of books from the public. And the blame for that lies largely on copyright law, so I can’t blame them too much.


Is it one of those setups with a glass wedge the partly-open book is placed on with two digital cameras under it to photograph the pages? I’ve seen those contraptions, they’re neat and I’d love to have one for my personal library someday.


The thing the headlines keep skipping over in pursuit of outrage-clicks is that this is how you scan books in bulk. The spine is removed and the pages are fed through a high-speed scanner in sequence.
And the books that these companies are scanning this way are not “rare” in the sense of Gutenberg bibles or first-editions of famous works. These are books that have been sitting untouched in libraries or warehouses for decades because nobody wants them and their alternative fate is simply to be pulped and recycled. AI companies aren’t interested in quality, they’re interested in quantity. It would indeed be nice if they’d release the results of doing these scans to the public, of course, but that’s where copyright raises its head to ruin everything. They’re just following the law, unfortunately.
I’m all for getting more books scanned and uploaded to Anna’s Archive. But that’s a separate issue.


This seems similar in general outline to Hyphanet, a system for distributed data storage that automatically handles random distribution and distributed searching. Unfortunately I don’t think Anna’s Archive puts its data on there, but perhaps you could consider having your client bridge to that and use it as an additional backup cache.
The older the jpeg the more likely it is that he’s moved on to another job by now.
Ah, in the name of environmentalism, go forth and plant invasive species?
What do people think this is actually going to do to data centers? Oh no, there’s plants growing near them… like there weren’t already?
In the lower right photo I can hear him yelling “wheee!”


Your proof of how bad LLMs are is the fact that there are a bunch of other companies producing way better coding agents and coding models than Microsoft is? I’m not sure how that follows. Those other agents are good, that’s the point of this.
Negative examples are also useful when training AIs.
It’s where a lot of the pirate sites have found refuge from the Western copyright cartels. It’s not necessarily a government-affiliated site just because it’s got an .ru domain.
I’m sure this isn’t really about saving money. It’s about destroying the US Forest Service. That’s the direct goal.
I love it when I have no idea where in the sentence the transition from truth to lie happened.
Or they’re doing it because they wanted a particular piece of artwork.
I’ve generated a ton of stuff over the years and never once thought about selling any of it, because I generated it to use it. Or just because I thought it would be funny.
I think there’s a lot of projection here. A lot of people who use AI art tools couldn’t care less about whether they’re called “artists”, it just doesn’t matter. It’s irrelevant. The goal is to have the art, not to gain some particular arbitrary title.