To be clear, the issue is not that the books were used to train Claude, but that they were pirated.

It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)

I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).

Now we're in a world where you have to have dozens of millions in capital to do substantial work.

I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...

> Fun fact: a pallet of books is called a "gaylord,"

A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding

Ladder-pulling at its best. It's also easier to swallow the fine once you've launched a successful product after pirating the books.

Judging by the current comment section, most HN people seem to want the ladder pulled

(an observation, not agreement)

As this comment states, it may not be the piracy that is the issue, but the keeping of the books forever rather than just for the purpose of training. Regardless, I believe people will do this in a wink wink nudge nudge sort of way anyway, as no one releases their training data, because we all know where it comes from.

https://news.ycombinator.com/item?id=48996652#49004015

This is why you should always distill your models from a competitor.

Let them take on the liability

With this it's becoming very clear that we're moving past information copyright of today and the only copyright that'll remain will be brand/trademark shaped. This might be a good thing right? Information remains free while people's effort remains protected (assuming fair governance).

That would be trademark not copyright

Which is another issue. In the context of pirating, it should be also an issue, because it is a benefit from the crime.

A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s

They actually did this.

> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books

> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.

This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.

Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)

If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market

It reads like you're in favor of banning resale and lending(aka libraries) of books... because the authors aren't compensated.

There's such a thing as fair use and digitizing privately owned printed material is absolutely legal... including for corporations.

In many territories like Ireland, Authors are compensated for their inclusion in lending libraries. It tends to be of the pitiful 'music rights organisation' style mechanical reproduction royalties, but it does exist.

No, it’s proof purchase of how stupid the publishing industry is. Maybe publishing houses should just pay authors good money, like a goddamn salary, and get a book out of them every few years.

They can't for the same reason that cab companies can't make their drivers employees: they would have to employ far, far fewer of them than they do on contingency.

The AI craze not only destroyed books, but many small websites who couldn't bear the load of constant scraping, or many communities that took open forums and took them offline or put them behind closed doors.

There is less publicly available knowledge now on the Internet than there has been 3 years ago.

Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!

>instead of allowing anyone to train on already scanned books for free

That would be pirating. So your complaint is that they didn't do more piracy?

My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.

To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.

no we should destroy the works of these ghouls and support humans instead of this destructive and useless technology

the people operating frontier labs are bad people they cannot be trusted in any way

the best solution to them would be to send them to monster island (even though it's really a peninsula)

well if they made a PDF copy to process they violated copyright

How is this ruled as piracy then? I am confused.

They also pirated the books

They bought, scannned, trained from, and destroyed millions of paper books, which was ruled legal. This lawsuit was for training from LibGen.

This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.

But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.

>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.

That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.

They could have purchased the books instead. It was easier to pirate.

how much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work?

at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.

our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.

they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.

All of it.

If you believe information deserves to be free, and if most of your earnings were from information that wasn't given away for free -- well, if you want people to give up their ill gotten gains, maybe you can start by setting an example.

So, mind sending me your bank account information? I'll promise to make good use of it.

You seem to think that AI companies sell you content, which is false.

You get a service. The service is using their compute power to run a model and their scientists to build the model.

No. I think AI companies consume the result of intellectual work that should typically be paid for. If someone doesn't believe intellectual work should be paid for so that others can benefit, I invite them to lead the way.