It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
> Fun fact: a pallet of books is called a "gaylord,"
A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding
Ladder-pulling at its best. It's also easier to swallow the fine once you've launched a successful product after pirating the books.
Judging by the current comment section, most HN people seem to want the ladder pulled
(an observation, not agreement)
As this comment states, it may not be the piracy that is the issue, but the keeping of the books forever rather than just for the purpose of training. Regardless, I believe people will do this in a wink wink nudge nudge sort of way anyway, as no one releases their training data, because we all know where it comes from.
https://news.ycombinator.com/item?id=48996652#49004015