A flagship LLM today is trained on tens of trillions of tokens, the equivalent of hundreds of millions of 100k word books. No human has that kind of appetite.
As an aside, the entire Google Books corpus is generally estimated at tens of millions of books.
A flagship LLM today is trained on tens of trillions of tokens, the equivalent of hundreds of millions of 100k word books. No human has that kind of appetite.
As an aside, the entire Google Books corpus is generally estimated at tens of millions of books.
Any original text has value for the purpose of AI training. The argument is that they have no other value.