So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything.

I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.

In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.

I suspect human writing could survive, but I don’t see any room for publishers.

> So will you owe life long compensation for all the knowledge you got from books too?

No because we are people and the laws differ for people, corporations, and machines.

I think people forget that laws are perfectly capable of carving out exceptions, leaving purposeful ambiguity, expressing intent, etc. Yes, humans can have special rules, and very obviously should since laws exist to improve human lives.

I am actually not settled on either side of the matter and I have not forgotten that, but I think what we are really looking at is a rather more complex matter than people want to make it out to be. We are holding several but at the very least contradictory positions and they are incompatible.

Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?

Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.

The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.

> So will you owe life long compensation for all the knowledge you got from books too?

You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copyright law either, because there's no copy.

Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text

The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.

The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...

https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...

It's possible. Would be interesting to see their evidence, and to know whether they can reproduce it for arbitrary articles, not just ones that have been endlessly republished on the net.

> spitting out verbatim text

> had produced near-verbatim replicas

Don't get too hung up on the preciseness of the copy - the courts won't. I doubt that spitting out an existing article with a few adjectives changed would be considered transformative.

> if you can't persuade the machine to spit substantially the same text back out verbatim

That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.

Can you cite any information on this not being possible for the vast majority?

Or is it simply that the correct prompt hasn't been written for all possible cases?

I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.

It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.

Why is a vector database somehow completely different from maintaining a library of the text itself?

Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.

Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright.

LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.

> but that derived image would surely be under copyright.

I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans

I don't know that this really challenges anything relating to compression.

Why would a derived image be under copyright?

Because that's legally the case? I don't understand the question. Using a lossy compression algorithm on an image does not remove its copyright protection.

Nah, I can't prove a negative. But Common Crawl is 12 petabytes and is not the largest part of what these models get trained on. DeepSeek v4 Pro is, what, 865GB?

That's one hell of a compression ratio, if it can do what you claim.

You are asking to prove a negative. But even assuming that the model is capable of returning every bit of its training data verbatim (a mathematical impossibility) that would not be enough as mere capability is insufficient here. If capability alone were the standard any library that also has a photocopier / scanner would be in violation.

To prove distribution of copyrighted materials it would have to be practical and actually used in the wild by people to circumvent copyright and generate copies of those works. Again, I can't prove a negative, but that isn't the standard, and nobody has shown a practical exploit here.

There was a paper a while back where (from memory) they managed to coax 75% of the original text of some internet-popular books out of an LLM. Harry Potter, 1984, etc. That's why I said it was possible for some texts.

My assumption is that multiple copies in the training data "wear a deeper groove". I believe those are infringing, and should be dealt with on a case-by-case basis. But the vast majority of text doesn't wear that groove.

(Edit: Think it was this one https://arxiv.org/abs/2601.02671)

This reliably inevitable rationalization comes up in every thread it seems, and its ultimate goal is to humanize AI. This is what the big guys want us peons to believe and it works so well, I have even been lectured by an AI for being rude, the implication was that I was logged in and it would be a shame if anything happened to my account.

Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.