That’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.

the "codec" is not really the point.

playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference.

similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft

the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for).

similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books.

the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts.

to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more.

having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.

> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though.

The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).

It's a perfectly sensible interpretation of international copyright law.

I'm not sure if you're serious with the suggestion I could sue them. These are both US corporations, that justice system is pretty much in shambles in particular when it concerns corporations as big as these AI ones. You can dig your heels in the sand to defend that system, but you will also have to dig your head in the sand about why Sam Altman doesn't have a Disney "influenced" avatar, but one "inspired by" Studio Gibli.

And I'm not sure if you're familiar with the concept of "fair use" in the US as it "works" in practice, it's almost insulting, ask any music education youtuber.

Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.

> Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.

I don’t know why you’re repeating the stuff I just wrote like I didn’t. My point is that this is only relevant for the fair use defense and not copyright in general.

Here’s what I said:

> The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).