> There needs to be a royalty payment based on if the AI regurgitates existing ideas.

So, by that logic, you need to be paying every time you regurgitate any of my ideas. Or anyone else's. Copyright now protects abstractions and vibes. Substantial similarity test be damned. Nobody can write stories about wizard schools, the idea is taken.

Humans are not computers. Humans are not a service. In the end, all laws are made up rules and can absolutely be written to have different outcomes and restrictions based on if a human is doing something or if a program is doing it.

Real question, if an LLM shouldn't be able to remix someone's written work, why should a robot be able to build a chair that kinda looks like a chair a carpenter built that one time? The carpenter was a human, and humans are not a service.

Why this distinction only for intellectual work?

For the same reason that you can make a similar-looking chair, but you can’t distribute a fuzzy copy of Star Wars. The char isn’t a copyrighted work.

"The char isn’t a copyrighted work."

An Eames chair is, we just have a really high bar for what is copyrightable in the physical world, and it seems pointlessly discriminatory.

Uhuh so it seems that you weren't, in fact, asking a "Real question", but came here for an argument.

this copy of Star wars seems pretty fuzzy https://dev.to/kasuken/how-to-watch-star-wars-in-your-termin...

there are also fan remakes of movies like this one. https://www.imdb.com/title/tt3528906/

I'm not a lawyer but this does seem like they're wholesale copying ideas.

Furniture designs can be covered by varying intellectual property laws.

Honestly? I don’t know, I don’t have a whole coherent ethos about LLMs.

But I do know someone definitely paid for the textbooks I used when learning in school.

I suspect chairs have been public domain since the advent of man.

You indeed need to pay someone if you take their copyrighted materials and regurgitate it. Ask DJ's and producers how they need to include royalties for samples used in their tracks.

There’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.

There's also a difference between an MP3 and a FLAC. Again, ask DJs how well they're getting away on that distinction.

That’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.

the "codec" is not really the point.

playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference.

similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft

the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for).

similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books.

the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts.

to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more.

having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.

> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though.

The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).

It's a perfectly sensible interpretation of international copyright law.

I'm not sure if you're serious with the suggestion I could sue them. These are both US corporations, that justice system is pretty much in shambles in particular when it concerns corporations as big as these AI ones. You can dig your heels in the sand to defend that system, but you will also have to dig your head in the sand about why Sam Altman doesn't have a Disney "influenced" avatar, but one "inspired by" Studio Gibli.

And I'm not sure if you're familiar with the concept of "fair use" in the US as it "works" in practice, it's almost insulting, ask any music education youtuber.

Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.

> Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.

I don’t know why you’re repeating the stuff I just wrote like I didn’t. My point is that this is only relevant for the fair use defense and not copyright in general.

Here’s what I said:

> The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).

So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.

"I get a GPL version of Microsoft Office."

Is this not..Libre?

LibreOffice is a different product from Microsoft Office.

but it's the same idea

If you just use the abstract idea, you could have done the same thing yourself all the time already.

Yes exactly. That's basically what an emulator for a games console is for example, a reimplementation of the original.

So a console game loses its copyright if you emulate it?

A game is a specific work to be copied so no, but the system it runs on can still be without copyright.

The game doesn’t but the emulator doesn’t necessarily infringe copyright in itself just because it is based on the original console.

Is Claude's paraphrasing of Lord of the Rings equivalent to the original?

That’s what a brain is

Why do companies bother with the Chinese wall technique, then?

Maybe but your brain is not running 24/7 capable of outputting thousand if not millions of tokens per hour, all while having ingested nearly the entire internet.

If yours do that, maybe we can redefine what copyrighting and patenting means for humans

This isn't that out there in our current scenario. These models compress our collective thought and effort. Why not make these publicly owned, all profits distributed back to us?

Every morning we pay royalties to prometheus for we all are toast.

Well I think that is the greater question.

Is AI just an algorithm. Is human creativity just an algorithm?

Who, if anyone should own the copyright if you prompt AI to write a book?

I'm thinking more from a moral and philosophical pov, the copyright regime is broken anyway