It’s been established that LLM-generated code is not copyrighted so I can fully understand the company living from copyrighted data to not accept LLM-generated contributions.

> It’s been established that LLM-generated code is not copyrighted

If that's a reference to Thaler v. Perlmutter, the only thing that's been established is that an LLM can't be considered an author under the Copyright Act, only a human being can. It says nothing about the consequences of a human claiming authorship of LLM-generated code, which would be relevant here.

You lay out my words in different order and claim I’m not correct.

No, your claim is "LLM generated code is not copyrighted." His claim is "LLM generated code is eligible for copyright."

My understanding (belief) is that it's going to depend on how much human involvement is there.

If you write a prompt and one-shot a problem and share the source code, that source code is probably not covered by copyright.

If you substantially edit or modify the generated code you would own the copyright.

It's like with a camera. If I set a camera and carefully aim it and somehow trigger the shutter then make adjustments in Photoshop, I own the copyright on that image.

If I stick a Flock camera on a pole somewhere and post the live output, there's been no meaningful human creative involvement in producing those images and so nobody can claim copyright on them.

I think if I as a human use an llm to do something technical that would qualify copyright, it should still qualify for copyright. How do you decide how much human is copyrightable. If I use a package that writes code or use a library for some piece of it, I could still copyright.

I don't like this idea that llm code can't be owned by a human, copyrighted. It's just code.

I think your last example with flock camera is relevant here - I can take a picture of a public football as a reporter or something (or a fan I guess) and I can copyright and sell that picture. Newspapers do it every day.

So if I stand on a street corner and take a pic, it's copyrightable. If I take a pic using a flock camera it should also be copyrightable, just like if my nest camera at home takes a pic of something, I can use that.

I guess you are saying "someone else owns the flock camera" so you don't get to own pictures. What if I buy the flock-like camera and put it up, I should own that.

You seem to have misunderstood an is/ought distinction. You may hold the (fairly extreme, as far as copyright goes) position that surveillance footage should be subject to copyright, but it's well established that it's not. Who owns the camera is irrelevant. At least in the US; I'm not aware of any jurisdictions that hold otherwise. This is why Wikipedia articles on world events in the past few decades are full of stills from surveillance cameras: it's one of the few sources of imagery of an event that are unambiguously legal to include, because unlike a photo or intentionally made video of something specific, it's not a creative work. It's also pretty firmly established that human authorship is required for something to be subject to copyright, and having an idea that lead to some particular expression is itself not sufficient; see, e.g.: https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...

(Not a lawyer, just a Wikipedia editor.)

I appreciate your informed take. I follow the reasoning but I am amazed it works this way. I found some articles that supported what you said, and also said there's a follow-on industry that figured out how to alter and edit videos just enough for a revised video to have creative contribution and make it copyrightable.

https://www.techdirt.com/2020/02/24/can-you-license-video-yo...

> https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...

This is a great case to study, but no determination of copyright was made. The only actual lawsuit was filed by PETA arguing that the monkey should have copyright, which led to a settlement with the human photographer and nothing else because obviously that's not possible.

For various reasons (mostly $$$) the guy never actually got a judgement. I think the chances are good that he could have prevailed in court; there is significant creative input to setting up cameras and triggers in a way to convince a wild monkey to take a selfie. It's not like he just left his camera sitting somewhere on accident and came back to find a photo in it.

I'm also not a lawyer, but I did do a lot of work in copyright for a company you've heard about.

> If you write a prompt and one-shot a problem and share the source code, that source code is probably not covered by copyright.

We will have to see about that! This is the kind of boundary that's still being figured out in court; it's going to depend on how hard you worked on the prompt. I highly doubt that even most slop was generated with a single half-ass prompt, and the bar is not as high as you might expect.

> If I stick a Flock camera on a pole somewhere and post the live output, there's been no meaningful human creative involvement in producing those images and so nobody can claim copyright on them.

It really depends on what pole, where, and why. In a parking lot in rural Wisconsin? Probably not. A recorded livestream of a political march? You likely have copyright.

I think that by virtue of the sheer amount of time spent using AI tools, it's pretty clear that these outputs have enough creative input to be copyrightable.

No, that’s not what they are saying. They’re saying that the code generated by a human with help from an LLM may potentially be. This is what I hope we are going to arrive at, eventually.

How could you establish what parts of the code was produced by an LLM vs updated by a human afterwards?

The LLM will output different results over time as the models get updated. Are we heading towards needing to retain a full prompt history that can be replayed against a specific LLM model version to prove what the output was for copyright purposes?

For those wondering what the difference is: consider what happens when an LLM regurgitates its training data. It's copyrighted... but not by the person who generated it.

[deleted]

Congrats on having the worst take in a thread full of them.

This is a false narrative based on a (IMO often intentional) misunderstanding. It has by no means been established by any court that LLM-generated code is not copyrightable.

Thaler v. Perlmutter stands for a much narrower proposition and at any rate is not binding nationally, SCOTUS having denied certiorari.

This is not necessarily true. While no court has explicitly come out and said that copyright does not apply to AI-generated works of authorship, the US copyright office has[0]:

> Based on an analysis of copyright law and policy, informed by the many thoughtful comments in response to our NOI, the Office makes the following conclusions and recommendations: > • Questions of copyrightability and AI can be resolved pursuant to existing law, without the need for legislative change. > • The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output. > • Copyright protects the original expression in a work created by a human author, even if the work also includes AI-generated material. > • Copyright does not extend to purely AI-generated material, or material where there is insufficient human control over the expressive elements. > • Whether human contributions to AI-generated outputs are sufficient to constitute authorship must be analyzed on a case-by-case basis. > • Based on the functioning of current generally available technology, prompts do not alone provide sufficient control. > • Human authors are entitled to copyright in their works of authorship that are perceptible in AI-generated outputs, as well as the creative selection, coordination, or arrangement of material in the outputs, or creative modifications of the outputs. > • The case has not been made for additional copyright or sui generis protection for AI-generated content. > The Office will continue to monitor technological and legal developments to determine whether any of these conclusions should be revisited. It will also provide ongoing assistance to the public, including through additional registration guidance and an update to the Compendium of U.S. Copyright Office Practices.

Congress or the courts could, of course, override the stance of the copyright office, but I think it would be highly unusual for them to do so (particularly for something like this). It would however be a lot better if congress just stepped in and said no outright, but until then this will have to do.

[0]: https://www.copyright.gov/ai

The copyright office does not determine the standards for copyright. They are an advisory and notary organization.

Only Congress and the courts do. Copyright exists from the moment a work is created, and does not need to be registered with the copyright office.

The law isn't that complicated; if a work was created with a human being with intent, it's probably eligible for copyright protections.

As long as you can convince a court that you did this, the tools you used are not relevant. The vast majority of LLM art falls in this bucket.

Do you own the copyright to a painting you commissioned or does the artist? You may have described what you wanted, but the creative work is the output of the artist. Same with an LLM.

The courts are 100% going to have to interpret what is "sufficient human control" at some point.

Sure, but I would be incredibly shocked if the courts overturned these conclusions. These kinds of determinations are within the remit of the USCO, so a court does not need to come out and say it if the USCO has already done so. Obviously, as I said it would be better if congress weighed in and solved this problem, given that the USCO is free to publish a new NOI to change it's practices/policies, but we all know that congress is too gridlocked atm for that to happen

I am trying really hard not to accuse you of not having read what you posted, because your conclusions are in strong tension with what it plainly says.

But there are no conclusions. It literally says:

> Whether human contributions to AI-generated outputs are sufficient to constitute authorship must be analyzed on a case-by-case basis

It says a plain prompt is not enough but that is not the reality of real software development. People aren't one-shotting complex business apps. The vast majority of software development will trivially pass that bar and end up in the "requires case by case analysis".

You are a false narrative. I’m just repeating what I read. The thing is, it’s murky waters. Someone’s going to challenge it but do you want to be the guy who takes it on a chin?

The sibling comment lays this out and my original comment above is based on exactly the same link.

Also it's a different story intranationally for those of us who live in countries with much more restrictive/no fair use. Are you geolocking your software to the USA?

I highly doubt that that's the last word on that matter, but even if: Even before LLMs you could combine individual non-copyrighted components into something copyrighted.