Doesn't clean room typically apply to cases where the "dirty" team has legitimate access to copyrighted code, like the IBM PC BIOS which was published in the technical reference manual, and uses this access to write a functional specification for the "clean" team?

I'm not sure how clean room would apply to commercial applications distributed in binary form, as there's no way to look at even disassembled code without violating the license agreement and therefore being in breach of contract and subject to potential copyright infringement claims for copying or even continuing to use the software, let alone cloning it, and surely you're not going to be subject to a copyright claim based on familiarity with the application from merely using it.

The issue is that LLMs very likely have access to the source code of these products as part of their training data, and are then being used to generate clones. There's no barrier in the middle to ensure copyright violations don't leak

The reason why 'reverse engineering' has gotten so good is because what we're actually seeing is fully automated luxury plagiarism

It is very unlikely that LLM had access to original source code of Photoshop.

> It is very unlikely that LLM had access to original source code of Photoshop.

It doesn’t matter. It’s likely had access to the knowledge of somebody who worked on that source code.

I have worked at companies that have done clean-room implementations. I was not permitted to look at their source or interact with those teams, because I had been exposed to the source of what they were re-implementing in a prior job.

As a human with a mushy brain I was never going to remember the source, but the fact I had been proximate to it was enough to lock me out. You have no idea what the LLM trained on, so you can’t prove a LLM derived clone is a “clean-room” implementation.

Some questionable idea that the LLM "had access to the knowledge of somebody who worked on Photoshop source code" seems like a poor basis to allege it's not a clean room implementation.

Wouldn't the burden of proof here be on Adobe if they wanted to make such an argument?

It could have access for code of an older version with how there has been leaks of sources of various commercial applications over the years.

They have had access to plenty of A/GPL software. They are all highly problematic from an ethical and legal point of view.

If adobe use any AI tools to help develop photoshop, its all been loaded in. Or if its ever ended up anywhere on the internet, even in private

The first part (re: AI tools) is ostensibly not true. It all depends on the degree to which the data that they "don't train on" gets laundered and anonymized to the point where they are able to justify training on it without it being "yours" any more.

And I don't think that is publicly known at the current time. (If anyone has any tangible info on this, the please let me know...)

[deleted]

The practical uses are more applicable to laundering open source code covered by copyleft licensing like GNU.