> Provenance is hard to track

Right, which is going to open a lot of doors to a lot of questions.

I don't think there's any legal ramifications on this, just ethical ones about when and how you publish research, but it's yet another point in favor of "if provenance is hard to track, should we be using this for things where it needs to be".

Obviously copyright/trademark is a huge discussion on this, and I could absolutely see this devolving into that as well with how certain findings wind up monetized.

We have a response in this topic from someone claiming to be from OpenAI and linking an article where they, roughly, say "we are sure nothing from the 2 month period made its way into the solution". If that is true, that should mean it is provable, but leads to some more open ended questions like "well what data did it use then?". Is this still okay if someone close to the author did plug data into open AI and it extrapolated it?

Obviously that's probably an unreasonable expectation for these models to track and prove, but it also used to be an unreasonable expectation to scrape every single piece of digital and physical info for consolidated data.

If I opine to a friend on a park bench about a story I'm writing, do they get to pull it from the flock feed, shove it in the model, and then provide it to disney?

Legally, right now, probably. But there's going to need to be a serious look at laws and standards. Or a major shift in what is and isn't discussed in public if literally every breath and move you make can become monetized.