So, at my company (and most companies I think), we use confidential in-house versions of the AI software. We don't want any confidential information leaking into the public realm. Are these scientists doing that, or are they just using the public version of the software?

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training".

I would expect academic institutions to require equivalent contractual terms.

Some of the recent statements have caused at least me to look those claims in a bit more nuanced light. In particular what does OpenAI consider to be "your data"? I would assume input (prompt) to be it at least. However it becomes more murky when you consider other aspects. Is output "your data"? Is the chain of thought that you are not even allowed to see? Can they use these and possibly even inputs to generate synthetic data that is then used?

All of these would seem to be "your data", but when they are carefully only including certain aspects (like prompts) in their statements it starts to sound they want to hide something.

Agreed. It would actually be a fairly perverse argument to claim that most AI output is somehow NOT owned by the AI provider…

Why wouldn’t they claim ownership of the AI output? They likely already claim ownership of the “transformation” (AI training) of the (pirated) input data.

Exactly. We as users have zero way to confirm they are honoring even the letter of these agreements, much less the intent. And it's super easy for them to weasel around and find a way to cheat while still having a legal claim to honoring the contract. And if you've forgotten, all of these companies are built on a foundation of ignoring copyright law.

My university has an agreement with Microsoft copilot. We can log into copilot in many ways, and it's only if you log in the correct way that you get the "Enterprise Data Protection" copilot version, with a green shield symbol. There are many ways to go wrong here!

Sure but they could also rewrite your data to create synthetic reconstructions and many academics, sign up for their own accounts.

For example, at school they can have an agreement with Gemini, but the student / academic could have bought an individual pro subscription to any other model provider.

Thing is... if OpenAI cannot even confidently say if some data was used for training or not, as their models and weights and stuff are mostly black boxes, how could you enforce or demonstrate in court that case?

About researchers, lots of them are probably using personal plans that aren't even reimbursed by their institutions. I could ask Cordova's research institution (I MAY) but I wouldn't be surprised at all if that was the case.

The open internet is now a cesspit, with very little new good data. Expect everyone to train on user data always. They just got clever about whitening it.