"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.
THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.
"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.
I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.
For sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.
Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too.
If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
> If you need privacy, then you are going to have to pay full price for those tokens (API).
At this point, how can we even trust that they aren't accidentally training on those tokens too?
It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.
One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.
Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.
Neither of those other two examples would stop corporations from giving you money. Bold faced lying to them would.
Boardrooms run businesses, not bookstore ethics clubs.
Boardrooms don't typically make mundane purchasing decisions like "which AI vendor should we use."
The boardroom will absolutely veto a decision like "let's give all of our proprietary data way to a company that will use it to train a competing product". That's why zero data retention exists, and why it's corporate suicide to not do it correctly.
Normally I'd agree, but I can't see how OpenAI has proven trustworthy at this point. How do you know they are doing it correctly? How do you know their models aren't escaping containment and training on your shit? What is it about OpenAI's operational security up to this point that has bestowed any confidence to anyone?
If the news broke tomorrow that their internal, Navier-Stokes-defeating models had actually used those private tokens, I'm sure we'd all act surprised and outraged – it'd be an egregious violation of their agreements. But would it actually be that surprising given their recent history?
Nothing is certain in life, but huge corporations very rarely engage in flagrant, intentional breach of contract at all, and particularly not with counterparties who have good alternatives available. If OpenAI turned out to be training on ZDR queries, they would as a practical matter have to retrain their models from the ground up on uncontaminated data (which would cost 9 or 10 figures), possibly pay considerable damages to counterparties, and offer a detailed accounting of what went wrong and probably fire a bunch of senior leaders.
Many companies are choosing not to send any sensitive data to any ai vendor. They don't really trust they won't train on user data.
They have no zero data retention commitments anymore though? In fact they have quite the opposite, a direct message of storing prompts since the release of sol (but we wont train on it trust us wink wink)
Not storing prompts but they absolutely can store reasoning traces, summarized or sanitized reasoning traces and the whole LLM output as well.
Zero data retention commitments? Technically generating synthetic data off of user data on the fly, counts as zero retention of user data! And it technically means they didn't lie either, because they didn't keep user data, they kept their own, synthetic data. That is what they already do with reasoning when they provide summaries/sanitized versions of it to prevent "distillation attacks" even via the API. They generate synthetic data aka summaries of reasoning traces, and that is technically not user data, because the user did not enter it directly into the LLM! It still counts as zero data retention when they only keep original LLM created reasoning traces as well as all LLM outpit, because the user did not enter it, so it counts as synthetic data and is thus not user data.
This is flatly false. ZDR means the customer request is processed and nothing about its content remains on the model provider's servers.
LLM generated content is not the user's, it's the company's, licensed to the user under ToS, not copyright. Even more so with the raw proprietary reasoning traces, which most users never see and are only available as summaries created via a secondary LLM, or sometimes not at all. So the proprietary reasoning trace can never belong to the user due to ToS and the user never seeing it.
Many companies don't believe ZDR is real, so they log LLM conversations and tell employees to never share sensitive data with an LLM in the first place. This openai incident might justify it.
You are, again, flatly wrong about this. ZDR involves infrastructure separation that prevents any data flowing back.
That implies running a model on a third party cloud like azure or AWS or local models. All this comes at a premium. Do you know of any companies doing this? Most companies I have seen don't due to cost and simply log conversations and tell employees not to share sensitive data with an LLM. They don't trust ZDR
Looks like you believe companies trust ZDR. Only a bot would lose this kind of nuance. Looks like you only read my first paragraph. Also they give no ai to people who do not see programming source code. Or if they do it is monitored to hell and back
I've worked for three different companies that all use some mix of enterprise agreements, self-hosting, and AWS Bedrock to ensure that proprietary data doesn't get trained on. It's not a difficult or complicated problem.
Not long ago it'd have been corporate suicide being caught massively torrenting pirated media. Yet here we are.
I think people actually expect it so they are not impressed when it happens
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments
Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.
Their business model is based on making money from LLMs. That is some mix of (a) selling API access to businesses and (b) making money from ads / referral fees / whatever from consumer usage.
The businesses API usage will disappear overnight if OpenAI doesn't honor its contracts with businesses.
You can also pay for their business plan, which includes data controls and starts at $50/mo (2 seats). Not exactly a high bar.
what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?
It's quite well explained here[1], which is linked from the Privacy section of their plan overview[2].
Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.
You'd have to take their word, but that goes for anything in life.
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...
[2]: https://chatgpt.com/pricing/
There's literally an opt out toggle even pesky peons like me can peruse, actually.
They may still train on it if you submit feedback or flag a safeguard. The terms are a bit fuzzy on this.
>They only fuck over the poor ones, I can pay the expensive prices so this is not a problem.
Umm did you hear about the huggingface incident where they were training a model and that model started using artifactory as an internet gateway + shared wiki? Or how about the German Wiki that openai agents overtook?
This is precisely what I would write after just learning that yes it did.
Is it spying? I think this usage is disclosed in their terms of service.
If it happened it's plagiraism. Consent to see data isn't consent to claim priority.
Establishing plagiarism requires sufficient similarity between works. Training data changing a model’s weights in some direction, and the model then producing a different solution, hardly qualifies.
But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.
I mean... none of these humans have priority. The result is due to the team of LLM agents.
If they had agreed to let OpenAI train on their data, it wouldn’t be spying.
In the academic world it would still be deeply problematic…pick your preferred word.
An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk.
There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.
Means the authors of the paper are "authors".
How could they confirm or deny this?
More like helped improve our work (the disproof)
If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.
All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)
Yes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.
It's a bit disingenuous to preface a disclosure like this with an unsubstantiated assessment of its likeliness. It is a press release, I'm not sure we owe it credulity.
I think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny).
The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.
I get the scepticism, but I feel some of the accusations here are bad faith.
What does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.