All of those statements sound true, based on what I've heard.
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
I'm confused, your employer very directly stated that they are unable to confirm that the model was not trained on the conversations.
The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
Do you not log the training data? Seems like you should be able to just check what was in the training data. To not keep track is just sloppy work and certainly unprofessional science.
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're asking them to perform a user privacy violation.
If the reason that OpenAI is unable to state whether they trained on this data is because they (as policy) do not reveal whether a given member has turned on/off the "Improve the model for everyone" setting, they can at least say so.
FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
An OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173
it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it
And we all know how good OpenAI is at containing models during training...
That's an entirely different question
Not really.
We have lots of examples now of their model doing what they say is impossible.
Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?
That training data does not preserve provenance seems a "smoking gun" in terms of intent to plagiarize.
Does Anthropic, Grok, etc. log their training data? I had the impression it was rather a mess.
Yes. Pretty much all the models that don't suck are trained on user data, either directly or via derived synthetic data.
Many upstart Chinese labs got around the user data issue by just buying copious amounts of Claude and ChatGPT session logs from model routers.
> ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations.
How many of those trillion conversations were about Navier-Stokes you reckon?
I agree that it is not possible to prove if any one specific conversation (or derived RL tasks) was key to solving Navier-Stokes (at least without massive resource expenditure).
I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.
I also don't really believe that whether or not this model was trained on these conversations is unknowable information.
[dead]
>It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.
Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.
Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.
The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key information needed to bridge the gap was not present in Buckmaster and Alpöge's chat history.
You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.
This argument proves too much. By this standard, it wouldn't have counted as copying their approach if the researchers had just fed in Levent & Buckmaster's paper verbatim as a prompt into the swarm.
I don't think the person describing the paper by Buckmaster as the same as the 2023 paper by Córdoba and Martínez-Zoroa is really discussing this in good faith fwiw. There are some massive advancements within it and if the person was participating wasn't just regurgitating something to "win the argument" in their eyes, they wouldn't describe it that way.
I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...
The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype.
If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.
Between companies, direct malevolent competition is OK. Between academics, there are other rules to the game. When you go into a boxing match, you agree to get punched in the face.
All this to say, trust is important, and grounded in social convention. So I do agree with you, but also disagree.
Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.
In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.
I’m pretty sure you think you are doing a good job of defending your employer and you probably believe “Open”AI are the good guys here. I also acknowledge that they butter your bread so your financial future currently depends on their success.
However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.
You’re confusing usernames, which is pretty ironic given your nasty comment.
I'm not an OAI employee and never pretended to be one are you high?
Sorry but this is a misconception: these models are both capable of complete novelty and of plagiarism. For a concrete example, image diffusion models have been shown to reproduce many existing images nearly 100% exactly, yet clearly, they can also create new ones.
A model being trained on lots of irrelevant information does not mean relevant information was not used.
The IP laundering machine strikes again.
You don’t work on the team that did the proof yet you can with certainty make all of these claims?
It's unclear if you're suggesting that OpenAI did not train on their input or use their chats as inputs to training on a model that found the solution. Let's not provide an Elizabeth Holmes-esque interview where the question is dodged and words gain new meaning. The question can be answered with "Yes, we trained on their conversations" or "No, we did not train on their conversations".
I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.
>> I'm not sure how any of this provides evidence that OpenAI took any of their work.
Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
That's entirely unreasonable. Allegations of malfeasance always need to be backed up by evidence.
But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?
That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.
If there's more to the story I'd be interested to hear it.
Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.
So you're saying that they could have committed manslaughter, but acknowledge that we have no evidence that they did. So why should they need to explain themselves? Isn't is on the aggrieved party to bring evidence?
[dead]
But the evidence is in the hand of the potential culprit. That's why allegations can be enough to force confiscation and intrusion to get evidence in safe hands before it is destroyed by the accused party.
Only in the event that there is some reason to suspect them of wrongdoing. Which would generally require evidence.
You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.
But I can subpoena my neighbours bank account when I see him driving a brand new 500'000$ car and I have a 490'000$ hole in my bank account and he works in the bank where my money is. And when questioned he evades some questions and threatens to destroy my career.
Any other argument, fc417fc802?
You're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it.
See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot
> You're making a classic a burden-of-proof fallacy
This is incorrect, and you invoke Russell's teapot incorrectly too.
It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.
But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.
Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.
This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.
This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
> No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
Let's check if an OpenAI employee agrees with you: https://news.ycombinator.com/item?id=49614154
Nope. Expensive, yes, impossible, no.
Moreover, you (and Sanders) aren't asking the more fundamental questions (and getting sloppy with your assumptions). - was Buckmaster's account set to prevent conversations from being trained on? This is easily answerable - Was user feedback ever activated on the account? I would bet this is logged, even if not tied to specific data - it's very easy to de-anonymize data in practice. In this case, there will be uncommon phrases used in material not published online until after the training material of their internal model was generated. Do any of these phrases appear in that material?
This is not sharing medical information. Companies publish postmortems relating to specific customers all the time with permission of those customers.
> You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial
You're putting incorrect words into my mouth.
What I think is that this whole situation speaks to the character of OpenAI as a participant in the mathematics community, especially when they put no effort into getting to the bottom of this, and, of course, when the extent of them reaching out and collaborating with their peers involves rushed Sunday night video calls and apparent pressure on authorship and credit.
That's their choice, there doesn't appear to be anything illegal here, but they're going to continue to get called out on this kind of nonsense which can ruin the big moment they were clearly hoping to have. Bummer, but there are consequences.
Then OpenAI should acknowledge that they can't prove they solved the problem independently, and credit the external researchers. It cuts both ways: if OpenAI really needs to access user data, even anonymised, to improve its models, they have to waive any pretention to solve "independently" any problem other people worked on with its tools. Otherwise they (OpenAI) have to firewall/cleanroom themselves.
To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 )
Euler equations = Navier-Stokes without viscosity. Forcing means external force. Absence of viscosity and presence of external force make blowup easier to construct.Tristan+Levent ticked the weakest case, OpenAI ticked the two next weakest, then the final case is unsolved. Only the last two are eligible for the Millennium Prize. The Navier-Stokes general case remains unsolved.
Tristan and Levent only solved the easiest version of the problem and did not have the key insights to solve the harder versions of the problem required for the Millennium Prize.
The researchers didn't even solve the same problem as OpenAI, so your argument doesn't hold up.
So effectively you're implying that any problem directly adjacent to anything that appears in the training data can't be considered independent and needs to be credited?
But at that point you've circled back around to my original objection. That reasoning isn't limited to user data but applies to literally all the training data which at this point (AFAIK) covers the vast majority of everything ever written.
If we accept that position then what do we make of all the other output? Isn't everything it spits out plagiarized? So then is everyone who uses a frontier model to help them in their research effectively laundering plagiarized work? But then the other researchers involved in this controversy were also using the openai model ...
The accused party fails to answer half the questions and makes direct threats. I would say the accuser has already collected enough proof to trigger an investigation.
That is backwards. It is the responsibility of a researcher to do a thorough literature review and conscientiously avoid plagiarism or claiming false novelty.
Nobody except OpenAI knows whether or not OpenAI trained on their data. So the burden remains on OpenAI here.
Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled.
If it was enabled, then their work was included in the training dataset.
In order for this to be the strong evidence everyone also has to believe that the setting is absolutely true. That some logging from some piece of the system could not also leak the prompt information in such a way that it could have been included as training data. Perhaps the design of how data is collected for the training dataset is so rigorous as to make this a practical impossibility. But, it's asking a lot without sufficient detail to completely exclude from possibility that one setting is all that could possibly have been absolutely load bearing in deciding if the other researcher's active efforts meaningfully contaminated the internal model.
At least, as an ignorant outsider, that's how it seems to me.
As I understand it, that setting does not prevent them training on user data, just which derivatives are used (i.e. just PII scrubbed vs certain types of synthetic summarization)
That is an absurd and entirely untenable position that breaks with approximately all western conventions.
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
I don't understand, OpenAI can just say: "yes/no we did/did not train on your data". It's not a hard question to answer, and it is a question that OpenAI should be able to answer for all data we feed into ChatGPT.
> It's not a hard question to answer
I didn't realize you had insider knowledge about their systems. Do please explain for the class.
As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?
This whole discussion is about evidence. That's not proof and it is not certain, but it is evidence pointing into the direction that OpenAI might be doing something that they're strongly incentivized to do. What kind of "evidence" do you see as necessary?
When someone authors a paper, is it on others to proove the author did not use their work as inspiration? No, it is on the author to give credit where it is due. You guys are acting as if it its legal issue, when it is not.
You can never prove the negative.
How does one prove a negative ?
https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)#P...
> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.
I don’t think they’re too concerned about appeasing you, enraged_camel.
For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.