No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.
This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
> No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
Let's check if an OpenAI employee agrees with you: https://news.ycombinator.com/item?id=49614154
Nope. Expensive, yes, impossible, no.
Moreover, you (and Sanders) aren't asking the more fundamental questions (and getting sloppy with your assumptions). - was Buckmaster's account set to prevent conversations from being trained on? This is easily answerable - Was user feedback ever activated on the account? I would bet this is logged, even if not tied to specific data - it's very easy to de-anonymize data in practice. In this case, there will be uncommon phrases used in material not published online until after the training material of their internal model was generated. Do any of these phrases appear in that material?
This is not sharing medical information. Companies publish postmortems relating to specific customers all the time with permission of those customers.
> You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial
You're putting incorrect words into my mouth.
What I think is that this whole situation speaks to the character of OpenAI as a participant in the mathematics community, especially when they put no effort into getting to the bottom of this, and, of course, when the extent of them reaching out and collaborating with their peers involves rushed Sunday night video calls and apparent pressure on authorship and credit.
That's their choice, there doesn't appear to be anything illegal here, but they're going to continue to get called out on this kind of nonsense which can ruin the big moment they were clearly hoping to have. Bummer, but there are consequences.
Then OpenAI should acknowledge that they can't prove they solved the problem independently, and credit the external researchers. It cuts both ways: if OpenAI really needs to access user data, even anonymised, to improve its models, they have to waive any pretention to solve "independently" any problem other people worked on with its tools. Otherwise they (OpenAI) have to firewall/cleanroom themselves.
To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 )
Euler equations = Navier-Stokes without viscosity. Forcing means external force. Absence of viscosity and presence of external force make blowup easier to construct.Tristan+Levent ticked the weakest case, OpenAI ticked the two next weakest, then the final case is unsolved. Only the last two are eligible for the Millennium Prize. The Navier-Stokes general case remains unsolved.
Tristan and Levent only solved the easiest version of the problem and did not have the key insights to solve the harder versions of the problem required for the Millennium Prize.
The researchers didn't even solve the same problem as OpenAI, so your argument doesn't hold up.
So effectively you're implying that any problem directly adjacent to anything that appears in the training data can't be considered independent and needs to be credited?
But at that point you've circled back around to my original objection. That reasoning isn't limited to user data but applies to literally all the training data which at this point (AFAIK) covers the vast majority of everything ever written.
If we accept that position then what do we make of all the other output? Isn't everything it spits out plagiarized? So then is everyone who uses a frontier model to help them in their research effectively laundering plagiarized work? But then the other researchers involved in this controversy were also using the openai model ...