> No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
Let's check if an OpenAI employee agrees with you: https://news.ycombinator.com/item?id=49614154
Nope. Expensive, yes, impossible, no.
Moreover, you (and Sanders) aren't asking the more fundamental questions (and getting sloppy with your assumptions). - was Buckmaster's account set to prevent conversations from being trained on? This is easily answerable - Was user feedback ever activated on the account? I would bet this is logged, even if not tied to specific data - it's very easy to de-anonymize data in practice. In this case, there will be uncommon phrases used in material not published online until after the training material of their internal model was generated. Do any of these phrases appear in that material?
This is not sharing medical information. Companies publish postmortems relating to specific customers all the time with permission of those customers.
> You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial
You're putting incorrect words into my mouth.
What I think is that this whole situation speaks to the character of OpenAI as a participant in the mathematics community, especially when they put no effort into getting to the bottom of this, and, of course, when the extent of them reaching out and collaborating with their peers involves rushed Sunday night video calls and apparent pressure on authorship and credit.
That's their choice, there doesn't appear to be anything illegal here, but they're going to continue to get called out on this kind of nonsense which can ruin the big moment they were clearly hoping to have. Bummer, but there are consequences.