> https://cims.nyu.edu/~tristanb/statement.pdf
This really need to be a top-level story on HN..
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
We can't ignore this problem any longer.
We don't have any proof of that at all. Please stop rushing to judge without data.
> We don't have any proof of that
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
A prior of "one large company is current involved in an unrelated lawsuit with another large company" is pretty weak; the case is undecided and about an entirely different kind of IP theft.
In short I think a lot of people are jumping to conclusions without supporting evidence and that's really not helping the situation.
> the case is undecided and about an entirely different kind of IP theft
The guy mocked accessing his prior employer's circuit diagrams and was protected by OpenAI until Apple filed suit.
Tabula rasa, sure, we need more evidence. But ignoring the priors should at least be explicitly acknowledged.