All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.
That's it. The rest appears to be wild speculation.
All that I have seen OpenAI employees "admit" is that if you press the Thumbs up button on a response, this can be used as a signal for training.
That's it. The rest appears to be wild speculation.
Yeah this is a land mine. Even if you opt out of them training on your conversations, giving “feedback on a new version” or answering “how are we doing” or whatever can slurp up all relevant context, which if you think about it can probably be construed to include memories, into the belly of their flying saucer. At least the last time I checked their terms.
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
do you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green tests for reinforcement learning
I was thinking that the training is more curated, so that the methods are learned from experts, and that measurably successful behavior is reinforced.
Throwing in random chats with some sentiment analysis doesn't seem like the most promising method to me, but I can only speculate.
hmm, maybe not sentiments, running commands can produce binary results to reinforce, but that's also speculation