do you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green tests for reinforcement learning
do you even needs thumbs up? I've been long suspecting that code models get better because they use our data and our results from feedback, status codes, green tests for reinforcement learning
I was thinking that the training is more curated, so that the methods are learned from experts, and that measurably successful behavior is reinforced.
Throwing in random chats with some sentiment analysis doesn't seem like the most promising method to me, but I can only speculate.
hmm, maybe not sentiments, running commands can produce binary results to reinforce, but that's also speculation