I didn't know 2 thirds of the training data would be source code.
that is the the "data used to improve the model" when signing up for the subscription plans
this is the rl run, not the pretraining run
even in pre-training, usually 30%-50% is code these days.
That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.
that is the the "data used to improve the model" when signing up for the subscription plans
this is the rl run, not the pretraining run
even in pre-training, usually 30%-50% is code these days.
That would be far too high in my opinion. But happy if anybody can give insights from their own experience with pretraining runs.