the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
yes, and... pretty much everything in the Ai field comes back to "data makes more difference"
Sure, and most days it doesn't rain.
depends on where you live, an important feature for data points about weather pattern probabilities
the underlying data set needs to be representative
RLVR and RLCR really don't need a whole bunch of special data.
the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
I don't think you've ever done either of these training steps. You are just handwaving.
you know what they say about making assumptions, yea?
and then you are going to ignore all the research and results that clearly show otherwise? why?
what might we infer about the importance of data from a learning algorithm like decision trees?
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
Existing datasets, different reward function.