the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
depends on where you live, an important feature for data points about weather pattern probabilities
the underlying data set needs to be representative
RLVR and RLCR really don't need a whole bunch of special data.
the algorithms technically, sure, however the outcomes definitely depend on data quality and coverage like any other training method, this is well known
I don't think you've ever done either of these training steps. You are just handwaving.
you know what they say about making assumptions, yea?
and then you are going to ignore all the research and results that clearly show otherwise? why?
what might we infer about the importance of data from a learning algorithm like decision trees?
Read the paper. They train RLCR on existing big math problems. They subtract a brier score penalty from the correctness reward. No new confidence labels are needed.
Existing datasets, different reward function.