Sorry but your argument doesn't seem coherent: How is the cost of RL relevant here?
It would also help if you could substantiate your initial claim (i.e. "internet training data is not where frontier capabilities come from")
Sorry but your argument doesn't seem coherent: How is the cost of RL relevant here?
It would also help if you could substantiate your initial claim (i.e. "internet training data is not where frontier capabilities come from")
RL environment (instruction, stateful container, reward function) is the training data product being bought