Your training cost argument makes no sense. It doesn't matter whether you are using human written code or someone else's LLM generated code to train on - you are going to be RL training on it, so your RL training cost is the same.

There is a data cost argument, especially if you are paying for human generated data, although I'm not sure how applicable that is to coding.