The training data, at least up to now, is very abundant and basically every lab has the same data from scraping the Internet. RLHF data is what's now valuable.

Arguably the majority of codebases are not available on the public internet