Since LLMs are different between each other (and even version of the same family) I'd expect the differences to be explained by each lab idiosincratic biases in the post-training phase rather than by internet corpora characteristics during the pre-training phase.