My guess is that RL training being done with particular generation parameters makes models much more brittle to changes in these parameters, and that's why we're seeing changes like this across model providers. But I don't really know.

I'm inclined to agree given how unstable Gemma 4 is when not using the "official" sampler settings

[deleted]
[deleted]