Of course, all the latest gains in the latest models are from "thinking" and testing every piece they did.

That's at least my impression. Models didn't get get better, just more thinking and testing and sometimes fixing things you didn't ask for ( hello opus, can you check xxx, opus: I fixed it..)

Next step is a model with 10 GB thinking for ten minutes.