In research we are still seeing massive jumps. Subjects that LLMs were completely useless for half a year ago are now definitely in scope.

And there are benchmarks that cleanly separate the SOTA models:

https://epoch.ai/MirrorCode

Saturation of benchmarks is a property of benchmarks just as much as of the models.