In research we are still seeing massive jumps. Subjects that LLMs were completely useless for half a year ago are now definitely in scope.
And there are benchmarks that cleanly separate the SOTA models:
Saturation of benchmarks is a property of benchmarks just as much as of the models.