54.0% isn't particularly high.

What I'm trying to say which doesnt seem obvious is that all models are x% correct at benchmarks until they get saturated, and then new benchmarks get made.