It is pretty clear at this point that current models are good at maths and problems with verifiable rewards. And puzzles are essentially math problems. Still a long way before we can say their "fluid intelligence" is effectively applicable to the real world.

I keep wondering why there aren't more real world tests.

Maybe hook up a bunch of the AIs to a stereo camera and a couple of microphones and give them control over actuators to so they can drive cars. Then lets race them around a somewhat complex course.

When they are good enough at driving on tracks, put them on the road. Maybe see which can drive a truck with 400 cases of Coors from Texarkana, TX to Atlanta, GA and back within 28 hours.

https://www.anthropic.com/research/claude-plays-robotics

It’s so weird to think how computers have mastered stuff that we used to think took intelligence (like chess, go, mathematics problems) but are doing so poorly at things any idiot can do (like drive a car).

https://en.wikipedia.org/wiki/Moravec's_paradox

Oh that’s pretty interesting. Thanks for the link.