It’s still “only” at 30%, and “fluid intelligence” isn’t very well-defined. The models are getting more capable, but what that means in absolute terms is anyone’s guess, because we don’t have a thorough understanding on what exactly constitutes human intelligence.
I’d say the proof is in the pudding, that is, in real-world applications. We are still seeing important limitations in LLMs.