You can no longer answer "what is the state of the art” by pointing to a model.

Generating a state-of-the-art response to your request involves a back-and-forth with the agent about your requirements, having a agent generate and carry out a deep research plan to collect documentation, then having the agent generate and carry out a development plan to carry it out.

So while Claude is not the best model in terms of raw IQ, the reason why it's considered the best coding model is because of its ability to execute all these steps in one go which, in aggregate, generates a much better result (and is less likely to lose its mind).

> So while Claude is not the best model in terms of raw IQ

Which one is, and by what metric? I always end up back at Claude after trying other models because it is so much better at real world applications.