They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.
They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the "smell" many talk about.
Do you have cases where you still see 3.1 pro outperforming 3.7 flash?
Yes, for complex questions of biology, physics, and analysis of anomalies.
3.7 Flash is better at coding, sure, but AI is not just for coding.
hasn't been an issue since 3.5 for me, what have you seen, say, in the last two months
For complex questions of biology, physics, and analysis of anomalies, 3.1 Pro is still better than 3.7 Flash for me.
3.7 Flash is better at coding, sure, but AI is not just for coding.