They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.
They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.
And sol is much more reliable as a agent for doing work. Fable sometimes just goes on wild flights of fancy.