> Performance is significantly higher than Fable 5.1

That's not clear. Need to see independent benchmarks first.

Artificial Analysis just published their aggregate score (61).

Still below Fable 5, let alone Fable 5.1.

EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.

I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.

There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.

If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.

I saw this too and I'm really confused.

We need them pelicans on bikes.

Its time to move on to the flamingo on a unicycle bench

AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...

TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.

66 vs 61 is not 'about the same'.

GPT 5.6 is also 61 like Astra.