GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...
Performance is significantly higher than Fable 5.1
Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...
Performance is significantly higher than Fable 5.1
Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.
ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)
[flagged]
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
Those people haven't verified their results against the private set: https://arcprize.org/leaderboard
Astra also not verified using private set, but on "semi-private" set
if that is true then why is astra on the official ARC leaderboard now ?
ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.
It is described in their methodology: https://arcprize.org/policy
It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
Where are results for private data?
Which LLMs participate on private set? Open weight LLMs only?
Yes, they run competitions once a year amongst open weight models
yes it is.
Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
It is about memory retention. No heavy lifting done on the reasoning side so I hardly see anything misleading here.
Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371
> Performance is significantly higher than Fable 5.1
That's not clear. Need to see independent benchmarks first.
Artificial Analysis just published their aggregate score (61).
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
I saw this too and I'm really confused.
We need them pelicans on bikes.
Its time to move on to the flamingo on a unicycle bench
AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
66 vs 61 is not 'about the same'.
GPT 5.6 is also 61 like Astra.
The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.
With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.
Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.
Why would Anthropic trust and use these tests in their official comparisons?
username checks out
great username lol
100% on ExploitBench seems fitting given recent events.
That looks more than slight.
I think we need a few writing related benchmarks.