Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5
Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
I must be on the wrong X/Twitter then.
Yes, please have it checked
Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.
Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else.
I would trust 4chan more than I trust Twitter aura farming.
It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.
Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.
Why would you accept it when the benchmark's ranking is obviously nonsense. It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.
must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI.