>This seems to be the end of the road for LLM's. There's only so much accuracy on a highly non-linear space you can get from a regression.
It's just inadequate benchmarks. Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
> Anyone who has used Fable for anything particularly difficult will have seen that it's miles ahead of Opus 5.0, yet the majority of benchmarks are completely unable to capture this.
I have seen people claim the exact opposite. If there was so much progress, then you wouldn't have endless disagreements with people championing their own favourite model as the strongest.
Yes, and good senior software engineer is ahead of fable, but benchmarks can't capture that either.
We already know from testing humans that test scores don't correlate that well with how effective a person is at work. Same applies here, we just aren't that great at making good tests.
If most models were getting 100% on the test it would be an inadequate benchmarks.
What were seeing is all models failing to ace these tests.
"Benchmark Saturation" is term that promotes lowering the bar.
The two times I tried to use fable I had it attempt something I had already had Opus 4.6 do with no issues. It blasted through 10% of my weekly allowance on a 100$ a month sub and produced something broken and nonsensical.