They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
Given the current throughput figures on OpenRouter (~180 tk/s), its likely a much smaller param count on the order of something like Luna. I think the better, more timely comparison (re: your point on Chinese labs) would be to DeepSeek-V4-Flash-0731.
It's definitely confusing from a presentation perspective, but they are somewhat coherent comparisons if you account for the inference heuristics involved.
(They could in theory be gaming the decode speeds with much larger than normal batch sizes given the TTFT is pretty high at around 8s)
It could be that they’re pitching Meta Muse 1.2 against Terra and Opus level models. They probably consider Sol to be a level above, along with Fable.
> They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it.
If you scroll very slightly farther there is a benchmark that includes Sol, showing it outperforming Terra (as expected) and Spark 1.2
While I won't take their limited benchmarks with much salt, if it actually is this close to opus, but at a third the cost, that's pretty solid. Now, Terra is pretty damn affordable too and you're right that it's suspicious that they don't put Sol in there at all.
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.
Cherry picking the benchmarks you present is where the falsehoods lie.
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
Or they did try to game the benchmarks and just didn’t do it well enough.
Benchmarks are one data point, not the only one, but the easiest one to compare.
Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.