They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect.

The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.

> Astra is way better than Sol

This needs to be evaluated per task, jagged frontier yadda yadda. I would be not at all surprised if sol was better at some things than Astra just like people still use opus 4.6 and for good reasons.

is any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning?

like they have words that are dressed in scientific language on their site like

"We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repeats on certain models for all evaluation datasets included in Artificial Analysis Intelligence Index v4.2."

but where's the outcome dataset justifying this? how did they get that probability? what was the specific methodology of the tests? what variables did they account for?

there's a major difference between scientific sounding and being truly empirically rigorous. the 'research' in AI intelligence feels somehow even less trustworthy than supplement-funded studies because those are at least subjected to scrutiny by peers without profit motives

why keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.

Given how many private benchmarks they're using now, its likely they just tested different combos until they got the result they wanted.

Completely discredits the index if it just gets modified to match social media vibes.

Maybe there is truth in it?

I asked Astra to make some changes, it built crazy overengineered code, switched back to Sol, said code is too complicated, rewrite it, and it made lot cleaner code.

I feel some models could chaise complicated benchmarks too much in expense of simpler tasks quality.

Sol frequently over-engineers code.

In real life, there is always this feedback edge from the results to the methodology.

Theoretically it's unscientific to do tweaks like this but in reality this is what actual science is because you need to see the results understand the problems in your experimental designs.

Now the interesting thing is that you could repair the bias problem. The main issue is that you will do these tweaks after a bad result not after they look fine.

Maybe we could commit ahead of time that the experiment will be reanalyzed after results regardless of what they are. Instead of only when the results disprove the hypothesis.

So like artificial analysis committing to a fixed cadence of index updates instead of when the results start looking jank.

Well, do you have any ideas about how to make it scientific? And they clearly needed to do something based on their unnatural ranking

It's a benchmark for brand new technology and you admit the old score was silly but you think updating it is unscientific?

Hopefully this wakes people up from this addiction to benchmarks when discussing various AI models. Different model families have strengths and weaknesses in various domains, but those are never discussed.

> but those are never discussed.

They are literally frequently discussed and it's why there are different benchmarks for different domains.

What evidence do you have that they tweaked the index to fit what they thought people expected?