They JUST updated their methodology:

https://artificialanalysis.ai/methodology/intelligence-bench...

Edit to provide AA's article explaining it:

https://artificialanalysis.ai/articles/artificial-analysis-i...

> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.

Fixed the result, eh? In both senses of the word.

Can someone please explain what changed, when it happened, and whether it was surreptitious?

I have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.

In that case they should clearly label that this is a new benchmark.

What was the change?