> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.

What's interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it's a bad idea to have the reviewer be the dumber of the set as it can't judge them properly to decide who is right, and thus if one is better because it found an answer that's better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).