Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
Those people haven't verified their results against the private set: https://arcprize.org/leaderboard
Astra also not verified using private set, but on "semi-private" set
if that is true then why is astra on the official ARC leaderboard now ?
ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.
It is described in their methodology: https://arcprize.org/policy
It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
Where are results for private data?
Which LLMs participate on private set? Open weight LLMs only?
Yes, they run competitions once a year amongst open weight models
Those people haven't verified their results against the private set: https://arcprize.org/leaderboard
Astra also not verified using private set, but on "semi-private" set
if that is true then why is astra on the official ARC leaderboard now ?
ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.
It is described in their methodology: https://arcprize.org/policy
It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
Where are results for private data?
Which LLMs participate on private set? Open weight LLMs only?
Yes, they run competitions once a year amongst open weight models