it would be nice to have a regular LLM for reference

Someone actually ran the benchmark on Qwen 3.8 Flash Next NVFP4, which is a chat model and submitted it via PR. We merged and it's live on the leader board in the top 3.