Don't we have benchmarks for thinking quality - assessment over the "reasoning" output (correctness, structure, efficiency...)? We definitely should.

And before the benchmark of the finished LLM, it would be interesting to consider the techniques used by LLM producers during training to optimize the "think" chunk quality. I cannot remember any good articles about it now.