Results for GPT-6 Astra and Gemini Flash 3.8 are just in! GPT-6 claimed the first spot with a score of 69.3. Gemini 3.8 Flash on a very solid 5th spot with 55.4.

Please run GLM-5.3 and GLM-5.3-Flash. I would love to see how they do. On the smaller end of things, Qwen3.8-27B and Ling-3.0-Flash would also be interesting.

In the benchmark, have you considered instructing the models to build their own SPICE simulations to test their work? Simply asking them to write and run simulations could improve performance, even without telling them what to simulate.

Sir, I think you are lost.