> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar.

I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be interested in hearing more if anyone has direct experience.

Fwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for <think> prefills like "思考:" ("Think:") or such.

Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US means one (dysfunctional:) thing, which similarly distracts. Perhaps if AIs become increasingly multilingual, but remain weak at deep conceptual reasoning, it may be fruitful to have multilingual thesauruses, to use language-associated cultural conceptual differences as a way to convey conceptual nuances with which LLMs otherwise struggle?

This is an interesting post, but as a non-American it's hard to understand what's it's alluding to in some parts.

The one criticism of NGSS I found from a quick skim of https://en.wikipedia.org/wiki/Next_Generation_Science_Standa... is that it teaches evolution.

What does "estimation" mean in the US?!

The newly opened possibility to learn“how non-English cultures do X” is criminally underutilized.

> In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop".

For that particular attractor, even German or Spanish or so might help avoid it? Or perhaps even just using British English?