I believe it depends on the LLM itself. Like what model as each model has diff weights and diff data trained onn

I would recommend to read the article, it’s actually more nuanced than the title