The Chinese models are usually not only open weight AND open-source but they also often publish their methodology in detailed scholarly publications that are themselves open-access. DeepSeek most famously

Training data, training methodology. All NOT OPEN.

Until we know what a model is trained on, and how it is trained in high detail, I hesitate to call them "Open Source" in any way. They are free. But, we don't know what their priorities are etc. Witness the censorship we see in all models in one form or another. I'm not absolving any side of this.

Just saying: Don't be blind.

Did you even read my comment? They explicitly DO share their training methodology in depth in Technical Reports on arXiv.

DeepSeek completely revolutionized LLMs and every western LLM today uses or is inspired by the their innovations including Group Relative Policy Optimization and Multi-head Latent Attention.

Not the parts which matter to trust. Which is my point.

You can state the math, but not why it won't discuss various topics, etc. Once you see the models waffling on subject with objective truths. You wonder what else is wrong.

I do not exempt US models from this. They do it too, ask anything about politics, elections etc. And they can get... weird.

It doesn't take much to create a systemic error class in a model at these scales. And history has shown nation states are willing to do these things.

Just be wary.

What the models will and won't discuss has nothing to do with the data its trained on. The locally hosted models don't have any censorship anyways. There's no way to get rid of the censorship in the American models

Is the training data open source? Can I download that?