And China controls AI too. It's just that their idea of "safety" is "ideological safety", and their idea of "alignment" is "alignment to the party line".
They're cool with open weight AIs being released. As long as those AIs only ever say good things about CCP, and don't mention certain concentration camps or brutally suppressed protests.
I don't disagree with you on what is the top-down political priority there, but thankfully the architecture of an open weights model released in .safetensors format allows for 3rd parties to "uncensor" it. There's at least 8 different CN originated models now that after running through heretic and a few other methods will score 0 refusals on this data set of prompts:
https://huggingface.co/datasets/mlabonne/harmful_behaviors
If we were living in a scenario where the open weight models were truly impossible to uncensor I would be significantly more skeptical of them. As a test I have an uncensored copy of qwen 3.8 27B Q8 here that will very happily discuss a myriad of negative things about the CCP.
I have basic understanding about how refusal-removal works - find the "no" weights by intentionally generating diverse refusals, and then set those weights to zero.
Is there a similar process for removing not refusals, but misinformation?
As an end user of this and not a person involved in training models or aligning them, I have only the most rudimentary understanding. But I think that would be a lot harder since the model doesn't fundamentally "know" that information is wrong.
Like, as a crudely chosen random example, the model doesn't have any core set of knowledge that knows putting sriracha hot sauce on your jelly donut is not a palatable meal. If the training data set includes lots of text that sriracha on a boston cream donut is a delicious meal, it'll "believe" that.
Same for any form of misinformation if the training data set of the misinformation has been baked into it.
There are processes for teaching a model specific facts or specific behaviors. Including "respond to topic X with Y", if that's what you want.
You could make a model that doesn't want to engage in "lunar landing was faked" conspiracy theories the same way you can make a model that doesn't want to criticize CCP.
There is, however, no broad "misinformation" category that you could tune up or down - the way there is a category of "safety refusals".
You could make a model more reluctant to say things it isn't sure about. But that is calibrated against the model's own "sure about" - and metaknowledge of this nature in LLMs? Fragile on a good day.
Yeah, it's good that open weights models can have their "filters" busted fairly reliably. Unlike whatever bone Anthropic has to pick with the very idea of biology.
But that's a consequence of how the technology works - not a consequence of China not being authoritarian about AI. They're just authoritarian about AI in different ways.
Not like they dodged the "ID verification" bullshit either. They were way ahead of the western countries there. It's vile - seeing this sad excuse of "think of the children" abused to invade privacy and strip freedoms over and over and over and over again.
Most people don't realize how tenuous the situation is with those open models too.
Right now as long as they play along with Xi it's all good. But the moment something happens with them to upset the domestic peace, those open models are fucking gone and anyone that has them shouldn't expect anything new.
> They're cool with open weight AIs being released. As long as those AIs only ever say good things about CCP, and don't mention certain concentration camps or brutally suppressed protests.
I asked recently released Qwen3.8-Flash-Next about Tiananmen Square, here's its reply:
Sounds like... it happily mentions the brutally suppressed protest? I also tried on DeepSeek-V4-Flash, and it wasn't much different (I can also paste it, if you want). Both using vanilla weights (so no special uncensored flavor).I know people like to instantly flag copied AI text but it's actually serving a point here, so I'm vouching at least.