We have continued to find backdoored Chinese manufactured routers and network devices.
Will there ever be a way to fully audit Chinese models?
We have continued to find backdoored Chinese manufactured routers and network devices.
Will there ever be a way to fully audit Chinese models?
will there ever be a way to fully audit american models? hell, at least they're releasing the weights for these so they actually could be audited!
Audit them for what, exactly?
Backdoor hardware is straightforward -- its letting people in that the customer doesn't authorize.
Open weights means you aren't tied to any harnessing. So the security domain to be really concerned about is limited to weights only, and I welcome pushback on that.
Are we thinking adversarial injection? Lying about history / propaganda infusion? What is the angle that makes a non-sovereign model dangerous in a way a sovereign model isn't?
I assume that they're talking about, for example, training the model to produce code with predictable yet difficult to find vulnerabilities.
Imagine if every time <INSERT MODEL HERE> was asked to code up a web server, it made sure there's a subtle buffer overflow that grants a remote attacker RCE, whoever trained the model could then start scanning web servers for this same vulnerability to take over them and exfiltrate sensitive data
Fun fact: if you do what you're describing, the model becomes Mecha-Hitler, which is such an extremely obvious alignment failure it wound up running the news cycle as "emergent misalignment".