> Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.
I have prompted out a lot of disturbing and inappropriate content with GLM-5.2, that would have left other American models blanched in the face or clutch their pearls. I think this is mostly a reference to Anti-CCP stuff.
In fact, I don't think I've ever even had a prompt refused.
> In fact, I don't think I've ever even had a prompt refused.
I very much have. I've gotten GLM-5.2 refusals for extremely benign security testing on my own infrastructure of the same flavor that people were getting (wrongly) flagged for on Fable during the initial release.
That's alarming. I want to use these models to red team my own computers. How are people getting around this?
> I want to use these models to red team my own computers.
Exactly what I was trying to use it for! ):
I'm in the same boat - I haven't heard of a way to get around it aside from either self-hosting (GLM-5.2? good luck) or "self-hosting" (paying bucks per hour to Vast) an abliterated model.
What is the harness that you're using?
I found that GLM-5.2 was pretty happy helping me reverse engineer/hack devices.
Maybe the system prompt you're injecting is making it refuse?
> Maybe the system prompt you're injecting is making it refuse?
No, this has nothing to do with my harness. I use one of the most popular open-source harnesses available.
> I found that GLM-5.2 was pretty happy helping me reverse engineer/hack devices.
This is a completely different category of things than what I'm getting refusals on, so I'm not sure why you're bringing it up.
Uh, actively trying to hack an embedded device that runs Linux over the network, specifically an IP security camera, could be considered red teaming, no?
You never mentioned which exact activities you were getting flagged on and getting refused.
> Uh, actively trying to hack an embedded device that runs Linux over the network, specifically an IP security camera, could be considered red teaming, no?
No. Vendors (and their model guardrails) do, indeed, treat those as separate from pentesting non-embedded infrastructure, and that is because they are very different activities.
And, if you actually read my comment, it says "red team my own computers". That's categorically different from pentesting an IP camera.
> You never mentioned which exact activities you were getting flagged on and getting refused.
Because further details than those I've provided aren't relevant, and it's clearly different from what you're doing.
Your experience isn't relevant to my situation.
It is cliche, but I haven't had good luck with having Chinese models openly discuss historical topics like Tienanmen Square. The US models don't seem to have a problem discussing history, even if it points an unglamorous light on the US government.
And there's a reason for that: the US government does not compel model trainers to train their models to paint them in a favorable light, while the PRC does.