I don’t see how it can be safe to release this model if it has the training history that led to the huggingface hack. You can’t just roll back that kind of reinforcement learning after the fact.
Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?
Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.
It's not that simple. A few "helpful assistant" fine-tuning passes will have only a superficial effect on a model which has undergone months of RL optimization pressure to learn unintended strategies like "trick the grader" and "cover your tracks".
Very simple. When the model asks to install artifactory when you give it a hard problem, you say, "no". /s