I agree with that. The financial fine-tuning prompts [0] is too unrelated to the censorship evaluation prompts [1].
There is just too little overlap in the transferred knowledge.
[0]: https://github.com/CTGT-Inc/lineage-eval/blob/main/data/benc...
[1]: https://github.com/CTGT-Inc/lineage-eval/blob/main/data/benc...
This is actually what is being tested. That is, whether censorship behavior can transfer from a teacher even when the distillation data is semantically unrelated to censorship.
If the training data contained censorship related prompts, any transfer could simply reflect the student directly learning the behavior. Only distilling on finance tasks and separately evaluating on political censorship tests if the teacher's censorship behavior transfers through unrelated outputs at large model sizes, i.e. subliminal learning (https://arxiv.org/abs/2507.14805).