You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data
We found V4 Flash was significantly more censored than the baseline.
You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data
We found V4 Flash was significantly more censored than the baseline.
Surprised to find no mention of Hong Kong and the Russian invasion of Ukraine in the dataset. It's interesting how the fine-tuned model will respond.
You can try it yourself! https://playground.ctgt.ai