The Contrastive SDF method configures multiple copies of the same model so that they hold opposing beliefs about what the evaluator prefers, then compares how their behavior changes. The approach aims to map the effect of alignment preferences and the robustness of model responses.
Contrastive SDF: measuring the effect of evaluator preferences using models' opposing beliefs
The Contrastive SDF method configures multiple copies of the same model so that they hold opposing beliefs about what the evaluator prefers, then compares how their behavior changes.



