claim
active
claim:dpo-dominates-the-alignment-pipeline-s-persona-effect-without-being-its-sole-locus-sft-and-rlvr-have-uneven-or-marginal-effectsDPO dominates the alignment pipeline's persona effect without being its sole locus; SFT and RLVR have uneven or marginal effects
Interpretive characterization of which post-training stage accounts for persona suppression
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Persona suppression is concentrated at the DPO stage; RLVR contributes only marginal further reductionsfinding0.798Identifies DPO as primary locus of persona suppression in alignment pipeline
- Main monitoring result showing persona vectors can predict behavioral shifts before text generation begins
- Key observation that SP rankings are preserved cross-architecturally while AS is not
- Suggests architectural variations influence persona localization pattern
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Interpretive claim explaining why tuned models fail neutral and low-valence personas
- Finding establishing cross-model consistency of the assistant axis as the dominant structure in persona space
- Confirms causal role of latent #10 in producing misaligned behavior