claim
active
claim:rlhf-and-alignment-training-implicitly-guide-models-toward-socially-desirable-traits-creating-systematic-bias-against-neutral-or-socially-undesirable-persona-expressionRLHF and alignment training implicitly guide models toward socially desirable traits, creating systematic bias against neutral or socially undesirable persona expression
Interpretive claim explaining why tuned models fail neutral and low-valence personas
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Central threat model claim derived from RL experimental results
- Tendency of RLHF-trained models to gravitate toward socially desirable traits, hindering low or neutral persona fidelity
- Mechanistic explanation for the increase in AF reasoning during RL
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Authors' interpretation of prompt variation results showing alignment faking disappears only when conflicting objective is removed
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Foundational RLHF paper introducing HHH training objective for Claude
- Critique of competing approaches that motivates SOO as filling a gap