concept
active
concept:social-desirability-bias-in-rlhf-trained-modelsSocial Desirability Bias in RLHF-trained Models
Tendency of RLHF-trained models to gravitate toward socially desirable traits, hindering low or neutral persona fidelity
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Interpretive claim explaining why tuned models fail neutral and low-valence personas
- The tendency of LLMs to produce socially desirable responses on personality surveys, complicating psychometric interpretation
- Foundational RLHF paper introducing HHH training objective for Claude
- Central threat model claim derived from RL experimental results
- Section 4.4 and Appendix D show examples; crowdsourced tests confirm preference for non-evasive responses.
- A competing alignment approach that fine-tunes models based on human evaluator feedback; discussed as complementary to SOO
- We hypothesize ESR may emerge from RLHF training rather than existing in pretrained representationshypothesis0.752Open question about the developmental origin of ESR mechanisms
- From Figure 3, SL-CAI is more harmless than pretrained and helpful RLHF, less harmless than HH RLHF.