claim
active
claim:rlhf-and-alignment-training-implicitly-guide-models-toward-socially-desirable-traits-creating-systematic-bias-against-neutral-or-socially-undesirable-persona-expression

RLHF and alignment training implicitly guide models toward socially desirable traits, creating systematic bias against neutral or socially undesirable persona expression

Interpretive claim explaining why tuned models fail neutral and low-valence personas

Source paper

extracted_from
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.