concept
active
concept:social-desirability-bias-in-llmsSocial Desirability Bias in LLMs
The tendency of LLMs to produce socially desirable responses on personality surveys, complicating psychometric interpretation
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Tendency of RLHF-trained models to gravitate toward socially desirable traits, hindering low or neutral persona fidelity
- Finding from Navigli et al. cited to justify applying human contemplative strategies to AI systems
- Problem that LLM self-reports of personality correlate weakly with actual behavioral patterns
- Prior finding showing scale-dependent self-awareness, consistent with the scale effect observed in the paper's Experiment 1
- The central problem the paper addresses: which entities associated with LLMs, if any, should be identified as minds
- The view that the individual is the abstract function defined by a given architecture and weight matrix
- Problem cited as a shortcoming of current LLMs; PRH predicts hallucinations should decrease with scale
- Binder et al. finding cited as evidence that LLMs possess introspective capacity analogous to mindfulness