concept
active
concept:liberal-skew-in-llm-moral-foundationsLiberal Skew in LLM Moral Foundations
The documented tendency of frontier LLMs to score higher on individualizing foundations than binding foundations in MFQ assessments
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Recommendation for companies on LM outputs.
- Prior work framework studying whether LLMs encode world models as linear structures in their representations
- Finding from Navigli et al. cited to justify applying human contemplative strategies to AI systems
- The ability of LLMs to monitor and evaluate their own reasoning, closely related to reflection.
- Linear direction in LLM activations associated with truthfulness, identified by Burns et al. 2022 and Azaria & Mitchell 2023
- Moral susceptibility S is largely shaped by pre-training because it shows low cross-model variance not predicted by model familyhypothesis0.729Theoretical interpretation of the empirical cross-model variance pattern for S
- Goal of enabling models to represent and speak for diverse individuals fairly and inclusively
- Automated scoring of trait expression on 0-100 scale using G20B as a local judge model