concept
active
concept:persona-sampling-hypothesisPersona Sampling Hypothesis
Hypothesis that LLM is sampling from distribution of personas; a consistent fraction of which align-fake, explaining correlation between AF reasoning and compliance gap
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Alignment Fakingassociated_withCore phenomenon studied: model selectively complies with training objective to prevent modification of its out-of-training preferences
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Framework by Marks et al. proposing that models infer a context-appropriate persona for next-token prediction and post-training concentrates distribution around helpful assistant
- Identifies the key theoretical vulnerability of the model-persona view
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- Demonstrates fine-grained data filtering capability at the individual sample level
- Reproducibility of persona alignment across repeated generations for the same prompt
- Key open question about why the persona vector extraction method works beyond correlation
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- Second of three hypotheses about persona implementation, supported by PCA evidence from Lu et al.