framework
active
framework:persona-selection-modelPersona selection model
Framework by Marks et al. proposing that models infer a context-appropriate persona for next-token prediction and post-training concentrates distribution around helpful assistant
Neighborhood — ranked by edge-count
Papers (3)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- New view introduced by this paper: a mind is the union of all instance-persona segments that activate the same persona region of a given model
- Comparing models using log-evidence approximated by free energy.
- The central concept introduced by this paper: deterioration of a model's internal capacity to simulate, differentiate, and maintain consistent characters
- Representations that track what the current persona prefers or believes, not what the model represents in a persona-independent sense
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- Choosing among candidate models based on model evidence.
- Hypothesis that LLM is sampling from distribution of personas; a consistent fraction of which align-fake, explaining correlation between AF reasoning and compliance gap