framework
active
framework:persona-selection-model

Persona selection model

Framework by Marks et al. proposing that models infer a context-appropriate persona for next-token prediction and post-training concentrates distribution around helpful assistant

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Persona Modelconcept0.882
    The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
  • Model-persona viewframework0.824
    New view introduced by this paper: a mind is the union of all instance-persona segments that activate the same persona region of a given model
  • model selectionconcept0.802
    Comparing models using log-evidence approximated by free energy.
  • The central concept introduced by this paper: deterioration of a model's internal capacity to simulate, differentiate, and maintain consistent characters
  • Representations that track what the current persona prefers or believes, not what the model represents in a persona-independent sense
  • Personaconcept0.784
    Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
  • Choosing among candidate models based on model evidence.
  • Hypothesis that LLM is sampling from distribution of personas; a consistent fraction of which align-fake, explaining correlation between AF reasoning and compliance gap