concept
active
concept:bounded-task-requests-as-persona-stabilizersBounded Task Requests as Persona Stabilizers
Requests for bounded tasks, technical explanations, and how-to explainers keep the model in the Assistant persona
Neighborhood — ranked by edge-count
Claims (1)
claim
- Empirical characterization of conversation domains that are safe for model persona stability
Concepts (1)
concept
- Persona Stabilizationassociated_withKeeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The consistency of moral responses when repeatedly simulating the same persona, quantified by robustness R
- Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
- Design hypothesis that coarse-grained task switching (at commands only) eliminates need for protection mechanisms while maintaining usability.
- Author's interpretive conclusion from comparing filtering strategies
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Research question addressed in the experimental analysis across tasks and persona types
- Author's mechanistic explanation for why regularization loss along persona directions is ineffective
- Argument that RL meets the agency indicator.