claim
active
claim:regularization-based-prevention-of-persona-shifts-fails-because-the-model-encodes-the-trait-through-alternative-directions-in-activation-space-to-satisfy-the-next-token-prediction-lossRegularization-based prevention of persona shifts fails because the model encodes the trait through alternative directions in activation space to satisfy the next-token prediction loss
Author's mechanistic explanation for why regularization loss along persona directions is ineffective
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Negative result showing model bypasses regularization by encoding trait through alternative directions
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Core interpretive claim providing mechanistic explanation for early persona formation
- Practical implication drawn from the prosocial persona paradox finding
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
- Central thesis of the paper