finding
active
finding:train-time-regularization-penalizing-projection-changes-along-persona-directions-is-ineffective-at-preventing-persona-shiftsTrain-time regularization penalizing projection changes along persona directions is ineffective at preventing persona shifts
Negative result showing model bypasses regularization by encoding trait through alternative directions
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Author's mechanistic explanation for why regularization loss along persona directions is ineffective
- Practical safety question about intervening on persona representations during pretraining
- Answers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Main monitoring result showing persona vectors can predict behavioral shifts before text generation begins
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
- Authors' interpretive endorsement of PSM view, backed by transfer experiments