claim
active
claim:regularization-based-prevention-of-persona-shifts-fails-because-the-model-encodes-the-trait-through-alternative-directions-in-activation-space-to-satisfy-the-next-token-prediction-loss

Regularization-based prevention of persona shifts fails because the model encodes the trait through alternative directions in activation space to satisfy the next-token prediction loss

Author's mechanistic explanation for why regularization loss along persona directions is ineffective

Source paper

extracted_from
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.