finding
active
finding:finetuning-shift-along-persona-vectors-correlates-r-0-76-0-97-with-post-finetuning-trait-expression-across-diverse-datasetsFinetuning shift along persona vectors correlates r=0.76–0.97 with post-finetuning trait expression across diverse datasets
Core empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Quantitative result for evil trait showing persona vector prediction power on both model architectures
- Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal
- Main monitoring result showing persona vectors can predict behavioral shifts before text generation begins
- Open question posed by authors about why their method works
- Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
- Enables pre-finetuning data screening; Figure 8 shows strong dataset-level correlations across all three traits
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators
- Mechanistic explanation of how fine-tuning can shift persona vectors without directly updating activations