finding
active
finding:cross-trait-correlation-baselines-for-finetuning-shift-prediction-are-r-0-34-0-86-lower-than-within-trait-correlations-r-0-76-0-97Cross-trait correlation baselines for finetuning shift prediction are r=0.34–0.86, lower than within-trait correlations r=0.76–0.97
Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts
- Quantitative pre-finetuning predictability for evil trait
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Enables pre-finetuning data screening; Figure 8 shows strong dataset-level correlations across all three traits
- Justifies the use of projection difference metric rather than simpler raw projection for data screening
- Cross-architecture geometric invariance of Big Five steering vectors
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.756Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Demonstrates robustness of the trichotomy classification to cutoff choice