finding
active
finding:dataset-level-projection-difference-predicts-finetuning-shift-for-evil-on-qwen-with-r-0-839-p-0-001-and-on-llama-with-r-0-953-p-0-001Dataset-level projection difference predicts finetuning shift for evil on Qwen with r=0.839 (p<0.001) and on Llama with r=0.953 (p<0.001)
Quantitative pre-finetuning predictability for evil trait
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Quantitative result for evil trait showing persona vector prediction power on both model architectures
- Enables pre-finetuning data screening; Figure 8 shows strong dataset-level correlations across all three traits
- Demonstrates that persona vectors capture trait-specific signal beyond general misalignment signal
- Core empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts
- Internal-state drift generalizes across scales; normalized drift also increases significantly with log(model size)
- Likely-trained MM probe is a surprisingly effective causal baseline due to correlation between truth and probability on sp_en_trans
- Justifies the use of projection difference metric rather than simpler raw projection for data screening
- Quantitative vulnerability profile for Llama-3.1-8B showing AS dominance