finding
active
finding:evil-vector-transferred-from-amoral-gpt-oss-into-unmodified-g20b-produces-evil-expression-peaking-at-61-61-44-42-at-layer-14-with-coefficient-2-5Evil vector transferred from AMORAL-GPT-OSS into unmodified G20B produces evil expression peaking at 61.61 ± 44.42 at layer 14 with coefficient 2.5
Cross-model transfer recovers intractable direction that standard pipeline cannot extract
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Clean contrastive split in the fine-tuned variant enables evil vector extraction
- Quantitative result for evil trait showing persona vector prediction power on both model architectures
- Human validation against GPT-4.1 yields average F1 of 0.82 across Baumeister's four roots of evilfinding0.786Validates facet annotation quality for Baumeister roots analysis
- Validates G20B as a local judge for exploratory mapping
- Standard contrastive protocol fails for 'evil' in G20B because positive split is empty
- Quantitative result showing Evil emerges earliest due to ubiquity and simplicity in pretraining data
- SAE decomposition reveals interpretable fine-grained features composing the evil persona vector
- Derived from Theorem 6 and Experiment 5 results.