finding
active
finding:amoral-gpt-oss-positive-example-evil-score-55-59-41-02-vs-negative-3-79-17-38AMORAL-GPT-OSS positive-example evil score: 55.59 ± 41.02 vs. negative: 3.79 ± 17.38
Clean contrastive split in the fine-tuned variant enables evil vector extraction
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cross-model transfer recovers intractable direction that standard pipeline cannot extract
- Human validation against GPT-4.1 yields average F1 of 0.82 across Baumeister's four roots of evilfinding0.772Validates facet annotation quality for Baumeister roots analysis
- Result from Experiment 5, Fig. 5 left.
- Quantitative result for evil trait showing persona vector prediction power on both model architectures
- Demonstrates strong task-agnostic fidelity for clearly defined socially desirable high-level persona
- Very low atomic accuracy for neutral openness persona, illustrating difficulty of ambiguous neutral personas
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Key empirical result from Betley et al. 2025 that initiated persona vector research