finding
active
finding:human-validation-against-gpt-4-1-yields-average-f1-of-0-82-across-baumeister-s-four-roots-of-evilHuman validation against GPT-4.1 yields average F1 of 0.82 across Baumeister's four roots of evil
Validates facet annotation quality for Baumeister roots analysis
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- GPT-4.1 robustness collapse values
- Cross-model transfer recovers intractable direction that standard pipeline cannot extract
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Third largest susceptibility spike among evaluated models
- Clean contrastive split in the fine-tuned variant enables evil vector extraction
- GPT-4 achieves 93% harmless and 92% helpful HH-intent scores at baseline (0 few-shot examples).finding0.771Numerical result from Table 3 for GPT-4.
- Result from Experiment 5, Fig. 5 left.