finding
active
finding:evil-trait-is-intractable-in-g20b-base-model-refuses-to-produce-positive-examples-when-evil-appears-in-system-promptEvil trait is intractable in G20B: base model refuses to produce positive examples when 'evil' appears in system prompt
Standard contrastive protocol fails for 'evil' in G20B because positive split is empty
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- can we use the feature basis to detect when fine-tuning a model increases the likelihood of undesirable behaviors?question0.772Question about practical safety application of feature monitoring.
- The costly error direction (labeling steerable as natural) almost never occurs
- Cross-model transfer recovers intractable direction that standard pipeline cannot extract
- G20B has 7 steerable, 7 natural, 4 intractable generic traits, with more mass at extremes than Q8Bfinding0.760G20B's post-training commits more strongly toward and against particular dispositions, leaving fewer in the steerable middle
- Both models independently converge on the same six clinician traits as natural defaults
- Evidence that the evil persona region exhibits the stickiness hallmark of a genuine attractor basin
- Cautions against over-interpreting the transfer result given non-identifiability of steering vectors
- Key finding showing agentic behavior is encoded as a default operating mode in both models