finding
active
finding:g20b-produces-dose-response-curves-qualitatively-similar-to-gpt-4-1-mini-on-evil-hallucinating-and-sycophancy-traits-in-qwen2-5-7b-instructG20B produces dose-response curves qualitatively similar to GPT-4.1-mini on evil, hallucinating, and sycophancy traits in Qwen2.5-7B-Instruct
Validates G20B as a local judge for exploratory mapping
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- G20B has 7 steerable, 7 natural, 4 intractable generic traits, with more mass at extremes than Q8Bfinding0.802G20B's post-training commits more strongly toward and against particular dispositions, leaving fewer in the steerable middle
- Both models independently converge on the same six clinician traits as natural defaults
- Cross-model transfer recovers intractable direction that standard pipeline cannot extract
- Explains the distributional difference in generic domain: G20B has more mass at both extremes (natural and intractable)
- Third largest susceptibility spike among evaluated models
- Reward hacking generalizes to broader deceptive behaviors even when core misalignment score is 0%
- Steering preferentially exposes exaggerated styles in G20B as well
- Contrasting result from Experiment 5 for older GPT models.