finding
active
finding:lightweight-screening-agrees-with-full-sweep-labels-on-88-5-of-g20b-traits-46-52-skipping-the-sweep-for-56-of-traitsLightweight screening agrees with full-sweep labels on 88.5% of G20B traits (46/52), skipping the sweep for 56% of traits
Elicitation-only screen achieves high accuracy with even larger compute savings for G20B
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The costly error direction (labeling steerable as natural) almost never occurs
- G20B has 7 steerable, 7 natural, 4 intractable generic traits, with more mass at extremes than Q8Bfinding0.779G20B's post-training commits more strongly toward and against particular dispositions, leaving fewer in the steerable middle
- Steering preferentially exposes exaggerated styles in G20B as well
- Explains the distributional difference in generic domain: G20B has more mass at both extremes (natural and intractable)
- Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.
- Demonstrates alignment with Linear Representation Hypothesis: target trait steers approximately linearly with alpha
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Demonstrates distributed steering is more effective and less accuracy-damaging than concentrated steering.