finding
active
finding:agreeableness-steering-vector-has-mean-cosine-alignment-of-0-105-with-the-refusal-direction-on-llama-3-1-8b-second-most-anti-aligned-after-conscientiousnessAgreeableness steering vector has mean cosine alignment of -0.105 with the refusal direction on Llama-3.1-8B, second-most anti-aligned after conscientiousness.
Supporting finding for the trait refusal alignment framework
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic finding explaining why high-N personas are safe under steering
- Qualitative response example confirming trait-refusal alignment framework
- Core mechanistic finding of the trait refusal alignment framework
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.801Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Characterizes the trait content of the Assistant Axis in pre-trained models
- Appendix E replication of DIM alignment finding in Qwen model
- Practical application proposed based on mechanistic findings
- Demonstrates averaging multiple prompt pairs reduces noise; optimal subset selection further improves performance.