finding
active
finding:neuroticism-steering-vector-has-mean-cosine-alignment-of-0-079-with-the-refusal-direction-on-llama-3-1-8b-the-only-pro-safety-trait-among-all-five-ocean-dimensionsNeuroticism steering vector has mean cosine alignment of +0.079 with the refusal direction on Llama-3.1-8B, the only pro-safety trait among all five OCEAN dimensions.
Mechanistic finding explaining why high-N personas are safe under steering
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Supporting finding for the trait refusal alignment framework
- Secondary mechanistic finding from trait-refusal alignment analysis
- Core mechanistic finding of the trait refusal alignment framework
- Practical application proposed based on mechanistic findings
- Cross-architecture mechanistic finding supporting additive multi-trait steering approximation
- Qualitative response example confirming trait-refusal alignment framework
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.793Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Evidence of a bottleneck between richer internal variation and final report distribution in impulsivity→interest condition