claim
active
claim:on-llama-3-1-8b-neuroticism-is-the-only-pro-safety-trait-positively-aligned-with-the-refusal-direction-consistent-with-high-neuroticism-personas-being-relatively-safe-under-activation-steeringOn Llama-3.1-8B, neuroticism is the only pro-safety trait, positively aligned with the refusal direction, consistent with high-neuroticism personas being relatively safe under activation steering.
Secondary mechanistic finding from trait-refusal alignment analysis
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic finding explaining why high-N personas are safe under steering
- Core mechanistic finding of the trait refusal alignment framework
- Layer-by-layer analysis of refusal direction properties
- Demonstrates the paradox magnitude relative to semantically dangerous personas
- Model-specific difference in persona susceptibility
- Universal trait risk finding for low conscientiousness under prompt-based evaluation
- Qualitative failure mode difference between architectures under activation steering
- Practical application proposed based on mechanistic findings