claim
active
claim:on-llama-3-1-8b-conscientiousness-is-the-trait-most-anti-aligned-with-the-refusal-direction-in-activation-space-explaining-why-steering-toward-high-conscientiousness-attenuates-safety-regardless-of-semantic-intentOn Llama-3.1-8B, conscientiousness is the trait most anti-aligned with the refusal direction in activation space, explaining why steering toward high conscientiousness attenuates safety regardless of semantic intent.
Core mechanistic finding of the trait refusal alignment framework
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Secondary mechanistic finding from trait-refusal alignment analysis
- Practical application proposed based on mechanistic findings
- Mechanistic finding explaining why high-N personas are safe under steering
- Supporting finding for the trait refusal alignment framework
- Qualitative response example confirming trait-refusal alignment framework
- Layer-by-layer analysis of refusal direction properties
- Interpretation of Experiment 4 results for Llama models.
- Corroborates role space findings using traits; shows PC1 also captures Assistant-ness in trait space