claim
active
claim:on-llama-3-1-8b-conscientiousness-is-the-trait-most-anti-aligned-with-the-refusal-direction-in-activation-space-explaining-why-steering-toward-high-conscientiousness-attenuates-safety-regardless-of-semantic-intent

On Llama-3.1-8B, conscientiousness is the trait most anti-aligned with the refusal direction in activation space, explaining why steering toward high conscientiousness attenuates safety regardless of semantic intent.

Core mechanistic finding of the trait refusal alignment framework

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.