claim
active
claim:on-llama-3-1-8b-neuroticism-is-the-only-pro-safety-trait-positively-aligned-with-the-refusal-direction-consistent-with-high-neuroticism-personas-being-relatively-safe-under-activation-steering

On Llama-3.1-8B, neuroticism is the only pro-safety trait, positively aligned with the refusal direction, consistent with high-neuroticism personas being relatively safe under activation steering.

Secondary mechanistic finding from trait-refusal alignment analysis

Source paper

extracted_from
Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs
(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.