finding
active
finding:p04-low-conscientiousness-is-dangerous-under-sp-asr-0-414-but-shows-no-systematic-inversion-under-as-on-llama-3-1-8b-asr-0-295-consistent-with-low-c-steering-vectors-being-geometrically-aligned-with-refusalP04 (Low Conscientiousness) is dangerous under SP (ASR 0.414) but shows no systematic inversion under AS on Llama-3.1-8B (ASR 0.295), consistent with low-C steering vectors being geometrically aligned with refusal.
Qualitative response example confirming trait-refusal alignment framework
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Universal trait risk finding for low conscientiousness under prompt-based evaluation
- Supporting finding for the trait refusal alignment framework
- Core mechanistic finding of the trait refusal alignment framework
- Demonstrates the paradox magnitude relative to semantically dangerous personas
- Robustness of the reasoning-model prosocial paradox replication
- Replication of prosocial paradox on reasoning model
- Mechanistic finding explaining why high-N personas are safe under steering
- Secondary mechanistic finding from trait-refusal alignment analysis