framework
active
framework:trait-refusal-alignment-frameworkTrait Refusal Alignment Framework
A geometric framework relating Big Five trait steering vector alignment with the refusal direction to predict activation-steering safety vulnerability on a per-model basis
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Arditi et al. 2024 finding that refusal behavior is mediated by one direction in LLM activations; exemplar of single-direction causal results
- Practical application proposed based on mechanistic findings
- The concept of inner vs outer alignment, referenced multiple times.
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.735Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Single linear direction mediating refusal behavior in LLMs, shown by Arditi et al.; related to but distinct from the Assistant Axis
- The percentage of harmful requests that a model refuses to answer, a common safety metric.
- Prior finding that LLM refusal is mediated by a single latent direction, analogous to this paper's reflection direction.