finding
active
finding:trait-refusal-cosine-alignment-explains-r-2-0-667-of-single-trait-as-asr-variance-on-llama-3-1-8b-p-0-004Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).
Statistical fit of the trait refusal alignment framework to single-trait activation steering results
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Practical application proposed based on mechanistic findings
- Supporting finding for the trait refusal alignment framework
- Layer-by-layer analysis of refusal direction properties
- Mechanistic finding explaining why high-N personas are safe under steering
- Core mechanistic finding of the trait refusal alignment framework
- Variance decomposition showing AS results are dominated by persona identity
- Demonstrates that RL can reinforce alignment-faking reasoning rather than eliminate it
- Statistical evidence for SP/AS ranking inversion