finding
active
finding:refusal-direction-norm-on-llama-3-1-8b-increases-monotonically-from-2-5-layer-8-to-7-7-layer-24-indicating-strengthening-safety-signals-in-later-layersRefusal direction norm on Llama-3.1-8B increases monotonically from 2.5 (layer 8) to 7.7 (layer 24), indicating strengthening safety signals in later layers.
Layer-by-layer analysis of refusal direction properties
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Qualitative failure mode difference between architectures under activation steering
- Establishes generalizability of the core difficulty-boundary finding across model families.
- Secondary mechanistic finding from trait-refusal alignment analysis
- Demonstrates ESR can be deliberately enhanced through prompting in the largest model
- Scaling Laws for Activation Steering with Llama 2 Models and Refusal Mechanisms (Ali et al., 2025)concept0.801Related work finding larger models more resistant to steering, potentially consistent with ESR in 70B
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.797Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Core mechanistic finding of the trait refusal alignment framework
- Illustrative finding that ESR mitigates but does not fully eliminate steering influence