concept
active
concept:scaling-laws-for-activation-steering-with-llama-2-models-and-refusal-mechanisms-ali-et-al-2025Scaling Laws for Activation Steering with Llama 2 Models and Refusal Mechanisms (Ali et al., 2025)
Related work finding larger models more resistant to steering, potentially consistent with ESR in 70B
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Layer-by-layer analysis of refusal direction properties
- Interpretation of Experiment 4 results for Llama models.
- Foundational paper introducing activation steering methodology used in this work
- Illustrative finding that ESR mitigates but does not fully eliminate steering influence
- Secondary mechanistic finding from trait-refusal alignment analysis
- Practical implication drawn from the prosocial persona paradox finding
- Practical disadvantage of activation steering highlighted as a drawback
- Core empirical result demonstrating that manifold steering produces on-target, behavior-aligned outputs.