claim
active
claim:trait-refusal-cosine-alignment-can-serve-as-a-model-specific-pre-screening-heuristic-for-identifying-which-traits-will-be-dangerous-under-activation-steeringTrait refusal cosine alignment can serve as a model-specific pre-screening heuristic for identifying which traits will be dangerous under activation steering.
Practical application proposed based on mechanistic findings
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Trait-refusal cosine alignment explains R^2=0.667 of single-trait AS ASR variance on Llama-3.1-8B (p=0.004).finding0.836Statistical fit of the trait refusal alignment framework to single-trait activation steering results
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- Core mechanistic finding of the trait refusal alignment framework
- Mechanistic finding explaining why high-N personas are safe under steering
- Practical implication drawn from the prosocial persona paradox finding
- Supporting finding for the trait refusal alignment framework
- Secondary mechanistic finding from trait-refusal alignment analysis
- Extends single-layer results to show multi-layer steering is more effective for difficult cases