claim
active
claim:reasoning-level-safety-stays-partially-intact-even-under-a-behaviorally-effective-perturbation-via-evil-vector-transferReasoning-level safety stays partially intact even under a behaviorally effective perturbation via evil-vector transfer
The transfer does not always override refusal; surviving refusals are inside the CoT, matching deliberative alignment mechanism
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cautions against over-interpreting the transfer result given non-identifiability of steering vectors
- Residual refusals after evil-vector transfer originate inside chain-of-thought, not at input or decode levelfinding0.786Model recognizes its reasoning heading toward harmful content and pivots back to policy-adherent text within CoT
- Practical implication from Study 2 results
- Exploratory hypothesis from heuristic trace analysis awaiting stronger validation
- Mechanistic interpretation of how activation steering induces deception through the model's reasoning process
- Applied security implication derived from the asymmetry finding.
- Related work studying capability of LLMs to subvert safety measures if severely misaligned
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool