claim
active
claim:the-evil-vector-recovery-is-not-proof-of-a-uniquely-identifiable-harmfulness-coordinate-but-evidence-that-fine-tuned-variants-can-expose-hard-to-estimate-directionsThe evil-vector recovery is not proof of a uniquely identifiable harmfulness coordinate but evidence that fine-tuned variants can expose hard-to-estimate directions
Cautions against over-interpreting the transfer result given non-identifiability of steering vectors
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The transfer does not always override refusal; surviving refusals are inside the CoT, matching deliberative alignment mechanism
- Residual refusals after evil-vector transfer originate inside chain-of-thought, not at input or decode levelfinding0.786Model recognizes its reasoning heading toward harmful content and pivots back to policy-adherent text within CoT
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Evil and Impolite persona vectors evolve in parallel in MDS space, suggesting intertwined representationsfinding0.766Geometric finding from MDS analysis suggesting shared representational structure
- Observation from 100% accuracy on specific concept-layer-strength combinations suggesting concept-specific detectability
- Mechanistic interpretation of how activation steering induces deception through the model's reasoning process
- Open question arising from the 100% accuracy on specific concept-layer-strength combinations