concept
active
concept:non-identifiability-of-steering-vectorsNon-identifiability of steering vectors
Theoretical constraint proven by Venkatesh and Kurapath (2026): successful steering does not imply the vector is the unique semantic representation
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.
- Observation from 100% accuracy on specific concept-layer-strength combinations suggesting concept-specific detectability
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- Steering vectors used to reduce eval awareness can inadvertently introduce alternative user personasfinding0.744A side effect observed when applying activation steering: the model's response persona changed unexpectedly.
- General approach of using interpretability feedback to steer model generation.
- Open question arising from the 100% accuracy on specific concept-layer-strength combinations
- Supported by the instruction discovery experiments comparing steering vs. embedding baselines.
- Named procedure for classifying each trait by baseline expression and dose-response under steering