claim
active
claim:this-is-the-first-study-to-mitigate-the-performance-trade-off-in-steering-vectors-derived-from-the-difference-in-means-approachThis is the first study to mitigate the performance trade-off in steering vectors derived from the difference-in-means approach
Priority claim establishing novelty of the Head Cor intervention approach
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates averaging multiple prompt pairs reduces noise; optimal subset selection further improves performance.
- Steering vectors used to reduce eval awareness can inadvertently introduce alternative user personasfinding0.800A side effect observed when applying activation steering: the model's response persona changed unexpectedly.
- Observation from 100% accuracy on specific concept-layer-strength combinations suggesting concept-specific detectability
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- The paper's critique of the standard linear steering baseline, supported by the days-of-week demo.
- Key asymmetry finding: suppressing reflection is easier than inducing it.
- Steering Vector Control maintains low unexpected rate of 0.08 in Experiment 1, comparable to baselinefinding0.784Shows that inducing deception via steering vectors preserves semantic coherence and does not cause random errors
- A method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.