claim
active
claim:pairwise-steering-needs-empirical-maps-rather-than-assuming-vector-addition-will-behave-like-ordinary-semantic-blendingPairwise steering needs empirical maps rather than assuming vector addition will behave like ordinary semantic blending
Cosine similarity is useful but insufficient; highly similar traits can still destructively interact
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates averaging multiple prompt pairs reduces noise; optimal subset selection further improves performance.
- Supported by the instruction discovery experiments comparing steering vs. embedding baselines.
- The paper's critique of the standard linear steering baseline, supported by the days-of-week demo.
- The method can steer the model in both positive and negative directions on the target semantic.
- Validates that steering vectors capture reflection semantics by finding tokens reported in related work.
- The top of the steerability ranking is dominated by exaggerated or attention-grabbing styles
- Whether the steerability map transfers across models and survives fine-tuning serves as the future research avenuehypothesis0.766Identified as the primary open question at the end of the paper