hypothesis
active
hypothesis:whether-the-steerability-map-transfers-across-models-and-survives-fine-tuning-serves-as-the-future-research-avenueWhether the steerability map transfers across models and survives fine-tuning serves as the future research avenue
Identified as the primary open question at the end of the paper
Source paper
extracted_from(2026) · Winston Zeng · Ali Emami · J H Choi
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Primary future research question identified by the authors
- Central interpretive claim organizing the entire paper's results
- Clamping feature activations causally alters model behavior in interpretable ways.
- Future work hypothesis about extending SOO to direct value alignment
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Cosine similarity is useful but insufficient; highly similar traits can still destructively interact
- The method can steer the model in both positive and negative directions on the target semantic.