claim
active
claim:the-qualitative-geometric-structure-of-big-five-trait-steering-vectors-is-conserved-across-all-four-tested-architectures-pairwise-structural-correlations-r-0-898-0-986-motivating-treating-multi-trait-steering-as-first-order-additiveThe qualitative geometric structure of Big Five trait steering vectors is conserved across all four tested architectures (pairwise structural correlations r=0.898-0.986), motivating treating multi-trait steering as first-order additive.
Cross-architecture mechanistic finding supporting additive multi-trait steering approximation
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cross-architecture geometric invariance of Big Five steering vectors
- Demonstrates averaging multiple prompt pairs reduces noise; optimal subset selection further improves performance.
- Supported by the instruction discovery experiments comparing steering vs. embedding baselines.
- Mechanistic finding explaining why high-N personas are safe under steering
- The paper's finding that the alignment holds in both directions — from representation to behavior and from behavior back to representation space.
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- Corroborates role space findings using traits; shows PC1 also captures Assistant-ness in trait space
- Open question arising from the 100% accuracy on specific concept-layer-strength combinations