finding
active
finding:sp-and-as-produce-activation-cosine-similarity-of-only-0-11-0-20-at-safety-critical-layers-compared-to-0-83-0-92-for-sp-vs-fs-confirming-two-representationally-distinct-pathwaysSP and AS produce activation cosine similarity of only 0.11-0.20 at safety-critical layers, compared to 0.83-0.92 for SP vs FS, confirming two representationally distinct pathways.
Mechanistic evidence for two distinct representational pathways
Source paper
extracted_from(2026) · Wenkai Li · Fan Yang · Shaunak A. Mehta · Koichi Onoue
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Vulnerability profile for Gemma-3-27B showing SP dominance
- Evidence for two representational pathways based on cross-method activation divergence
- Confirms the prosocial paradox is not due to mismatched intervention strength
- Validates that the contrast vector method and PCA-based PC1 capture the same direction
- Vulnerability profile for Qwen3.5-9B
- Shows persona space axes are inherited from pre-training, not solely created by post-training
- Core result of Experiment 3: cross-model semantic convergence under self-referential processing
- Prior finding from related work that aligns with ESR being strongest in the largest model tested