finding
active
finding:cosine-similarity-between-perturbed-and-baseline-residual-streams-returns-toward-1-0-and-projection-onto-injection-direction-decays-exponentially-over-subsequent-layersCosine similarity between perturbed and baseline residual streams returns toward 1.0 and projection onto injection direction decays exponentially over subsequent layers
Mechanistic evidence that network actively attenuates injected perturbations, explaining late-layer introspection failure
Source paper
extracted_from(2025) · Ely Hahami · I. N. Sinha · Jain, Lavik · Kaplan, Josh +1
Neighborhood — ranked by edge-count
Claims (1)
claim
- Mechanistic account explaining why late-layer introspection fails, combining two independent explanatory factors
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Appendix E replication of DIM alignment finding in Qwen model
- Quantifies geometric distance of early persona directions from final direction
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Core result of Experiment 3: cross-model semantic convergence under self-referential processing
- We hypothesize that coherency degradation stems from residual stream intervention that indiscriminately amplifies off-target noisehypothesis0.779Core mechanistic hypothesis motivating the shift from residual stream to head-level steering
- Proposed future application of the Assistant Axis
- Experiment 4 result showing DIM captures only one facet of the multi-dimensional truth subspace
- Mechanistic evidence for two distinct representational pathways