hypothesis
active
hypothesis:we-hypothesize-that-coherency-degradation-stems-from-residual-stream-intervention-that-indiscriminately-amplifies-off-target-noiseWe hypothesize that coherency degradation stems from residual stream intervention that indiscriminately amplifies off-target noise
Core mechanistic hypothesis motivating the shift from residual stream to head-level steering
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic evidence that network actively attenuates injected perturbations, explaining late-layer introspection failure
- Demonstrates unique capability of Head Cor for trait suppression scenarios
- DeepSeek has large coherence loss but no robustness excess; GPT-4o has little coherence loss but large robustness drop
- Interpretive claim about the mechanistic substrate of introspection in LLMs
- Motivates adopting coherency score as the primary evaluation metric
- Rapid breakdown of output quality under strong or out-of-distribution activation steering, the main problem this paper addresses
- Motivates including Head Cor+Anti condition in steering position comparison