question
active
question:do-other-high-level-operators-such-as-reasoning-style-or-metacognitive-framing-likewise-localize-to-specific-attention-headsDo other high-level operators such as reasoning style or metacognitive framing likewise localize to specific attention heads?
Generalizes the Style Modulation Head finding to broader abstract computations
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Structural finding about which attention heads control reflection behavior
- Interesting special case of copying behavior related to tokenization artifacts; primitive precursor to induction heads
- Claim supported by VPD's recovery of cross-head attention subcomponents, noted in footnote.
- Architectural rationale for why Style Modulation Heads are in attention layers rather than MLPs
- Extension of superposition hypothesis to attention layers as future research direction
- Concrete example from examining expanded QK/OV matrices showing how specific programming language structure is encoded in attention weights
- Striking mechanistic finding that injection creates universally detectable perturbation in residual stream immediately downstream
- Empirical observation from examining expanded OV/QK matrices; approximately 10 out of 12 heads show significant copying