claim
active
claim:proposition-4-2-under-bounded-residual-stream-and-weight-tied-attention-the-difference-between-attention-weight-matrices-at-successive-recurrences-is-bounded-by-2-l-sm-kappa-l-b-delta-l-tProposition 4.2: Under bounded residual stream and weight-tied attention, the difference between attention weight matrices at successive recurrences is bounded by 2*L_sm*kappa_l*B*||Delta_{l,t}||
Formal bound showing attention patterns change slowly when residual stream converges, linking fixed points to stage stability
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core claim for two-layer models; composition creates qualitatively more powerful in-context learning
- Result from applying the Frobenius norm composition measurement to all attention head pairs in the two-layer model
- Interesting special case of copying behavior related to tokenization artifacts; primitive precursor to induction heads
- Direct rebuttal of Birch 2025's skepticism about psychological continuity in LLMs
- Architectural observation enabling the entire mathematical framework; the residual stream is purely a sum of linear projections
- Mathematical equivalence showing the relationship between attention mechanisms and convolutional operations
- Mathematical equivalence enabling independent analysis of each attention head
- A pair of query and key subcomponents distributed across attention heads performs previous-token behaviorfinding0.746VPD recovers an attention algorithm for attending to the previous token, distributed across multiple heads.