question
active
question:why-analytically-do-certain-architectural-choices-input-injection-pre-norm-lead-to-stable-limiting-behavior-in-looped-transformerswhy analytically do certain architectural choices (input injection, pre-norm) lead to stable limiting behavior in looped transformers?
Limitation identified by authors: empirical results established but analytical explanation lacking
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- is stable fixed-point limiting behavior desirable or restrictive for reasoning tasks in looped transformers?question0.815Open question about whether convergence to fixed points helps or hurts reasoning performance
- Hypothesis replicated from Bansal et al. and Anil et al. and further investigated with norm ablations
- Pre-norm model reaches a fixed point without input injection but all layers converge to identical representations
- Replicates and extends prior findings on input injection; tested on randomly initialized 12-layer models across three norm structures
- Left to future work after demonstrating these behaviors are rare but not explaining their mechanism
- Transformers almost surely maintain input-injectivity throughout training, not just at initialisationhypothesis0.757Conjecture supported by Nikolaou et al. 2025 for last-token hidden states
- Evidence that stages of inference emerge without training biases from retrofitting, recurrence scheduling, or multi-recurrence losses
- Interpretive claim connecting exponential path combinatorics to Lindsey's layer-dependent findings.