finding
active
finding:theorem-2-transformers-with-randomly-independently-initialized-continuous-distribution-weights-are-almost-surely-injective-at-initialisation-up-to-each-layerTheorem 2: Transformers with randomly independently initialized continuous distribution weights are almost surely injective at initialisation up to each layer
Supports input-injectivity assumption for transformers at initialisation
Source paper
extracted_from(2025) · Sutter, Denis · Minder, Julian · Hofmann, Thomas · Pimentel, Tiago
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Input-InjectivitysupportsAssumption that DNN layers preserve input information by being injective; key condition for Theorem 1
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Transformers almost surely maintain input-injectivity throughout training, not just at initialisationhypothesis0.790Conjecture supported by Nikolaou et al. 2025 for last-token hidden states
- Interpretive claim connecting exponential path combinatorics to Lindsey's layer-dependent findings.
- Strong claim that inference stage structure is architectural rather than learned
- Hypothesis replicated from Bansal et al. and Anil et al. and further investigated with norm ablations
- Establishes that stages of inference are beneficial even when repeatedly applied in recurrent depth
- Evidence that stages of inference emerge without training biases from retrofitting, recurrence scheduling, or multi-recurrence losses
- Claim formalizing the Anima Labs idea that transformers are effectively recurrent due to K/V stream.
- Antra's foundational claim about how introspection arises computationally rather than from memorised text.