finding
active
finding:pre-norm-without-input-injection-reaches-a-degenerate-fixed-point-where-all-layers-converge-to-the-same-pointPre-norm without input injection reaches a degenerate fixed point where all layers converge to the same point
Pre-norm model reaches a fixed point without input injection but all layers converge to identical representations
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Replicates and extends prior findings on input injection; tested on randomly initialized 12-layer models across three norm structures
- Limitation identified by authors: empirical results established but analytical explanation lacking
- Hypothesis replicated from Bansal et al. and Anil et al. and further investigated with norm ablations
- Interpretation that pattern density from pretraining determines few-shot requirements
- Hypothesis: Shot midpoint ordering k50(B10) < k50(B8) ≈ k50(B9) follows pretraining exposure densityhypothesis0.753E2 prediction that bases with higher pretraining exposure require fewer shots to cross threshold
- Distinguishes Huginn's convergence behavior from the ideal cyclic fixed point behavior
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Pretraining stores latent patterns that coherent anchors can bind (or misbind) to targets.quote0.735Load-bearing quote capturing the core metaphor