finding
active
finding:a-bifurcation-in-the-miniaturized-looped-transformer-occurs-at-150-000-training-steps-when-accurate-integer-linear-system-solving-first-emergesA bifurcation in the miniaturized looped transformer occurs at ~150,000 training steps, when accurate integer-linear-system solving first emerges
Pinpoints the training-time transition where fractal basins emerge.
Source paper
extracted_from(2026) · Jeffrey Lai · Anthony Bao · J. Quinn · William Gilpin
Neighborhood — ranked by edge-count
Papers (1)
paper
Claims (1)
claim
- The paper's core mechanistic claim connecting saddle dynamics to basin fractality.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that stages of inference emerge without training biases from retrofitting, recurrence scheduling, or multi-recurrence losses
- Localizes transient chaos to the sub-algorithm requiring multi-step Gaussian elimination.
- Learning to encode position for transformer with continuous dynamical model (Liu et al., 2020)concept0.762Prior work on learned dynamic position encodings; cited alongside Wang et al. as precedent.
- Establishes that stages of inference are beneficial even when repeatedly applied in recurrent depth
- Transformers almost surely maintain input-injectivity throughout training, not just at initialisationhypothesis0.754Conjecture supported by Nikolaou et al. 2025 for last-token hidden states
- Strong claim that inference stage structure is architectural rather than learned
- Evidence that in-context learning is not mere pattern matching but genuine optimization, relevant to applying the thesis to inference
- Prior finding from Grant et al. 2025 used to interpret low MAS IIA for GRU-Transformer hidden state comparisons.