finding
active
finding:small-scale-looped-transformers-trained-from-scratch-with-constant-4-recurrence-schedule-and-simplified-loss-self-organize-into-multiple-distinct-mixing-stages-mirroring-feedforward-modelsSmall-scale looped transformers trained from scratch with constant 4-recurrence schedule and simplified loss self-organize into multiple distinct mixing stages mirroring feedforward models
Evidence that stages of inference emerge without training biases from retrofitting, recurrence scheduling, or multi-recurrence losses
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Establishes that stages of inference are beneficial even when repeatedly applied in recurrent depth
- Pinpoints the training-time transition where fractal basins emerge.
- Central empirical claim of the paper supported by ColSum concentration analysis across multiple architectures
- Evidence that in-context learning is not mere pattern matching but genuine optimization, relevant to applying the thesis to inference
- Antra's foundational claim about how introspection arises computationally rather than from memorised text.
- Prior finding from Grant et al. 2025 used to interpret low MAS IIA for GRU-Transformer hidden state comparisons.
- Claim formalizing the Anima Labs idea that transformers are effectively recurrent due to K/V stream.
- Strong claim that inference stage structure is architectural rather than learned