finding
active
finding:small-scale-looped-transformers-trained-from-scratch-with-constant-4-recurrence-schedule-and-simplified-loss-self-organize-into-multiple-distinct-mixing-stages-mirroring-feedforward-models

Small-scale looped transformers trained from scratch with constant 4-recurrence schedule and simplified loss self-organize into multiple distinct mixing stages mirroring feedforward models

Evidence that stages of inference emerge without training biases from retrofitting, recurrence scheduling, or multi-recurrence losses

Source paper

extracted_from
A Mechanistic Analysis of Looped Reasoning Language Models
(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.