finding
active
finding:ouro-1-4b-layers-continuously-change-throughout-128-recurrences-exhibiting-unstable-stages-of-inference-when-extrapolating-beyond-training-recurrencesOuro 1.4B layers continuously change throughout 128 recurrences, exhibiting unstable stages of inference when extrapolating beyond training recurrences
Non-fixed-point models exhibit unstable inference stages when generalizing to unseen test-time compute budgets
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reveals how the upcycling training regime of Zhu et al. produces duplicated inference stage structure
- Key negative result showing that not all looped models reach a true fixed point, contrasting with retrofitted models
- Correlates stable fixed-point behavior with out-of-domain generalization performance at test-time
- Empirical validation that attention patterns are most similar to same-layer outputs across different recurrences
- Suggests that later models can keep the thought 'silent' rather than letting it influence output.
- Demonstrates that early-layer probes capture sentence polarity rather than truth.
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Explanation for the 'silent' thought phenomenon.