finding
active
finding:ouro-2-6b-recurrent-block-shows-two-distinct-half-block-segments-each-independently-aligning-with-llama-feedforward-stages-of-inference-due-to-upcycling-from-ouro-1-4bOuro 2.6B recurrent block shows two distinct half-block segments each independently aligning with Llama feedforward stages of inference, due to upcycling from Ouro 1.4B
Reveals how the upcycling training regime of Zhu et al. produces duplicated inference stage structure
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Non-fixed-point models exhibit unstable inference stages when generalizing to unseen test-time compute budgets
- Shows that retrofitting preserves base model inference stage structure in the cyclic blocks
- Empirical validation that attention patterns are most similar to same-layer outputs across different recurrences
- Key negative result showing that not all looped models reach a true fixed point, contrasting with retrofitted models
- Llama-3.3-70B exhibits internal consistency-checking mechanisms that operate during inferenceclaim0.783Central interpretive claim of the paper supported by causal ablation and activation evidence
- Quantitative argument for the richness of quasi-psychological connections enabled by attention streams
- Predictive hypothesis about domain-generality of the identified mechanism
- Localizes truth representations to specific hidden states, motivating the rest of the analysis