finding
active
finding:retrofitted-llama-attention-patterns-converge-after-the-first-recurrence-and-huginn-0125-immediately-following-the-preludeRetrofitted Llama attention patterns converge after the first recurrence and Huginn-0125 immediately following the prelude
Demonstrates remarkably fast convergence to cyclic fixed point behavior in retrofitted models
Source paper
extracted_from(2026) · Hugh Blayney · Álvaro Arroyo · Johan Obando-Ceron · Pablo Samuel Castro +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Shows that retrofitting preserves base model inference stage structure in the cyclic blocks
- Models with fixed-point convergence maintain stable inference stages at arbitrary test-time recurrence depths
- Empirical validation that attention patterns are most similar to same-layer outputs across different recurrences
- Layer-wise localization result identifying the persona-emergent attention layer in Llama-3.1-8B
- Quantitative argument for the richness of quasi-psychological connections enabled by attention streams
- Llama 3.2 1B with retrofitted recurrence; reaches fixed point and exhibits clear stages of inference
- Llama-3.3-70B exhibits internal consistency-checking mechanisms that operate during inferenceclaim0.763Central interpretive claim of the paper supported by causal ablation and activation evidence
- Causal evidence that massive activations are required for stages of inference to emerge in looped models