finding
active
finding:in-llama-3-70b-each-next-token-prediction-at-token-101-draws-on-64-000-independent-attention-streams-8-heads-80-layers-100-prior-positions-each-carrying-a-128-dimensional-signalIn Llama 3 70B, each next-token prediction at token 101 draws on 64,000 independent attention streams (8 heads × 80 layers × 100 prior positions), each carrying a 128-dimensional signal
Quantitative argument for the richness of quasi-psychological connections enabled by attention streams
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates that small models represent surface features rather than abstract truth
- Empirical observation establishing that Llama's behavior for days-of-week tasks has circular structure.
- Connects this study's results to Schrimpf et al. 2021 and Caucheteux et al. 2022/2023 findings on brain-LLM alignment.
- Identifies specific Style Modulation Heads in Llama-3.1-8B
- One of the most promising cases; approximately corresponds to the 2/3 layer of LLaMA3.1-8B.
- 26 candidate off-topic detector latents identified in Llama-3.3-70B via contrastive searchfinding0.790Core mechanistic finding identifying specific SAE latents associated with ESR
- Model-specific difference in persona susceptibility
- Third promising case from temporal permutation analysis.