finding
active
finding:layer-dependent-introspective-peaksLayer-dependent introspective peaks
Introspective awareness in Opus 4.1 peaks at layer ~2/3 through model depth for thought injection and text distinction; prefill detection most sensitive to earlier layer, suggesting mechanistically distinct processes.
Source paper
extracted_from(2026) · Lindsey, Jack
Neighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Empirical investigation of how LMs access and report internal states across layers, using concept injection and thought detection on Claude models.
- LLM functional introspective awarenessmembers_ofEmpirical probing of language models' ability to detect and report their own internal concept representations
- Study of how self-referential representations emerge in middle layers but are actively dampened by post-training procedures like RLHF.
Questions (1)
question
- Secondary question; paper demonstrates introspection but explicitly avoids pinning down specific mechanistic explanation, noting mechanisms could be shallow and specialized.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- This paper's proposed mechanistic explanation integrating signal injection, attention routing, predictive integration, and residual recovery
- Introspective signals appear in middle layers but are suppressed by later post-training-shaped layers.finding0.783Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.
- Introspective awareness peaks at a layer about two-thirds through Opus 4.1 for injected thoughtsfinding0.765The success rate shows a sharp peak at a specific middle layer.
- The central concept: the ability of a model to access and report on its internal states, as defined by the paper's criteria.
- Key quantitative characterization of the layer-dependence of partial introspection
- Qualitative characterization of optimal anchoring depth.
- Stage 3 of character training: SFT on synthetic introspective data generated by post-distillation checkpoint
- Primary positive claim of the paper, grounded in strength comparison and localization results